Photo 1 · left
Photo 2 · rightBest friends
Two friends trade verses under the mic, made from two separate selfies.
Upload one photo of each person, or pet, and they become a rap duo in a bright orange studio, sharing one hanging microphone. Use yourself and a friend, your partner, your grandparents or your dog.
You don't need to write a prompt or film anything. Pick a length, hit generate, and download the MP4 with sound.
No photos handy? Try a sample pair:
Sign in to generate. Credits are refunded if a render fails.
Use one clear, front-facing photo per performer, and only photos you have permission to use.
The routine your duo will perform
Your two people stand in an orange studio and take turns at the mic: the one on the left points at the camera and raps first, then the one on the right takes over. The moves come from a built-in reference clip.
Each run comes out a little different. The examples below were made the same way, and show the photos that went in.
Choose your two performersMade with Froging AI
Each clip was made with Froging AI from the two photos shown. The people and pets in the photos are AI-generated too. Turn up the volume to hear the audio, or load any pair into the generator above.
Photo 1 · left
Photo 2 · rightTwo friends trade verses under the mic, made from two separate selfies.
Photo 1 · left
Photo 2 · rightGrandma and grandpa trade verses, made from two separate portraits.
Photo 1 · left
Photo 2 · rightA couple version for an anniversary post.
Photo 1 · left
Photo 2 · rightA golden retriever and a tabby cat perform the same routine.
Photo 1 · left
Photo 2 · rightTwo coworkers, for a team intro or launch post.
Photo 1 · left
Photo 2 · rightAn astronaut and a knight share the mic.
Photo 1 · left
Photo 2 · rightA mother-and-daughter duet for Mother’s Day or a birthday surprise.
Photo 1 · left
Photo 2 · rightA pug and a husky, made from two pet photos.
You pick the two people. The moves are already set.
One JPG, PNG or WebP photo per performer. They don't need to be in the same photo. You can also start from a sample pair.
Photo 1 becomes the left performer, who opens the verse; Photo 2 becomes the right. Use the swap button to change roles.
Choose 10 seconds to see both performers, 5 for just the opening verse, or 15 for the full routine. Pick 768P or 2K; the button shows the exact credit cost.
Rendering takes a few minutes. Watch it through, then download the MP4 with sound. Failed renders are refunded automatically.
Your two photos are sent to MiniMax H3 alongside a fixed orange-studio performance clip. The photos decide who is in the video. The clip decides the setting, timing and gestures, which is why every pair does the same routine.
That's also why there is no prompt box.
We made the reference clip for Froging AI: two performers in front of an orange backdrop with one microphone hanging between them, trading lines with points, an arm across the face and a shush.
Pairs that tend to work, with a caption you can post them with.
Pair the quiet friend with the one who never stops talking in the group chat. Whoever is in Photo 1 raps first. Works well as a birthday message.
Caption idea
“When the group chat finally drops a track.”
Watch the exampleGood for an anniversary post. Pick photos with the outfits you want in the video, because clothes usually carry over: see the red jacket and white shirt in our couple example.
Caption idea
“Our anniversary photos took an unexpected turn.”
Watch the exampleSend grandma and grandpa rapping to the family chat. Recent, clear photos work best, so everyone can tell who it is.
Caption idea
“Grandma took the first verse.”
Watch the exampleA dog and a cat, or two dogs, work too. Pets end up standing and gesturing like people, so it looks more cartoonish than real.
Caption idea
“They have opinions about dinner being late.”
Watch the exampleUse it for a launch post or an end-of-project thank-you. Ask both people first and use a separate portrait of each.
Caption idea
“The launch meeting got a new opening act.”
Watch the exampleOur example pairs an astronaut with a knight. Original characters or your own costume photos work as long as the face is visible.
Caption idea
“Nobody asked for this crossover.”
Watch the exampleUse two clear photos with one person in each. They can be from different days. Clothes in the photo usually show up in the video.
Crop out other people and pets. If a photo shows two faces, the model may pick the wrong one or blend them.
Front or three-quarter views with both eyes visible keep identity best. Avoid sunglasses, hands over the face and heavy shadows.
A selfie or waist-up shot works better than a distant full-body photo. Daylight or even indoor light beats a dark night photo.
For animals, pick a photo where the whole head is visible and the pet is looking toward the camera. Distinctive markings help the model keep them recognizable.
Small, blurry or side-on faces give the model too little to work with.
Try next
Use a sharper, closer photo of that person and generate again. Shorter clips usually hold identity more reliably.
Photo 1 always becomes the left performer and photo 2 the right.
Try next
Press the swap button between the two photos before generating.
Fast gestures and objects held in the hands can produce distortions.
Try next
Use a photo without handheld objects, try a shorter clip, or generate again; results vary between runs.
The reference is a human performance, so the model may give pets upright poses and human-like limbs.
Try next
Choose a face-forward photo and enjoy it as a stylized performance. The pet examples above show what to expect.
Every video comes with an AI-generated audio track, and the download keeps it. Previews on this page start muted; turn up the volume to listen.
The audio is not the original “Hotel Lobby” recording. Keep the generated sound, or add the song from your social app's licensed music library when you post.
Pay per video. 10 seconds gives both performers a turn; 2K renders sharper output.
| Length | 768P | 2K |
|---|---|---|
| 5s | 50 credits | 80 credits |
| 10s | 100 credits | 160 credits |
| 15s | 150 credits | 240 credits |
The price is shown on the Generate button before you start. If a render fails, the credits are returned automatically. See plans and credit packs.
Photos, credits, music and getting faces right.
It is a photo-to-video trend named after the song "Hotel Lobby". Two people, pets or characters stand in a bright orange studio with one microphone hanging between them and trade rap verses with big hand gestures and head nods. Despite the name, the scene is a music studio, not a hotel.
Upload one photo per performer. Froging AI sends both photos and a built-in orange-studio performance clip to MiniMax H3. Your photos decide who performs; the reference clip decides the gestures, timing and framing. No prompt or video upload is needed.
Two: one photo per performer. Photo 1 becomes the performer on the left and photo 2 the performer on the right. Clear, front-facing photos with one subject each work best. You can swap the two before generating.
A 5-second clip costs 50 credits at 768P or 80 at 2K. 10 seconds costs 100 or 160 credits, and the full 15-second routine costs 150 or 240. The Generate button always shows the price before you start.
MiniMax H3 adds its own AI-generated audio to the clip, but it is not the original song. For the real track, add it from your social app's licensed music library when you post, or put any track you have the rights to under it in a video editor.
You can try friends, couples, siblings, grandparents, pets or original characters. See the examples for generated results using each kind of pairing. For animals, choose a photo with the whole head visible and the pet facing the camera. Pets end up standing and gesturing like people, so it looks more cartoonish than real.
Only use photos of yourself, people who have agreed, or characters you own. Requests that impersonate real public figures or use someone else's photo without permission may be blocked by moderation, and you are responsible for how you share the result.
Use sharp, well-lit photos where the face is large and unobstructed: no sunglasses, hats over the eyes or heavy filters. Crop out other people. Shorter clips usually keep identity more stable.
Separate photos are what it's built for. They don't need the same background or location. Put one person in each slot. If you only have a group photo, crop it into two individual portraits before uploading so each role has one clear face.
The routine opens with the left performer, so 5 seconds shows mainly the opening verse. Choose 10 seconds to see both performers take a turn, or 15 seconds for the full routine. 2K gives sharper output than 768P and costs more credits.
Browsing the examples is free. Making your own video uses credits: 100 credits for the default 10-second 768P video, or 50 for a 5-second clip. Failed renders are refunded automatically.
Usually a few minutes; busy queues can make it longer. The page shows whether your video is queued or rendering, and finished videos also appear in History, so there's no need to submit the same pair again.
Previews start muted so several videos don't play over each other. Turn on the volume in the player. The example videos include audio, and downloads keep the audio track.
The generator outputs landscape 16:9 video. In your editor or social app, you can place it on a vertical canvas, add a background, or crop it. Preview any crop carefully: both performers and the hanging microphone need room. Use photos and music you have permission to share.