Blog/Troubleshooting··9 min

Why Do Faces Swap in Hotel Lobby AI Videos?

Learn why two-person AI videos can mix identities, how one-photo and two-photo inputs behave differently, and what to change before regenerating.

The short answer

Identity mixing happens because MiniMax H3 must preserve two people while also generating motion, the selected scene, synchronized rap vocals, and a 10-second performance. The risk rises when faces are small, covered, similar, blurred, or heavily overlapping. If one shared photo is ambiguous, switch to two clean portraits. If two separate portraits are weak, replace the weaker reference. The guided workbench can reduce setup mistakes, but it cannot guarantee perfect identity preservation in every frame.

A two-person AI video has a harder identity problem than a single-person animation. The model must know there are two distinct people, keep their features separated through changing poses, generate a coherent rap performance, and maintain the selected visual environment at the same time.

The current HotelLobbyAI workflow adds another useful troubleshooting choice: you can change not only the photo itself, but also the input mode. Sometimes the fastest fix for an ambiguous duo photo is to stop asking one image to define two identities and give the model two separate portraits instead.

Best for

  • ✓ Users who saw face drift, identity blending, or one performer absorbing features from the other
  • ✓ People deciding whether to switch from one-photo mode to two-photo mode
  • ✓ Anyone who wants realistic expectations before spending another 100 Credits

Not for

  • × A promise that every identity error can be eliminated
  • × Manual frame-by-frame face correction
  • × Advanced identity-strength sliders or masks that the current workbench does not expose

The model is doing several difficult jobs at once

For the full video, MiniMax H3 receives the source image or images, the selected scene direction, and the short verse written by MiniMax M3. It is also asked to create native synchronized audio with a modern hip-hop beat and clear rap vocals while keeping both performers visible and distinct.

That is much more demanding than producing a single still portrait. During movement, faces can turn, become smaller, overlap, or change expression. Each of those moments gives the model less stable identity information, so a performer can drift toward the other person’s features even when the first frame looked strong.

One-photo mode can fail for a different reason than two-photo mode

In “One photo together” mode, the model must separate two identities from one shared image. That can work well when both faces are clear and each person occupies a readable part of the frame. It becomes harder when one person is behind the other, the faces overlap, one subject is much smaller, or the image is really a crowded group photo.

In “Two photos” mode, each performer gets a dedicated source image, which can reduce that ambiguity. The tradeoff is that the two photos may have very different lighting, angles, age, hairstyle, or image quality. The weaker portrait can then become the source of identity instability.

Common source-photo problems that increase identity drift

The most useful fix is usually not a longer personalization message. The personalization field changes what the verse is about; it does not strengthen the facial reference. Fix the image input first.

  • A face is very small relative to the whole image.
  • The image is soft, blurry, or heavily compressed.
  • One person is hidden behind hair, sunglasses, hands, or another person.
  • The two people have very similar visible features and clothing in the chosen references.
  • One shared photo contains extra people that the model may also interpret as subjects.
  • The only clear identity cues disappear when the performer turns or moves.

What to change before a second generation

Step 1

Identify which performer is drifting

If one person is consistently weaker, replace that reference first instead of changing every variable at once.

Step 2

Switch input mode when the shared image is ambiguous

If one duo photo contains overlap or extra people, crop or replace it with two clear individual portraits.

Step 3

Use a tighter, clearer crop

Give the model more facial and upper-body detail without enlarging a tiny source until it becomes blurry.

Step 4

Keep the scene simple while troubleshooting

Use the scene that gives the pair the clearest readable composition before experimenting with a different visual mood.

Step 5

Regenerate after you change the underlying input

A new sample can differ, but changing a weak reference is more meaningful than repeatedly submitting the same ambiguous photos.

Does the Hotel Lobby scene behave differently?

Yes. The Hotel Lobby preset uses the recognizable orange performance setup together with a motion reference. Rooftop Night, Warehouse Session, and Neon Garage use their own H3 scene directions without that Hotel Lobby motion reference. That means the visual challenge is not identical across all four scenes.

If your main goal is the classic trend, keep Hotel Lobby selected and optimize the identity references around that motion. If your main goal is simply a personalized rap clip, another scene may give you a visual composition that works better for the specific pair.

Where the guided workflow stops

The current workbench does not expose face masks, per-subject identity-strength controls, manual keyframes, or frame-by-frame repair. It is designed to make the workflow simple: choose the photo mode, choose the scene, add personalization, and generate the 10-second result.

That boundary matters for recommendations. If precise identity control is more important than speed, a specialized motion-transfer or face-editing workflow may be the better category. If the goal is a quick personalized social clip, the guided workbench trades some control for a much shorter setup.

Try the focused workflow

Preview your duo in a rap-video scene first

Use one photo together or two separate portraits, choose one of four scenes, and create one free static preview before sign-in. For the full video, add an occasion or personal detail so MiniMax M3 can write the short verse before MiniMax H3 generates the 10-second rap performance.

Frequently asked questions

Should I switch from one photo together to two separate photos?

Switch when the shared photo makes the two identities hard to separate, especially if the faces overlap, one person is small, or extra people are present.

Will changing the personal details fix face drift?

Usually not. Personal details are used to write the short verse. Identity drift is more directly affected by the image references and the difficulty of the generated motion.

Are failed generations charged?

If a generation reaches a failed or canceled terminal state, the generation Credits are returned automatically. A completed video with a creative imperfection is different from a technical failure.