Blog/Photo Guide··9 min

Best Photos for Hotel Lobby AI: One Duo Photo or Two Portraits?

Choose better source images for Hotel Lobby AI, including when to use one photo together, when to use two portraits, and how to preview the selected scene.

The short answer

Use “Two photos” when each performer has a clean individual portrait; use “One photo together” when you already have a strong image with both people clearly visible and not heavily overlapping. In either mode, prioritize visible faces, sharp identity details, useful lighting, and subjects large enough to recognize. The source background does not need to match the final scene because you choose Hotel Lobby, Rooftop Night, Warehouse Session, or Neon Garage separately.

The current HotelLobbyAI workbench supports two valid input modes, and that changes the photo advice. You no longer need to force every project into two separate portraits. A good photo of both people together can be the better input when both faces are clear and the relationship between the two subjects is already easy to read.

The decision is not about which mode is theoretically better. It is about which source gives MiniMax H3 cleaner identity information. The goal is to reduce ambiguity before the model also has to follow a scene, generate motion, and perform a personalized verse.

Best for

  • ✓ Friends, couples, siblings, family members, coworkers, or other two-person casts
  • ✓ Users choosing between one strong duo photo and two separate portraits
  • ✓ Anyone trying to reduce identity drift before a 100-Credit video render

Not for

  • × A guarantee that any clear portrait will preserve identity perfectly in motion
  • × Three-person or larger group videos
  • × Manual face reconstruction or frame-by-frame identity correction

Choose the input mode that contains the clearest identity information

Two separate photos give the model a dedicated reference for each performer. That is useful when the best available images were taken at different times or when a shared photo contains clutter. One photo together has a different advantage: the model sees the pair in the same image and can preserve their relative look from one real moment.

Neither mode removes the difficulty of two-person video generation. The right question is simply which source makes the two identities easiest to read.

Input modeUse it whenAvoid it when
Two photosEach person has a clear individual portraitOne portrait is tiny, blurred, covered, or much weaker than the other
One photo togetherBoth faces are visible, reasonably large, and not heavily overlappingOne person blocks the other or the pair is tiny inside a group photo

Prioritize identity information over a perfect-looking source photo

A dramatic portrait is not automatically a better generation reference. A simple phone photo in even light can expose more useful identity information than an artistic image with deep shadows, motion blur, or a tiny face.

  • Keep eyes, nose, mouth, jawline, hairline, and other recognizable features visible when possible.
  • Prefer front-facing or three-quarter views over extreme profiles.
  • Use a sharp image instead of a heavily compressed screenshot.
  • Avoid very dark lighting that removes facial detail.
  • Avoid sunglasses, masks, hands, or hair covering most of the face unless those features are essential to the identity.
  • Keep both subjects large enough to identify when using one photo together.

You do not need to match the final scene in your source photo

The source image provides identity; the workbench supplies the performance setting separately. After uploading the performers, you choose Hotel Lobby, Rooftop Night, Warehouse Session, or Neon Garage. That means you do not need an orange background to make the Hotel Lobby version or a city skyline in the source photo to make Rooftop Night.

Choose photos for identity clarity first. Let the scene selector handle the environment. Matching the original background is much less important than giving the model a clear view of the people it needs to preserve.

Use the free preview to test casting and scene choice

Before sign-in, you can generate one free static 9:16 preview from the selected input mode and scene. This is particularly useful now that the product supports four looks. A pair that feels ordinary in the orange Hotel Lobby setup may work better on Rooftop Night or in the Warehouse Session, and vice versa.

The preview does not include your personalized verse and does not test H3 motion or audio. Treat it as a visual casting check: are these the right photos, do both people read clearly, and does this scene match the tone you want?

If the free still already makes one person difficult to recognize, change the input before spending Credits on the full video.

Personalization comes after photo quality

Names, occasions, milestones, and inside jokes can make the final clip more personal, but they cannot repair a weak identity reference. MiniMax M3 uses the text detail to write the short original verse; MiniMax H3 still depends on the image input to know who should perform it.

Keep the jobs separate in your head: photos answer “who is in the video,” the scene answers “where and how it looks,” and the personalization field answers “what the short verse should be about.” Better results start when each input is doing one clear job.

A quick photo checklist before the full video

CheckGood signWarning sign
PeopleExactly two intended performersExtra people competing for attention
Face visibilityBoth faces clear and readableCovered, tiny, or extreme-profile faces
SharpnessUseful facial and hair detailMotion blur or heavy compression
One-photo modeBoth people visible with limited overlapOne person blocks the other
Two-photo modeOne clear performer per fileGroup photo used as an individual reference
SceneChosen for the desired moodTrying to match the source background unnecessarily

Try the focused workflow

Preview your duo in a rap-video scene first

Use one photo together or two separate portraits, choose one of four scenes, and create one free static preview before sign-in. For the full video, add an occasion or personal detail so MiniMax M3 can write the short verse before MiniMax H3 generates the 10-second rap performance.

Frequently asked questions

Is one photo together worse than two separate photos?

Not automatically. A strong duo photo can be a good input if both faces are clear and both people are easy to distinguish. Use two portraits when they provide cleaner individual identity references.

Do the source photos need matching backgrounds or lighting?

No. The final environment comes from the selected scene. Similar image quality can help, but matching backgrounds are not required.

Does the free preview use my personalization details?

No. The free preview is a static scene preview based on the photo input and selected scene. The personalized verse is generated as part of the signed-in full video flow.