Input
Text or image
Start from a prompt, or upload one image to guide the opening frame.
Your video will land here
Write a prompt, then generate
Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash combines text and visual input in one video workflow. It is useful when you want to move quickly from an idea or reference image to a short clip with motion and generated sound.
The underlying model supports broader workflows such as first/last-frame interpolation, reference images, editing, and extension. This page intentionally exposes the simplest production path first: prompt plus an optional start image.
Input
Text or image
Start from a prompt, or upload one image to guide the opening frame.
Output
720p
This page uses the 720p tier for a balanced production cost.
Length
Auto · 3–10s
The provider decides clip length; there is no exact duration parameter.
Formats
16:9 · 9:16
Choose landscape or vertical output before starting generation.
Capabilities
Use it for rapid iteration when text, image context, motion, and audio all belong in the same creative brief.
Turn a concise creative brief into a short video without preparing a source image first.
Use one start image to anchor the subject or opening composition.
Describe dialogue, ambience, music, or effects in the same prompt as the visual direction.
Generate 9:16 creative drafts for Shorts, Reels, and TikTok-style placements.
Test product shots, camera ideas, and visual hooks before committing to a longer production workflow.
Use the fixed-cost page to explore several directions without managing an exact clip duration.
How it works
The workflow stays simple even though the model itself supports more advanced multimodal controls.
Describe the subject, action, framing, camera movement, lighting, mood, and any audio direction.
Upload one image when the opening composition or subject identity matters to the result.
Select 16:9 or 9:16. The model decides the final clip length automatically.
Prompt ideas
Specific camera and sound direction usually produces a clearer creative brief than a vague one-line prompt.
Cinematic
“A cyclist crosses a rain-soaked city bridge at blue hour, cinematic tracking shot, reflections on wet asphalt, soft traffic ambience and distant thunder.”
Product
“A matte-black smartwatch floats above a glass pedestal, slow orbiting camera, clean studio reflections, restrained electronic sound design.”
Portrait
“Animate the person in the reference image turning toward camera and smiling naturally, preserve identity and clothing, subtle handheld motion, quiet room ambience.”
Vertical ad
“A creator opens a package in a bright apartment, quick social-video pacing, natural reactions, crisp close-ups, vertical framing, light room sound.”
Use cases
The current page is optimized for quick, short-form generation rather than long editing sessions.
Prototype a visual hook, product moment, or opening shot before producing a full campaign asset.
Create quick vertical or landscape clips from a text idea or creator portrait.
Turn a still frame into a moving shot to test pacing and camera direction.
FAQ