Google | Gemini Omni 1.1 Flash | Reference to Video
Gemini Omni 1.1 Flash Reference-to-Video generates short videos from text prompts with image or video references, synchronized audio, aspect ratio controls, and resolution up to 4k.
- Runtime (p50)
- 1m
- Estimated price
- Usage-based
Overview
Google | Gemini Omni 1.1 Flash | Reference to Video Overview
Google | Gemini Omni 1.1 Flash | Reference to Video is a multimodal Google text-to-video model that turns prompts, images, and short video clips into cinematic videos with synchronized audio. It belongs to the Gemini Omni Flash family, Google’s lightweight video generation line designed for fast, controllable clip creation from mixed media inputs. A key differentiator of Google | Gemini Omni 1.1 Flash | Reference to Video is its ability to use video references alongside text and images to maintain character, style, and motion consistency across shots. The model supports short clips in the 3–10 second range at resolutions from 720p up to 4K, depending on the host API. On each::labs, this model is ideal for creators and developers who want rapid, reference-driven video generation without managing low-level rendering pipelines themselves.
Capabilities
Capabilities
- Text-to-video generation from a standalone prompt, producing 3–10 second clips with synchronized audio for social, product, or cinematic content.
- Reference-to-video workflows that accept text plus one or more images and optional short video references to preserve character identity, style, and motion.
- Image-to-video “first-frame-to-video” animation, where a static image defines the initial frame and the model animates the scene.
- Video-to-video editing, transforming an input clip based on a natural language description while keeping core structure and motion.
- Multi-aspect ratio support, including 16:9 landscape and 9:16 vertical, tailored to social feeds, trailers, or story formats.
- Resolution scaling from 720p to 1080p and up to 4K in some integrations, providing pathways from draft to near-production quality.
- Audio-aware generation, where the model renders audio and video in a single pass to keep timing and atmosphere aligned.
- Configurable duration presets (e.g., 4, 6, 8, 10 seconds) and reference-following durations for video transformations.
Use cases
Use Cases for Google | Gemini Omni 1.1 Flash | Reference to Video
Google | Gemini Omni 1.1 Flash | Reference to Video suits creators, marketers, designers, and developers who need reference-driven short-form video. A content creator can turn a detailed prompt and a character illustration into a consistent social clip: “Generate a 9:16 6-second video of this mascot dancing in a neon arcade, using my character image as reference, with upbeat synth music.” Marketers can remix existing footage by feeding a product clip as reference and prompting: “Transform this 5-second product shot into a winter holiday scene with matching camera motion and gentle bells in the audio.” Designers can animate storyboards, while developers can build pipelines that pass text, images, and video into the Google | Gemini Omni 1.1 Flash | Reference to Video API via each::labs for programmatic campaign generation and A/B tested variants.
Tips & tricks
Tips and Tricks
To get the most from Google | Gemini Omni 1.1 Flash | Reference to Video, treat the text prompt as the director’s note and the visual references as your storyboard. Clearly describe subject, action, camera movement, lighting, and tone, and then use reference images or a short clip to lock identity and style. For example, combine three reference images (subject, style, setting) and a concise directive such as “cinematic dolly shot at sunset” to help the model maintain consistency. When using reference video, keep the clip short, high quality, and close to your desired motion so the model can map new content onto existing movement.
Example prompts:
- "Create a 6-second 16:9 cinematic shot of a cyberpunk woman walking through neon rain, matching the style and character in my reference images, with subtle electronic background music."
- "Using this 3-second reference video of a dancer, generate a 720p clip where the same choreography happens on a futuristic glass stage, with soft ambient lighting and orchestral music."
- "Animate this illustrated character reference into a 9:16 vertical video doing a friendly wave and smile, in a cozy living room, with gentle lo-fi background audio."
Technical spec
Technical Specifications
- Provider / Family: Google, Gemini Omni Flash (Omni 1.1 Flash; often exposed as
gemini-omni-flash-previewin APIs). - Input modalities: Text prompts, single or multiple reference images, optional reference video clips (typically up to about 3–10 seconds, depending on integration).
- Output: Short video clips with embedded audio; typical lengths 3–10 seconds.
- Resolution: Commonly 720p base output, with options for 1080p and 4K via some APIs and aggregators.
- Aspect ratios: 16:9 and 9:16 are widely supported for text-to-video and reference-guided generation.
- Max reference video length: Reference videos generally limited to around 3 seconds in some official messaging and 10 seconds in upstream APIs.
- Processing time: Typical generation windows are on the order of tens of seconds to a couple of minutes per clip, varying by duration and resolution.
- Architecture: Proprietary multimodal Gemini architecture accepting text, image, audio, and video inputs to produce video with sound.
- Provider / Family: Google, Gemini Omni Flash (Omni 1.1 Flash; often exposed as
Things to be aware of
Things to Be Aware Of
Because Google | Gemini Omni 1.1 Flash | Reference to Video is optimized for short clips, it may not suit long-form narratives without stitching multiple generations together. Reference videos typically must be short (around 3–10 seconds), and longer sources may be rejected by upstream limits. Some APIs require explicit configuration, such as setting a reference_to_video task when combining images and video; missing this can lead to validation errors. Users also report that vague prompts or low-quality references lead to inconsistent faces, branding, or motion, so high-quality images and concise, specific instructions are important. Running at higher resolutions or longer durations increases cost and latency, which matters in batch or API-driven workflows.
Key considerations
Key Considerations
Google | Gemini Omni 1.1 Flash | Reference to Video is optimized for short, high-impact clips instead of long-form video. It works best when you provide a clear text description plus strong visual references, such as character or style images, or a short reference video. Many APIs limit clips to 10 seconds and cap reference videos at about 3–10 seconds, so workflows requiring longer timelines must be chained or extended. Because it is a Google text-to-video model with audio baked into a single pass, it is ideal when you want cohesive video and sound quickly rather than running separate visual and audio pipelines. On each::labs, you should consider this model when you need multimodal control and fast iteration, while heavier models may suit long or extremely photorealistic sequences.
Limitations
Limitations
Google | Gemini Omni 1.1 Flash | Reference to Video is limited to short-duration clips, generally capped around 10 seconds per generation through current APIs, and does not replace full-length editing suites. Some integrations cap base output at 720p, with upscale paths rather than native 4K for every workflow. Reference inputs must respect strict size and duration limits, and multi-image setups may only allow specific counts or layouts. The model may struggle with highly complex multi-scene stories, exact lip-sync to arbitrary external audio, or precise replication of fine-grained typography and logos, so it is best viewed as a powerful generative starting point rather than a frame-perfect production tool.



