Google | Gemini Omni 1.1 Flash | Text to Video

Video·gemini-omni-flash·by Google

Gemini Omni 1.1 Flash Text-to-Video animates a source image into a short AI video with prompt-guided motion, synchronized audio, portrait or landscape framing, and selectable resolution up to 4k.

Runtime (p50)
1m
Estimated price
Usage-based
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "google-gemini-omni-1-1-flash-text-to-video",
    "version": "0.0.1",
    "input": {
        "prompt": "Ultra cinematic macro nature film, hyperrealistic, soft depth of field, slow motion, atmospheric, dreamlike, volumetric lighting, Fibonacci-inspired transitions, seamless morphing between natural spirals, elegant camera movement, documentary-quality realism",
        "duration": "10s",
        "aspect_ratio": "16:9"
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Google | Gemini Omni 1.1 Flash | Text to Video Overview

    Google | Gemini Omni 1.1 Flash | Text to Video is Google’s latest multimodal video generation model that turns text prompts and reference media into short, high-quality clips with synchronized audio. It is part of the Gemini Omni family from Google DeepMind, designed specifically for controllable video generation and editing. Compared with earlier Gemini Omni Flash releases, Gemini Omni 1.1 Flash adds support for higher resolutions up to 4K and more advanced scene extension controls, giving creators a faster way to prototype motion while keeping visual quality competitive with frontier text‑to‑video models. On each::labs, the Google | Gemini Omni 1.1 Flash | Text to Video model focuses on animating a single source image into short videos, guided by prompts that define motion, framing, and sound so users can move from static visuals to dynamic, share‑ready clips in a single workflow.

  • Capabilities

    Capabilities

    • Animates a single source image into short videos, preserving subject identity while adding prompt‑guided motion and camera behavior.
    • Generates text-to-video clips directly from natural language prompts, without requiring reference media.
    • Outputs synchronized native audio aligned to visual action, removing the need for a separate sound‑design step for most short clips.
    • Supports multiple resolutions, including 720p, 1080p, and up to 4K in the Omni 1.1 Flash generation path, with 16:9 and 9:16 aspect ratios for horizontal and vertical formats.
    • Accepts short reference videos (≤10 seconds) for scene extension or editing, enabling iterative refinement of existing clips with textual instructions.
    • Provides scene extension capabilities in Google’s own tools, allowing users to continue a video in 10‑second increments using up to 10 seconds of prior context.
    • Integrates safety policies that filter disallowed content, including violence‑adjacent actions and restricted real‑person scenarios, reducing risk in production use.
    • Offers a unified Google | Gemini Omni 1.1 Flash | Text to Video API across text‑to‑video, image‑to‑video, and video editing, simplifying workflow integration for developers.
  • Use cases

    Use Cases for Google | Gemini Omni 1.1 Flash | Text to Video

    For creators, Google | Gemini Omni 1.1 Flash | Text to Video is ideal for animating static character art or concept illustrations into short motion tests with preserved style and identity, using the image‑to‑video capability to validate designs before full production. Example: “Turn this comic panel into a 6‑second clip where the hero’s cape moves in the wind and the camera slowly pushes in, dramatic orchestral hit at the end.” For marketers, the higher resolution output and native audio let teams quickly generate social ads or product teasers from a single hero image, optimized for 9:16 vertical feeds. Example: “Animate this product photo into a 5‑second vertical video with a slow spin, clean white background, soft whoosh sound, brand logo at the end.”

    Developers can integrate the Google | Gemini Omni 1.1 Flash | Text to Video API into content pipelines to generate preview motion from prompts and source images, useful for automated campaign generation and A/B testing. Example: “Generate three 4‑second variations from this landing page hero image: subtle parallax, text fade‑in, and logo hover, calm ambient audio.” Designers can use short reference videos plus text overrides to iteratively extend or tweak sequences while keeping layout and motion language consistent. Example: “Extend this 10‑second UI animation by another 10 seconds with the same camera style, slower pacing, and softer audio.”

  • Tips & tricks

    Tips and Tricks

    For Google | Gemini Omni 1.1 Flash | Text to Video, prompt structure matters more than length. Effective prompts usually specify subject, action, setting, camera behavior, sound, timing, and what to exclude, in that order. When animating a source image, describe how the subject should move relative to the existing framing, rather than re‑describing the still image; this helps the model preserve identity and composition while adding motion. Keep sequences simple: one or two key actions over a 3–8 second clip tend to produce more coherent results than complex multi‑shot instructions. Avoid ambiguous camera language and instead request clear moves like “slow push‑in” or “static camera” to reduce geometry distortions.

    Example prompts:
    "Animate this portrait photo into a 6‑second video where the character turns their head toward the camera, soft studio lighting, subtle slow‑motion, gentle ambient synth music, 16:9."
    "From this landscape image, create a 10‑second Google text-to-video clip with the camera slowly gliding forward through the valley, evening golden hour, light wind on the grass, no people, 9:16 vertical."
    "Generate a 4‑second Google | Gemini Omni 1.1 Flash | Text to Video clip of the illustrated robot waving to the viewer, clean white background, static camera, cheerful electronic chime audio, 1080p."

  • Technical spec

    Technical Specifications

    • Provider / family: Google Gemini Omni 1.1 Flash video model for text-to-video and image-to-video.
    • Supported resolutions: 720p native, with options for 1080p and 4K output in the updated Omni 1.1 Flash release.
    • Max duration: Typical generation length 3–10 seconds per clip; scene extension can chain up to ~40 seconds total in Google’s own tools.
    • Aspect ratios: 16:9 and 9:16 are commonly supported for short‑form and vertical video use.
    • Inputs: Text prompt plus optional reference image or short reference video (≤10 seconds) depending on the API route.
    • Outputs: MP4 video with native synchronized audio, typically at 24 fps.
    • Processing time: Short clips usually render within tens of seconds, depending on resolution and load (varies by platform).
    • Architecture: Proprietary multimodal Gemini Omni architecture optimized for video generation and editing from text, image, audio, and video inputs.
  • Things to be aware of

    Things to Be Aware Of

    Gemini Omni Flash‑based video generation still has practical limits. Complex camera moves or aggressive angle changes can distort scene geometry, so camera prompting should stay simple and clearly defined. Motion involving many objects or multiple people speaking in the same frame remains challenging, and outputs can lose coherence when prompts demand dense action in very short clips. Text inside video, such as signage or UI copy, may render imperfectly, which is important for brand‑critical layouts. Some implementations of Omni Flash currently cap output at 720p and 3–10 seconds, and daily generation limits or cooldowns can apply depending on the host platform’s policy. Safety filters can block certain real‑person edits and contact‑based actions, so workflows should avoid borderline content to reduce failed generations.

  • Key considerations

    Key Considerations

    Google | Gemini Omni 1.1 Flash | Text to Video works best for short, high‑impact clips rather than long‑form video, so plan stories around 3–10 second segments or scene‑by‑scene chains. Users must supply a well‑structured prompt and, for image‑to‑video workflows, a high‑quality source image that clearly shows the subject, composition, and style they want to preserve. Compared with heavier cinematic models, Gemini Omni 1.1 Flash trades absolute realism for speed and controllability, making it a strong fit for rapid prototyping, social‑ready motion, and iterative editing where users want fast turnaround and integrated audio. On each::labs, cost is tied to clip duration and resolution, so shorter and 720p or 1080p outputs will typically be more cost‑efficient than repeated 4K generations.

  • Limitations

    Limitations

    Google | Gemini Omni 1.1 Flash | Text to Video is optimized for short clips, not long continuous narratives, and typical APIs restrict per‑clip duration to about 10 seconds, with scene extension handled as separate steps. Some deployments still expose only 720p output, and even when 4K upscaling is available, textures and fine detail may fall short of true cinema‑grade models. The model does not offer low‑level controls like seeds or negative prompts in standard video APIs, limiting reproducibility and precision for advanced workflows. Reference video support is constrained to short inputs and certain platforms note that multi‑video prompting or mis‑configured references can degrade results or fail outright. Safety restrictions block specific real‑world identity use cases and certain edit types, so it is not suitable for every high‑risk scenario.

Related models

4 models
* FAQ

About Google | Gemini Omni 1.1 Flash | Text to Video

01 / 03

What inputs does Gemini Omni 1.1 Flash Text-to-Video use?

This integration uses a text prompt and one source image URL. The image is used as the first frame, and the prompt describes how the scene should move or sound.