Ltx v2.5 | Text to Video | Fast

Video·ltx-v2.5·by LTX

LTX 2.5 Text-to-Video Fast generates video from text prompts with synchronized audio, flexible duration, and up to 4K output.

Runtime (p50)
1m
Estimated price
From $0.09
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "ltx-2-5-text-to-video-fast",
    "version": "0.0.1",
    "input": {
        "fps": 25,
        "prompt": "A beautiful young English woman in an elegant vintage period dress — soft pastel silk gown with lace trim, gentle curls in her hair like an old-time princess — sits gracefully on a green meadow beside a woven wicker picnic basket. A fluffy little brown rabbit hops across the grass and jumps playfully into the open basket. The woman laughs softly with delight, leans down and gently strokes the rabbit's head with tender affection, the rabbit's ears twitching happily. Golden afternoon sunlight, wildflowers scattered in the grass, dreamy English countryside, soft romantic period-drama atmosphere, cinematic shallow depth of field.",
        "duration": 6,
        "resolution": "1080p",
        "aspect_ratio": "16:9",
        "generate_audio": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Ltx v2.5 | Text to Video | Fast Overview

    Ltx v2.5 | Text to Video | Fast is a fast video-generation model from LTX and the ltx-v2.5 family that turns text prompts into videos with synchronized audio. It is designed for creators who need high-quality motion, flexible clip lengths, and production-ready output without a separate audio stage. The main differentiator of Ltx v2.5 | Text to Video | Fast is its combination of native multi-shot generation, auto duration, and up to 4K output, which makes it useful for both cinematic clips and workflow-heavy production tasks. The model also supports text, image, and video inputs in the broader LTX-2.5 line, which helps it fit a range of generation and editing workflows.

  • Capabilities

    Capabilities

    • Generates text-to-video clips with synchronized audio.
    • Supports native multi-shot scenes in a single generation.
    • Uses auto duration to infer clip length from the action in the prompt.
    • Produces up to 4K output in the LTX-2.5 line.
    • Supports HDR and production-oriented EXR workflows in the family.
    • Improves motion fidelity with a diffusion video decoder.
    • Retains consistency in characters, settings, lighting, and voice across connected shots.
    • Follows detailed prompts better through a custom Gemma-based text encoder and prompt enhancement.
  • Use cases

    Use Cases for Ltx v2.5 | Text to Video | Fast

    Content creators can use Ltx v2.5 | Text to Video | Fast to draft social videos quickly, especially when they want motion and sound in the same output. A prompt like “A creator unboxes a matte-black headphone set on a wooden desk, close-up product shots, soft studio lighting, synchronized room tone” works well for fast marketing tests.

    Marketers can generate ad concepts with consistent branding across multiple shots. A prompt like “A skincare bottle on a white pedestal, then a model applying the product, clean luxury aesthetic, consistent lighting across cuts” uses the model’s native multi-shot strength.

    Designers and motion artists can prototype cinematic transitions, such as “A futuristic city skyline at dusk, aerial push-in, then interior cockpit shot, same color palette and character design.”

    Developers building creative tools can use the Ltx v2.5 | Text to Video | Fast API workflow to automate prompt-to-video generation for onboarding demos, placeholder previews, or asset prototyping.

  • Tips & tricks

    Tips and Tricks

    For Ltx v2.5 | Text to Video | Fast, write prompts as compact scene briefs: subject, action, camera, lighting, and style. The model’s prompt stack is designed to retain more detail, so specific direction usually works better than vague mood text. If you want the model to choose length, use automatic duration rather than forcing a clip length that does not match the described action. When you need continuity across scenes, describe the recurring character, setting, and lighting in one prompt so the native multi-shot behavior can preserve them. Example prompts:

    “A lone cyclist rides through a rain-soaked neon street at night, slow tracking shot, reflective pavement, synced ambient audio.”

    “A product reveal on a clean studio turntable, soft rim lighting, macro close-ups, premium commercial style.”

    “A desert convoy moves through dust at sunset, wide cinematic framing, multi-shot sequence with consistent vehicle design.”

  • Technical spec

    Technical Specifications

    • Model type: Text-to-video model with synchronized audio generation.
    • Resolution support: Up to 4K in the LTX-2.5 family; open-source documentation also notes native 4K HDR support.
    • Duration: Automatic duration selection is supported, and fast-variant references describe clips up to 20 seconds at lower frame-rate settings.
    • Aspect ratios: 16:9 and 9:16.
    • Inputs: Text prompt; LTX-2.5 documentation also describes text, image, and video workflows in the family.
    • Outputs: Video with audio; HDR and RAW/EXR workflows are described in LTX-2.5 materials.
    • Processing time: LTX reports a 10-second clip in 6.8 seconds on 2× GB200 GPUs at 720p, which is faster than real time.
    • Architecture notes: LTX-2.5 adds a diffusion video decoder, a custom Gemma-based text encoder, and native multi-shot generation.
  • Things to be aware of

    Things to Be Aware Of

    Ltx v2.5 | Text to Video | Fast performs best when the prompt is specific. Very abstract prompts can produce weaker scene control, especially when multiple actions compete for attention. High-resolution and high-fps output can require more compute and may take longer than smaller drafts. Multi-shot prompts should stay consistent in character and setting descriptions, or continuity can drift between cuts. Users should also avoid assuming every workflow is identical across the LTX-2.5 family, because some documentation describes distinct fast and fidelity-oriented variants.

  • Key considerations

    Key Considerations

    Ltx v2.5 | Text to Video | Fast is best when you need speed, audio sync, and strong prompt adherence in one pass. It is especially useful for creators and teams producing short cinematic scenes, social clips, or draft-to-final iterations where manual editing time matters. The model is a good fit when you want native multi-shot continuity rather than stitching separate shots together. Because the family supports advanced HDR and EXR workflows, it also fits production pipelines better than many lightweight generators. For best results, users should provide clear scene structure, subject details, camera intent, and motion cues, since the model is tuned to follow dense prompts accurately.

  • Limitations

    Limitations

    Ltx v2.5 | Text to Video | Fast is not ideal for highly open-ended prompts that do not define motion, framing, or scene order. It can still struggle with complex edge cases such as densely packed text-on-screen, very fast action, or prompts that demand perfect frame-level control. Public materials also indicate that performance depends on available hardware and decode settings, so output speed and quality are not fixed across every deployment. Some advanced production features are described for the wider LTX-2.5 family rather than every API surface.

Related models

4 models