Flux 3 API

Video·flux-3·by Black Forest Labs

FLUX.3 is Black Forest Labs’ frontier text-to-video model, transforming text prompts into high-quality videos with realistic motion, cinematic composition, and fast generation.

Runtime (p50)
4m
Estimated price
From $0.06
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "flux-3-text-to-video",
    "version": "0.0.1",
    "input": {
        "prompt": "A majestic black stallion gallops through an ancient pine forest at sunrise, its mane flowing naturally in the wind as golden rays of sunlight stream through the towering trees. Soft morning mist drifts across the forest floor while leaves and dust rise gently beneath each stride. The camera begins with a close-up tracking shot beside the horse, then smoothly pulls back into a wide cinematic view revealing the endless forest. Realistic muscle movement, elegant slow-motion moments, natural breathing, authentic hoof impact, rich deep green and charcoal color palette, ultra realistic wildlife cinematography, premium nature documentary aesthetic, cinematic lighting, subtle film grain, immersive atmosphere",
        "duration": "15",
        "resolution": "hd",
        "aspect_ratio": "16:9",
        "generate_audio": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Flux 3 | Text to Video Overview

    Flux 3 | Text to Video is Black Forest Labs’ frontier text-to-video model that turns natural language prompts into short, high-quality video clips with synchronized audio. Flux 3 Video is part of the FLUX 3 multimodal foundation family, trained jointly on images, videos, audio, and action data to understand and generate rich visual motion and sound from a single prompt. Its primary differentiator is native audio generation inside the same pass as video, producing up to 20-second clips with realistic motion, cinematic composition, and coherent sound without separate audio tools. On each::labs, Flux 3 | Text to Video gives creators, marketers, and developers a powerful way to prototype sequences, storyboards, and social-ready content directly from text.

  • Capabilities

    Capabilities

    • Generates text-to-video clips up to 20 seconds with native, synchronized audio from a single prompt.
    • Supports image-to-video modes, animating a reference image into a moving clip with consistent style.
    • Handles multiple aspect ratios from vertical 9:16 to ultra-wide 21:9, plus square and standard 16:9/4:3 formats.
    • Outputs 480p and 720p video suitable for social media, web experiences, and rapid iteration workflows.
    • Understands camera motion and scene dynamics thanks to training on image, video, and audio within one backbone.
    • Can be configured with explicit duration presets (5, 10, 15, or 20 seconds) or auto duration selection.
    • Provides coherent soundtracks that match visual content, reducing the need for separate audio-generation tools.
    • Integrates into automated pipelines via the Flux 3 | Text to Video API, enabling programmatic video generation from text.
  • Use cases

    Use Cases for Flux 3 | Text to Video

    Content creators can use Flux 3 | Text to Video to generate 10–20 second vertical clips with built-in sound for platforms that favor short-form video, relying on its aspect ratio presets and audio synchronization. For example: “9:16 street fashion montage, 15 seconds, upbeat pop track, quick cuts between outfits.” Marketers can rapidly prototype campaign teasers or product hero shots, leveraging 720p output and cinematic camera motion for polished yet fast iterations. Example: “21:9 10-second car reveal, dramatic lighting, slow pan, cinematic orchestral hit.” Designers can storyboard interactions or UI concepts as animated sequences, describing transitions and ambient sound to communicate feel. Developers can integrate the Flux 3 | Text to Video API into apps that convert user text into clips, setting duration and aspect ratio parameters dynamically. Example: “square 1:1 explainer of app feature, 5 seconds, calm narration-style music.”

  • Tips & tricks

    Tips and Tricks

    To get the most from Flux 3 | Text to Video, write prompts that combine scene description, movement, and audio cues. The model responds well to explicit camera instructions (for example, “slow dolly in,” “handheld tracking shot”) and temporal structure across the 5–20 second window. Include atmosphere and sound design (“soft piano”, “city ambience”, “crowd cheering”) so the native audio track aligns with the visual narrative. When using the Flux 3 | Text to Video API, set duration and aspect ratio up front to match your target platform and avoid unnecessary re-renders. For reference-image modes, briefly describe both the source image and the desired transformation over time.

    Example prompts:

    • “Cinematic 15-second shot of a futuristic city at sunset, slow aerial orbit, neon lights turning on, gentle synthwave soundtrack.”
    • “Vertical 9:16, 10-second product demo of a smartwatch on a rotating stand, clean studio lighting, soft electronic background music.”
    • “Loopable 5-second cozy coffee shop scene, camera panning past wooden tables, steam from mugs, quiet jazz playing.”
  • Technical spec

    Technical Specifications

    • Provider: Black Forest Labs (BFL), FLUX 3 multimodal foundation model family.
    • Generation modes: Text-to-video and image-to-video in a single model, with native audio.
    • Max duration: Up to 20 seconds per generation; common presets at 5, 10, 15, and 20 seconds.
    • Resolution: 480p and 720p output in current public implementations.
    • Aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, 9:21, plus auto selection.
    • Audio: Synchronized soundtrack generated by the same flow-matching backbone, audio on/off toggle.
    • Architecture: Unified flow-matching / self-flow backbone for images, video, and audio.
    • Inputs: Text prompt; optional reference image or video depending on API mode; duration and aspect ratio parameters.
    • Outputs: Video file with embedded audio (format may vary by Flux 3 | Text to Video API implementation).
  • Things to be aware of

    Things to Be Aware Of

    Flux 3 | Text to Video currently targets short clips and moderate resolutions, so it may not meet requirements for 4K or long-form content. Complex multi-scene stories packed into a single prompt can lead to muddled motion or rushed pacing; it performs best when the time window and main action are clearly defined. Audio is generated automatically and can be disabled, but highly specific lyrics or speech are not a primary focus and may require post-production. Early-access deployments of Black Forest Labs text-to-video models have evolving performance characteristics, so API latency and throughput can vary by infrastructure. Users should also be mindful of content policies and avoid unsafe or disallowed prompts when automating generation at scale.

  • Key considerations

    Key Considerations

    Flux 3 | Text to Video is designed for short-form clips rather than full-length films, with an upper limit of around 20 seconds per generation. It excels when prompts clearly describe motion, camera behavior, and sound, leveraging its multimodal training on image, video, and audio. For teams needing tight control over shot framing, aspect ratio presets from vertical 9:16 to ultra-wide 21:9 help optimize content for social feeds, web, or product demos. Because FLUX 3 Video is still in early access, production SLAs, pricing, and very high resolutions are not yet standardized, so it is best used for rapid ideation, prototyping, and marketing assets rather than heavy offline rendering pipelines.

  • Limitations

    Limitations

    Flux 3 | Text to Video is constrained to clips of around 20 seconds per generation and to 480p–720p resolutions in current public tooling. It is not a full video editor; fine-grained frame-by-frame control, multi-shot sequencing, and complex timelines must be handled outside the model. While Black Forest Labs text-to-video outputs include audio, they are not optimized for high-precision speech synthesis or long-form music composition. Highly technical scenes, dense text overlays, or exact brand assets may require additional passes or manual refinement to achieve pixel-perfect results.

Related models

4 models