Flux 3 API

Video·flux-3·by Black Forest Labs

FLUX.3 is Black Forest Labs’ frontier image-to-video model, transforming a single still image into high-quality video with natural motion, smooth animation, and coherent scene dynamics.

Runtime (p50)
4m
Estimated price
From $0.06
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "flux-3-image-to-video",
    "version": "0.0.1",
    "input": {
        "prompt": "Preserve the exact black horse, landscape, mountains and composition from the reference image. Maintain full identity consistency of the horse throughout the entire sequence. Begin with a slow cinematic push-in as the horse stands silently on the rocky cliff above a sea of clouds during sunrise. Gentle mountain wind flows through its mane and tail while the clouds slowly drift beneath the cliffs. The horse calmly raises its head and looks toward the distant mountain range, breathing naturally. The camera smoothly circles around the horse, revealing the vast alpine landscape glowing with warm golden light. Birds glide far in the distance while soft rays of sunlight break through the clouds, creating an epic, emotional atmosphere. Finish with the horse standing proudly against the sunrise as the camera slowly pulls back into a breathtaking wide shot. The final 2 seconds fade to black, and centered on screen, elegant white serif typography appears reading only: \"THE END\". Keep the text perfectly centered, crisp, correctly spelled, and on a completely clean black background. Ultra realistic cinematography, premium nature documentary aesthetic, cinematic lighting, subtle film grain, smooth camera movement, emotional orchestral atmosphere",
        "duration": "10",
        "image_urls": [
            "https://cdn-us.eachlabs.ai/defaults/3084ea403375453ab0f78611b19d72f1.webp"
        ],
        "resolution": "hd",
        "aspect_ratio": "16:9",
        "generate_audio": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Flux 3 | Image to Video Overview

    Flux 3 | Image to Video is Black Forest Labs’ frontier image-to-video model that turns a single still image into coherent, high-quality video with natural motion, smooth animation, and consistent scene dynamics. Built on the FLUX 3 multimodal foundation, it extends the established FLUX family from still image generation into native video with audio, using a unified flow-matching backbone across image, video, and sound. The primary differentiator of Flux 3 | Image to Video is its ability to generate up to ~20-second clips with synchronized audio in a single pass, starting from text, an input image, or existing footage, while maintaining temporal consistency and realistic camera motion. On each::labs, this model is exposed specifically for image-to-video workflows, giving creators, marketers, and developers a focused way to animate static visuals using the Flux 3 | Image to Video API.

  • Capabilities

    Capabilities

    • Image-to-video generation: Converts a single still image into a coherent, animated video clip with natural motion and scene dynamics.
    • Text-conditioned motion: Uses text prompts to control how subjects move, how the camera behaves, and how the story unfolds over time.
    • Native audio synthesis: Generates synchronized audio (music, ambience, effects, dialogue in multiple languages) in the same pass as video.
    • Video continuation and keyframe transitions: Extends existing footage, transitions between keyframes, and chains clips into longer sequences.
    • Flexible duration presets: Supports 5, 10, 15, and 20-second clips, plus auto-selection to match content dynamics.
    • Multiple aspect ratios: Handles cinematic widescreen, standard landscape, square, and vertical formats suitable for social feeds and mobile.
    • Multilingual content support: Can produce dialogue and on-screen typography across languages when audio and text are combined.
    • Unified multimodal backbone: Shares one Self-Flow architecture across image, video, audio, and action, improving consistency and cross-modal alignment.
  • Use cases

    Use Cases for Flux 3 | Image to Video

    Flux 3 | Image to Video gives creators, marketers, designers, and developers a practical way to animate static visuals into narrative clips. Content creators can turn concept art or thumbnails into teaser videos by leveraging text-conditioned motion and flexible duration presets; for example, “Animate this podcast cover into a 10-second intro with slow zoom and subtle particle effects.” Marketers can start from product stills and generate short promotional videos in multiple aspect ratios, using native audio for quick soundtracks: “From this shoe photo, create a 15-second vertical ad with rotating product and upbeat music.” Designers can prototype motion graphics from UI mockups by specifying camera moves and typography animations. Developers integrating the Flux 3 | Image to Video API into apps can offer users a one-click “animate this image” feature, chaining clips for longer explainers or tutorials while keeping temporal coherence and audio aligned with visual events.

  • Tips & tricks

    Tips and Tricks

    Flux 3 | Image to Video responds best to prompts that precisely describe motion, camera behavior, and timing. Include verbs and temporal cues such as “slowly pans,” “camera dolly forward,” or “character turns to face the viewer over 5 seconds” so the model can plan the motion across the full clip. When using the Flux 3 | Image to Video API, start with moderate durations (5–10 seconds) to validate motion and style, then extend to 15–20 seconds once you are satisfied. For image-to-video, choose source images with clear lighting and uncluttered subjects; complex overlapping motion or extreme perspective changes from a single frame may be harder to synthesize reliably. Consider disabling audio if you plan to add custom sound design later, or keep it on to quickly prototype synced music and ambient effects.

    Example prompts:

    • “From this still of a cyberpunk street, create a 10-second night-time tracking shot with neon reflections and slow camera pan to the right, cinematic 720p video.”
    • “Animate this product photo into a 15-second hero shot: subtle rotation, soft studio lighting changes, and macro close-up at the end, no audio.”
    • “Turn this illustration of a dragon flying over mountains into a 20-second epic flight sequence, dynamic camera moves, dramatic orchestral audio.”
  • Technical spec

    Technical Specifications

    • Provider / Family: Black Forest Labs FLUX 3 multimodal foundation model.
    • Video length: Clips up to approximately 20 seconds per generation; common presets at 5, 10, 15, and 20 seconds.
    • Resolution: Typical outputs at 480p or 720p for video clips, with 720p used in reference examples.
    • Aspect ratios: From 21:9 and 16:9 through 4:3, 1:1, 3:4, 9:16, and 9:21, plus auto selection.
    • Inputs: Text prompt plus one still image or existing footage, depending on the workflow.
    • Outputs: Video file with native synchronized audio generated in the same pass; audio can be disabled in some implementations.
    • Architecture: Self-Flow / unified flow-matching backbone trained jointly on image, video, audio, and action.
    • Processing time: End-to-end generation is typically several seconds per clip on production GPU backends; exact latency depends on provider infrastructure (each::labs, cloud tier, and clip length).
  • Things to be aware of

    Things to Be Aware Of

    While Flux 3 | Image to Video is powerful, it has practical constraints. Very complex scene changes or radical camera moves inferred from a single still can produce less stable results, especially when motion contradicts the original perspective. Highly detailed images with overlapping objects may lead to minor artifacts or temporal flicker as the model interpolates occlusions over time. Audio is generated from the same semantic signal as video, so unexpected sound effects or dialogue may appear if the prompt is ambiguous; always review outputs before publishing. Users sometimes under-specify motion, resulting in clips that feel static—being explicit about actions and camera behavior is crucial. Resource-wise, longer and higher-resolution clips consume more compute and may increase latency or cost on each::labs.

  • Key considerations

    Key Considerations

    Flux 3 | Image to Video is best used when you need high-quality motion and temporal coherence from a single image, rather than frame-by-frame manual animation. The model expects a clear text prompt that explains motion, camera behavior, and desired style, plus a well-composed input image with sufficient detail in the subject and background. Because Flux 3 is natively multimodal, enabling audio means the soundtrack will follow the same semantic scene, which is valuable for dynamic marketing clips or social content but may require more review for brand safety. For longer productions, users often chain multiple Flux 3 clips into sequences instead of generating a single very long output. Cost and performance are tied to clip length, resolution, and concurrent usage of the Flux 3 | Image to Video API on each::labs.

  • Limitations

    Limitations

    Flux 3 | Image to Video cannot yet generate arbitrarily long videos in a single call; current behavior centers around clips up to about 20 seconds. Ultra-high resolutions beyond 720p, such as native 4K image-to-video, are not part of the public reference specs and may require separate upscaling workflows. The model relies on the semantic content of the input image and text; it does not guarantee pixel-perfect physical accuracy for complex simulations or scientific visualization. Fine control over every frame, such as keyframe-by-keyframe editing, is limited compared to traditional animation tools. As with any generative system, outputs may inherit dataset biases, and strict brand, legal, or safety review is recommended before deployment.

Related models

4 models