Flux 3 Video Continuation API

Video·flux-3·by Black Forest Labs

FLUX.3 is Black Forest Labs’ frontier video extension model, continuing existing clips with smooth, coherent motion while preserving the original scene, style, and visual consistency.

Runtime (p50)
4m
Estimated price
From $0.12
Call the API
prediction.sh
sh
curl -X POST \
  -H "Authorization: Bearer $EACHLABS_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "flux-3-video-continuation",
    "version": "0.0.1",
    "input": {
        "prompt": "The dancer must continue performing the pantomime dance.",
        "duration": "10",
        "video_url": "https://cdn-us.eachlabs.ai/defaults/a6cbf41d1ca948f3ac8c18f66f5d5e8c.mp4",
        "resolution": "hd",
        "aspect_ratio": "auto",
        "generate_audio": true
    },
    "webhook_url": ""
}' \
  https://api.eachlabs.ai/v1/prediction/
Documentation8 sections
  • Overview

    Flux 3 | Video Continuation Overview

    Flux 3 | Video Continuation is a frontier video-to-video extension model from Black Forest Labs, built on the multimodal FLUX 3 Video foundation for synchronized image, video, and audio generation. It continues existing clips by predicting future frames and soundtrack in a single pass, creating smooth motion while preserving the original scene composition, style, and temporal coherence. Unlike traditional video upscalers or stylizers, Flux 3 | Video Continuation leverages the unified Self-Flow backbone of FLUX 3, which is trained jointly across image, video, and audio tokens, enabling highly consistent visual and auditory continuity over extended sequences. Integrated on each::labs, it focuses specifically on extending user-provided videos, rather than generating standalone clips from scratch.

  • Capabilities

    Capabilities

    • Extends existing video clips with smooth, temporally consistent motion, using FLUX 3 Video’s joint video prediction pipeline.
    • Preserves the original scene layout, style, and visual identity, thanks to FLUX 3’s unified flow-matching backbone trained across video tokens.
    • Generates synchronized audio continuation alongside video, including ambient sounds and simple dialogue-style patterns, in a single pass.
    • Supports multiple aspect ratios (e.g., 16:9, 9:16, 1:1) for social, cinematic, and vertical formats, with auto matching of source ratio.
    • Handles text-conditioned video-to-video, allowing users to steer future motion, atmosphere, or camera movement via prompts.
    • Can perform keyframe-to-video style continuation when the input clip acts as a visual keyframe sequence, enabling controlled transitions.
    • Maintains scene-level coherence across up to ~20 seconds of extended footage per generation, making it suitable for short narratives and ads.
    • Designed to integrate with the Flux 3 | Video Continuation API on each::labs, enabling programmatic workflows for batch video extension.
  • Use cases

    Use Cases for Flux 3 | Video Continuation

    For content creators, Flux 3 | Video Continuation can extend short cinematic shots into longer sequences while preserving the original color grading and camera language, ideal for story-driven reels or film previsualization. A creator might prompt: “extend the tracking shot as the protagonist walks deeper into the alley, fog getting denser and lights flickering.” For marketers, the model can lengthen product clips into full 15–20 second ads, keeping brand visuals consistent while adding motion beats and ambient audio. Example: “continue the product spin shot, camera circling wider as soft spotlight moves and subtle music builds.” For designers and art directors, Flux 3 | Video Continuation enables mood tests by stretching motion graphics or concept animatics into longer loops without re-rendering from scratch. Prompt: “keep the abstract geometric animation evolving, shapes morphing slowly and ambient synth sound growing.” For developers, the Flux 3 | Video Continuation API supports automated video extension pipelines, such as generating alternate cuts or platform-specific durations directly from a base asset.

  • Tips & tricks

    Tips and Tricks

    To get the best results from Flux 3 | Video Continuation, anchor the continuation with a concise, descriptive prompt that reinforces the scene identity, style, and desired motion of your input clip. For example, specifying camera behavior (“slow dolly forward”) or action progression (“the character turns and walks toward the window”) helps the Self-Flow backbone maintain smooth temporal dynamics. When using the Flux 3 | Video Continuation API on each::labs, match the output aspect ratio to the source video or leave it on auto to avoid unwanted cropping or letterboxing. Audio generation works best when the source track is clean and consistent, enabling the model to predict plausible future ambience or dialogue; disable audio if you plan to layer custom sound design later. Example prompts include: “continue the scene as the drone flies further over the forest, golden hour light getting warmer,” “extend the office walkthrough, camera moving down the hallway while employees keep working,” and “keep the animated character running through the city as neon lights grow brighter and rain intensifies.”

  • Technical spec

    Technical Specifications

    • Base model: FLUX 3 Video multimodal flow model from Black Forest Labs, with unified Self-Flow architecture for image, video, audio, and actions.
    • Generation type: Video-to-video continuation (optionally with text conditioning) from an input clip; audio continuation generated in the same pass when enabled.
    • Max duration per continuation: Up to approximately 20 seconds of additional video, aligned with FLUX 3 Video’s single-generation limit.
    • Output resolution: Typically 480p or 720p for video generations, depending on configuration; higher resolutions are not guaranteed.
    • Aspect ratios: Supports landscape, square, and portrait ratios such as 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and 9:21, with an auto option to match the source.
    • Input formats: Short video clips (plus optional text prompt) as primary input; audio track used for conditioning when present.
    • Output formats: Encoded video file with optional synchronized audio; audio can be disabled for silent continuations.
    • Inference behavior: Single-pass generation for the entire extended segment; exact processing time depends on clip length, resolution, and provider infrastructure.
  • Things to be aware of

    Things to Be Aware Of

    Because FLUX 3 is an early-access multimodal foundation model, behavior can vary more than mature production systems, especially on edge-case footage. Very fast cuts, heavy motion blur, or complex overlays may reduce the model’s ability to maintain consistent semantics from frame to frame. Audio continuation is learned jointly but uses a relatively small fraction of training tokens, so intricate musical structure or highly specific speech content may not match professional sound design expectations. Users should avoid over-constraining prompts; overly detailed instructions can conflict with the visual dynamics of the input video and lead to artifacts or abrupt scene changes. Longer continuations also increase compute requirements, so it is often more efficient to generate several shorter segments and assemble them in post.

  • Key considerations

    Key Considerations

    Flux 3 | Video Continuation is best used when you need coherent narrative extension of an existing clip, not frame-by-frame editing. Because FLUX 3 Video predicts motion and audio jointly, the model works particularly well when the source video has clear subject focus and stable composition; noisy or rapidly changing footage can reduce temporal consistency. Users should be aware that FLUX 3 remains an early-access, frontier model, so service-level guarantees, performance characteristics, and pricing for Flux 3 | Video Continuation API may evolve over time. For tasks that require precise per-frame control or ultra-high-resolution outputs, image-oriented FLUX 2 or FLUX 1 variants may be more suitable than this continuation-focused setup.

  • Limitations

    Limitations

    Flux 3 | Video Continuation inherits FLUX 3 Video’s maximum generation length of about 20 seconds per clip, which constrains very long-form storytelling to multi-pass workflows. Current resolutions focus on 480p and 720p; native 4K or higher outputs are not part of the public FLUX 3 Video specs and should not be assumed. The model does not guarantee frame-perfect continuity for highly technical sequences, such as complex typography animations or precise brand-safe character acting. Text-to-audio remains generative and approximate, so users needing exact scripted voiceover or licensed music should layer external audio tools on top of the visual continuation.

Related models

4 models