Kling 4.0 Is Coming: 30-Second Shots, 4K and Keyframes
Kling 4.0 doubles native video length to 30 seconds and adds keyframes, 4K and a much bigger reference system. Here's what Kling has confirmed and how to prepare on each::labs.

Fifteen seconds was always the quiet ceiling. You could plan a beautiful shot, nail the character, get the camera move right, and then watch the clip end just as the idea started to breathe. Every longer story became a stitching job: generate, trim, match the next shot, hope the face still looked like the same person.
Kling 4.0 is Kling's answer to that ceiling. Kling published the announcement on September 30, 2026, opened early access to a limited group of users on September 28, and says the full launch is scheduled for October 2026. It isn't on each::labs yet. What follows is what Kling has actually confirmed, what it means for the way you build video, and how to get your pipeline ready so you can run it on each::labs the day it lands.
What Kling Has Confirmed About Kling 4.0
Start with the official list, because it's specific. According to Kling's launch post, Kling 4.0 generates video natively from 3 to 30 seconds. It supports 4K output and a 21:9 ultrawide format, and Kling says 10-bit HDR at 1080p and 4K is "coming soon". Audio gets an upgrade to higher quality two-channel stereo, with improved lip-sync accuracy.
Control is the other headline. Kling lists up to 10 flexible keyframes on the timeline, prompts of up to 8,000 tokens, and a much larger reference system: up to 15 combined reference items, of which up to 10 can be images, up to 5 can be videos (30 seconds combined), and up to 7 can be subjects, including up to 3 built from video. Editing tasks accept up to 5 video inputs.
What Kling hasn't said matters just as much. The post doesn't mention API availability, and it doesn't publish benchmarks. Kling's page does list a lighter version, Kling 4.0 Flash, next to the full model, and that's the version early testers are running right now.

Thirty Seconds Changes the Unit of Work
Doubling the length sounds like a spec bump. It's actually a change in what a single generation is for.
The Kling 3.0 family on each::labs generates 3 to 15 seconds per clip. That's a shot. Thirty seconds is a scene: a product reveal, a short dialogue exchange, a full social ad with setup and payoff. When the model holds a scene in one pass, it also holds the things you used to lose between cuts: the light, the wardrobe, the way a character moves. Stitching is where consistency goes to die, and less stitching means fewer chances for it.
It also changes how you prompt. A 5-second clip wants one clear action. A 30-second clip wants a structure, a beginning, a turn and an end, and Kling's 8,000-token prompt limit is a clear signal that it expects you to write one.

Keyframes Turn Prompting Into Directing
Here's the feature that will separate casual users from teams that ship. Kling says creators can use up to 10 keyframes to "define important visual states at different points on the timeline", controlling when characters appear or when a scene transitions.
That's the difference between describing a video and blocking one. Instead of writing "the camera reveals the bottle and then the logo appears" and hoping the timing lands, you pin the visual state you need at second 4, the one you need at second 12, and the one you need for the final frame. The model fills the motion in between. Anyone who has tried to force a product to be on screen at exactly the right moment will understand why this matters.
Today on each::labs, the closest thing is the start and end frame control in Kling 3.0 Pro image-to-video, plus its multi-shot mode with up to five prompts. If you build your storyboards around those now, moving to more keyframes later is an upgrade, not a rewrite.
Omni Reference Gets a Lot Bigger
Most consistency problems are reference problems. You described the character in words, the model guessed, and the guess drifted. Kling 4.0's reference limits are the bet that you'll stop describing and start showing.
Fifteen references in one job is a different kind of brief. Picture a brand film: a few images of the product from different angles, a couple of images of the presenter, a video reference for the camera rhythm you want, another for the pacing of the edit. Kling's FAQ says reference videos can guide "pacing, shot composition, and camera movement", and that multimodal editing lets you modify "subjects, backgrounds, visual styles". That's a director's toolkit, not a text box.
If you've been fighting identity drift, our guide to keeping the same character across AI video shots covers the habits that will carry straight over: clean frontal references, consistent lighting in your reference set and one subject per reference.

Sound, Text and the Details That Sell a Shot
The upgrades that get the least attention are often the ones that decide whether a clip is usable. Kling says Kling 4.0 improves motion stability in fast scenes, preserving "subject structure, spatial relationships, and motion continuity", and renders more natural facial detail and expressions. It also claims better rendering of multilingual text, logos and emojis during camera movement.
That last point is worth testing hard when it arrives. On-screen text that survives a camera move is exactly what product and brand ads need, and it has been a weak spot across AI video. Stereo audio and tighter lip-sync push in the same direction: fewer clips you have to fix in an editor before they can go anywhere.
Kling 4.0 vs Kling 3.0: What Actually Changes
Put the two side by side and the shape of the upgrade is clear. Kling 3.0 on each::labs tops out at 15 seconds, offers 16:9, 9:16 and 1:1 framing on text-to-video, has a 4K variant, generates native audio, supports multi-shot sequences with up to five prompts, and takes up to three character or object elements on image-to-video. Kling 4.0, per Kling, doubles the length to 30 seconds, adds 21:9, moves to stereo audio, adds up to 10 keyframes and takes up to 15 references in one job.
So is Kling 3.0 suddenly obsolete? No. For a single punchy shot, a fast iteration loop or a high-volume pipeline, a proven model with known behavior is often the better call, and Kling 3.0 Turbo exists exactly for speed. The new model earns its place on longer, reference-heavy, story-driven work. Expect most teams to run both.

How to Get Ready on each::labs Today
You don't have to wait for launch day to prepare. The work that makes Kling 4.0 useful is mostly work you can do now.
First, build your reference library. Collect clean product shots, character sheets and a few motion references you like. When the model arrives, the teams with organized references will be generating useful clips on day one while everyone else is still hunting for assets.
Second, write longer scripts on today's models. Use Kling 3.0 Pro text-to-video with multi-shot prompts to test how your 30-second ideas break into beats. A structure that works as five shots will translate naturally into keyframes later.
Third, keep your integration model-agnostic. On each::labs every text-to-video model sits behind the same API, so switching to a new Kling version is a model id change, not a new integration. If you chain steps (image, then video, then upscale), build that chain as an each::labs flow now and swap the video step when the time comes.
What Kling 4.0 Early Access Is Showing
Hype is loud. Footage is quieter, and more useful. Early access to Kling 4.0 started rolling out on September 28, and what people are running today is Flash, not the full model. Hands-on write-ups describe the early-access build as capped at 720p and limited to Kling Ultra subscribers on the annual plan. So read every clip you see on X with that in mind: it's the light version, at a fraction of the resolution the full model promises.
Even so, the early tests point somewhere specific. Lip sync is the clearest win: testers ran full dialogue exchanges without the sync drifting apart, which was a known pain with Kling 3.0. Spatial memory looks better too, with objects placed early in a clip staying where they were seconds later. The misses are just as telling. The model sometimes adds small actions nobody asked for, and performances can feel emotionally flat even when the lighting and motion are right. Plan your prompts around that: state what should not happen in the shot, and direct the emotion explicitly instead of hoping the model reads it from context.
The full model, with 30-second clips, 4K and the bigger reference system, is the one that lands in October. That's the version coming to each::labs.
Frequently Asked Questions
When is Kling 4.0 coming out?
Kling says the full launch is scheduled for October 2026, with early access rolling out to a limited group since September 28. No exact day has been published. It's coming to each::labs, and you'll be able to run it through the same API you use for Kling 3.0.
How long can Kling 4.0 videos be?
Up to 30 seconds natively, starting from 3 seconds, according to Kling. That's double the 15-second maximum of Kling 3.0 on each::labs, and it pairs with up to 10 keyframes to structure the timeline.
Is Kling 4.0 better than Kling 3.0?
On paper it does more: longer clips, 21:9, stereo audio, keyframes and far more references. Whether it's better for your work depends on the job. For quick single shots, Kling 3.0 remains a strong, proven option. For longer, reference-heavy scenes, Kling 4.0 is the one to test first.
Can I use Kling 4.0 on each::labs now?
Not yet. Kling 4.0 isn't live on each::labs today. You can prepare with the Kling 3.0 models already available, and you'll be able to use Kling 4.0 on each::labs as soon as it's publicly released.