all dispatches
Best OfOct 7, 20269 min read

Best AI Models for Mask Editing in 2026

Ideogram edits black, Bria and Flux Fill edit white, GPT Image edits transparent. Here's every model on each::labs that takes a mask, how each reads it, and which one to pick for text, products, fills and video.

Best AI Models for Mask Editing in 2026

You paint a careful mask over the coffee cup. You ask for a glass of orange juice instead. The model hands back a new kitchen, a new table, a new window, and the coffee cup, untouched, right where you left it.

Nothing was broken. The model did exactly what your mask told it to do. You just spoke the wrong color.

That's the quiet trap in mask editing with AI models. A mask is the most precise instruction you can give an image model, a pixel-level "here, and nowhere else", but the models that accept one don't agree on what the colors mean. On each::labs we checked the request schema of every model in the catalog to find out which ones take a mask at all, and the list is shorter and stranger than you'd expect. Here's which AI models support masks, what each mask is for, how each one reads it, and which one we'd pick for which job.

What a Mask Does That a Prompt Can't

A prompt describes. A mask points. When you write "replace the sign on the left", the model has to work out which sign and where "left" ends. When you paint the sign, there's nothing to interpret: these pixels change, those don't.

That matters most on client-approved images, on text inside a design, and on any job you'll run more than once. A mask is a contract you can reuse. A prompt is a request the model reads a little differently every time.

Editors that skip the mask find the region themselves, and they're good at it, often. "Often" is the word a mask exists to remove.

Black means edit here. White means edit there. Read the manual first.
Black means edit here. White means edit there. Read the manual first.

Black, White or Transparent: The Mask Convention Clash

Three conventions, and they contradict each other.

Black means edit in the Ideogram family: Ideogram 4.5 Edit, Ideogram V3 Turbo and Ideogram Character all inpaint the black pixels and keep the white ones.

White means edit in Bria Fibo Edit 1.5 and Flux Fill Pro. Black is preserved, white gets rewritten. The exact opposite.

Transparent means edit in GPT Image 2.5. There the mask is a PNG with an alpha channel, and the fully transparent areas are the ones that change. Color doesn't enter into it.

So a mask that works perfectly in Ideogram edits everything except your subject in Flux Fill Pro. An inverted mask is the classic mistake when you switch models, and it fails silently: you get a clean, confident, completely wrong image. Keep one master mask where white marks the edit, then convert at the edge of your pipeline:

from PIL import Image, ImageOps

m = Image.open("mask_white_edits.png").convert("L")
m.save("mask_for_bria_and_flux_fill.png")          # white = edit
ImageOps.invert(m).save("mask_for_ideogram.png")    # black = edit
rgba = Image.new("RGBA", m.size, (0, 0, 0, 255))
rgba.putalpha(ImageOps.invert(m))                   # edit = transparent
rgba.save("mask_for_gpt_image.png")

One more rule all of them share: the mask must match the source image's size. Only Ideogram V3 Turbo and Ideogram Character resize a mismatched one for you.

The best edit is the one nobody notices outside the lines.
The best edit is the one nobody notices outside the lines.

Ideogram 4.5 Edit: Best for Text and Multi-Turn Precision

Ideogram 4.5 Edit is the model here built most squarely around the problem masks solve. Ideogram calls it "the most precise edit model" and aims it at "edit drift": the artifacts, color shifts and softened details that pile up when you edit the same image five times in a row. Its docs make the promise concrete. The output always keeps the source's width and height, and pixels the edit didn't meaningfully change are copied exactly from the original.

The mask goes in mask_url. Black marks the area to edit, white the area to keep, and gray values get rounded to whichever is nearer. Two details trip people up. The mask has to contain both black and white after rounding, so an all-black "edit everything" mask is rejected. And masked edits keep the source geometry: only unmasked edits accept a new size. You can add up to four reference images, or three when you send a mask, and references are never edited themselves, only the main image is. The full contract lives in Ideogram's API docs.

Ideogram's long typography record shows here. Launch coverage highlights editing stylized lettering in place and translating copy while the layout around it holds still. Paint a black box over the headline and write:

Replace the headline with "SUMMER DROP" in the same condensed white serif, same size and letter spacing, centered in the masked area.

GPT Image 2.5 Flare and Sunburst Edit: Masks Written in Alpha

GPT Image 2.5 Flare Edit and GPT Image 2.5 Sunburst Edit share the same mask contract. The mask is optional, a PNG with an alpha channel, the same dimensions as the first source image, and fully transparent areas mark the edit. That "first" matters: these models take up to 16 source images, and the mask only ever applies to image one.

That combination is the reason to pick them. You can hand GPT Image a room, a sofa, a lamp and a rug, cut a transparent hole where the sofa should go, and describe how the pieces come together. Quality runs from low for drafts up to max for finals.

Place the sofa from image 2 into the transparent area of image 1. Match the window light from the left, add a soft contact shadow on the floor, keep the wall color unchanged.

Note the version line. GPT Image v1.5 Edit and v2 Edit on each::labs have no mask input. If you need OpenAI with a mask, it's 2.5 or the older GPT-1 Image Edit.

Precision is mostly about what you refuse to touch.
Precision is mostly about what you refuse to touch.

Bria Fibo Edit 1.5: Masks for Repeatable, Controlled Edits

Bria Fibo Edit 1.5 flips the Ideogram convention: black preserved, white edited. The mask is optional, has to match the input size, and is only valid when you send exactly one image (the model takes up to four without one).

The interesting part isn't the mask. It's what Bria lets you put next to it. Instead of a plain instruction you can pass structured_instruction, a JSON edit instruction meant for reproducible, programmatic edits. Pair that with a fixed seed and a mask, and an edit becomes something your backend can run the same way on the ten-thousandth catalog image as on the first. The schema also ships with content moderation on prompts, input images and outputs switched on by default, plus an optional ip_signal flag that warns when an instruction may touch IP-sensitive content. For a team that answers to legal, those defaults are the feature.

Change the background inside the white area to a seamless light gray studio sweep. Keep the product, its shadow and its label exactly as they are.

Flux Fill Pro: The Classic Inpainting Specialist

Flux Fill Pro from Black Forest Labs is the dedicated fill model of the group. It takes an image, a mask and a prompt, and fills. White is inpainted, black preserved, the mask matches the image size, and you can skip the separate mask entirely if your source PNG already carries an alpha mask.

It gives you more dials than the newer editors: steps (default 50) trades time for finer detail, guidance (default 3) balances prompt adherence against image quality, and prompt_upsampling rewrites your prompt for more creative fills. For object removal, keep the prompt boring on purpose. Describe what should be behind the object, not the object.

Empty wooden boardwalk continuing into the distance, same afternoon light, same weathered planks.

Outpainting works the same way with one manual step. Extend the canvas yourself, paint the new border white in the mask, and let Flux Fill Pro build the rest of the scene.

A mask that moves with the shot is a different kind of promise.
A mask that moves with the shot is a different kind of promise.

Ideogram V3 Turbo and Ideogram Character: Fast Fills and Placed Faces

Ideogram V3 Turbo is a text-to-image model that becomes an inpainter when you pass image plus mask. Black pixels are inpainted, white preserved, and the mask is resized to fit the image. It's a quick way to patch a region with Ideogram's lettering skills when you don't need 4.5's multi-turn guarantees.

Ideogram Character uses the same black-equals-edit mask for a narrower, very useful job: putting a consistent character into an existing picture. You pass a character_reference_image, the scene as image, and a mask over the spot where the character should appear. Mask a seat at the café table, and the person from your reference sits down in it.

PixVerse Modify: Masks That Live in Video

Masks aren't only for stills. PixVerse Modify edits a region of a video, and it handles masks by name rather than by color. You send up to three masks in mask_urls, and the prompt refers to them as @selection0, @selection1 and @selection2. Reference images, up to ten, become @img0, @img1 and so on. A keyframe_id sets the frame where the edit applies.

@selection0 subject is replaced with @img0, keep the camera move and the background unchanged.

That covers subject swaps, adding or removing objects, relighting, on-screen text and restyling, at 360p, 540p or 720p. Because the schema names masks instead of defining a color rule, run one short test clip before you batch.

Worth a brief mention: Infinitalk, the talking-avatar model, also takes a mask_image. In the image-to-video version it picks which person in a group photo the audio animates. In video-to-video it marks which regions are allowed to move.

Here's the surprise for a lot of people: several of the most-used image editors on each::labs don't take a painted mask at all. Nano Banana, Nano Banana 2 and Nano Banana Pro Edit, Seedream v4, v4.5 and v5 Edit, the Qwen Image Edit family, Wan 2.7 Image Edit and the Flux Kontext models all edit by instruction and reference images. You describe the region in words and the model locates it.

Flux 3 Edit takes a third route: bounding boxes. Instead of a mask image, you tag elements in the prompt and give each a box as [top, left, bottom, right] on a 0 to 1000 grid, so the layout survives any resolution. Black Forest Labs notes that pixels outside the boxes typically stay unchanged, though shadows, reflections or nearby lighting may still shift. Boxes are coarser than a mask and quicker to write. Browse the full image-to-image models to see both camps side by side, and our guide to using AI image editing models covers the instruction-only workflow.

When to Pick Which Mask Editing Model

These verdicts are our opinion, built on what each model's documented inputs let you do, not on a benchmark.

Typography and text edits: Ideogram 4.5 Edit. Ideogram's text heritage plus exact pixel copying outside the mask is the right combination for changing a headline, a sticker on a packaging mockup or a translated label without disturbing the design. For quick, single-pass text patches, Ideogram V3 Turbo is the lighter option.

Product shots and precise multi-turn edits: Ideogram 4.5 Edit, with GPT Image 2.5 as the pick when the edit is a composite. If you'll edit the same hero image again and again, 4.5's drift-free design is the point. If the job is "put these three products into that scene", GPT Image 2.5's 16 sources and alpha mask win.

Classic fill, removal and outpaint-style extension: Flux Fill Pro. It's purpose-built, accepts an alpha mask straight from your PNG, and exposes the steps and guidance controls you want when a fill has to blend.

Enterprise-safe, structured edits at scale: Bria Fibo Edit 1.5. Structured JSON instructions, seeds and moderation on by default make it the one we'd wire into a catalog pipeline that has to behave the same every run.

Video region edits: PixVerse Modify. It's the one true video editor on the list, and the @selection syntax makes multi-region edits readable.

Placing a consistent character into a scene: Ideogram Character.

Best overall for still images? Ideogram 4.5 Edit. It covers the widest range of jobs with the strictest promise about what it won't touch. Just remember it speaks black.

Frequently Asked Questions

Which AI models support mask editing on each::labs?

After checking every model's schema, here's the list. For images: Ideogram 4.5 Edit, Ideogram V3 Turbo, Ideogram Character, GPT Image 2.5 Flare Edit and Sunburst Edit, GPT-1 Image Edit, Bria Fibo Edit 1.5, Flux Fill Pro and the older Realistic Vision V3 Inpainting. For video: PixVerse Modify, plus Infinitalk using a mask as a selector.

Why did my mask edit the wrong area?

Almost always an inverted convention. Ideogram edits black, Bria and Flux Fill Pro edit white, and GPT Image 2.5 edits fully transparent pixels. A file that's correct for one is backwards for the next, so convert it per model.

Do Nano Banana, Seedream and Qwen Image Edit accept a mask?

No painted mask on any of them. They find the region from your instruction and references. Flux 3 Edit doesn't take one either; it uses bounding boxes on a 0 to 1000 grid instead.

Is a mask better than a bounding box?

For irregular shapes and small details, yes. For moving or recoloring whole objects, a box is faster and usually precise enough. Pick by the edge you need to protect.