Coordinates Instead of Free Text: Flux 3 Image Edits Photos via JSON Grid from 0 to 1000

Black Forest Labs launched Flux 3 Image on October 1, 2026, completing the image branch of its multimodal Flux 3 family. The company promises multi-step editing in which the rest of the image stays untouched — but its own published…

A photograph divided into a precise grid of thin cobalt lines, with a single cell lifted slightly out of place revealing a different image beneath. Editorial still life in cream, black and cobalt.
Gift article

Coordinates Instead of Free Text: Flux 3 Image Edits Photos via JSON Grid from 0 to 1000

Black Forest Labs launched Flux 3 Image on October 1, 2026, completing the image branch of its multimodal Flux 3 family. The company promises multi-step editing in which the rest of the image stays untouched — but its own published measurements show that between 67.8 and 89.7 percent of pixels remain unchanged in the released examples. That is substantial, but far from complete preservation, and the tension between the marketing and the numbers may be the most interesting part of the launch.

What the model covers

Flux 3 Image is the image side of the Flux 3 family and, according to Black Forest Labs, covers text-to-image, image-to-image, text rendering, and photorealism. Multi-step editing — that is, the ability to make successive changes without other parts of the image changing — is BFL's own claim, not something independently verified (The Decoder).

In practice, this means users can compose scenes with bounding boxes, include up to ten reference images, and produce output up to 4K. A free demo is available.

How the coordinate system works

The technical core is a structured JSON scene layout. Every element in the scene has an ID, a description, and coordinates. The coordinates are mapped to a normalized grid from 0 to 1000, measured from the image's top-left corner, in the order [top, left, bottom, right] — according to an MSN-syndicated article describing the documentation in BFL's materials (MSN).

Instead of describing a desired change in free text alone, users can specify exactly where in the image the element should sit, with machine-readable identifiers. This makes composition deterministic at a level that pure text prompts rarely achieve.

What BFL's own numbers actually show

This is where the story becomes nuanced. BFL has itself published pixel-difference measurements of editing operations, as relayed by the MSN article:

  • Removal of a cat: 80.7 percent of total image pixels identical
  • Addition of a bird: 86.4 percent unchanged
  • A print on a shirt: 89.7 percent intact
  • Hair recoloring: 67.8 percent unchanged — the lowest figure in the published examples

The individual frames in the article present this more strongly, with wording that pixels outside the selected region are preserved "bit-identically." But the numbers BFL itself has published show something different: large, but partial, preservation. The MSN article itself points out that hair recoloring necessarily "bleeds" color into immediately adjacent regions, and calls the 67.8 percent figure a meaningful revelation about the limits of the guarantee.

Why the figures are not 100 percent — whether they reflect total frame area, model drift, or something else — is not explained in the available material. The key point for the reader is this: the measurements are company-published, relayed through secondary reporting, and not independently verified. At the same time, they show an unusual degree of transparency: BFL could have chosen not to publish them.

The agentic angle: JSON as project state

The MSN article suggests a further reading: because element IDs and coordinates are stable JSON state, an LLM agent could treat the element table as a persistent project file across conversation turns — and edit individual elements without regenerating untouched regions.

This is the article's analysis of the architecture, not a documented capability or a benchmark. But it points to why JSON-based scene layout could matter more than it appears: an image described in machine-readable structure can in principle be edited programmatically, not just by humans writing instructions.

Pricing, licensing, and the open-weights race

The time window is tight. API access is 50 percent discounted through October 8. Companies can license commercial weights to run and fine-tune the model on their own infrastructure. An open-weight version is expected "in the coming weeks," according to The Decoder — no concrete date has been given.

The context is an active race around editing-focused image models. Shortly before the launch, Ideogram announced its own editing model, version 4.5, which is also expected as open weights soon. Both companies opening their weights means that comparisons on one's own machines — including independent tests of the preservation claims — should soon be possible.

What remains to be answered

Three questions remain open. First, the technical: what does 67.8–89.7 percent unchanged pixels actually mean in practice, and why does it not match the "bit-identical" wording in the same material? Independent measurements may clarify this once the weights arrive. Second, the commercial: whether the open-weight version will match the API model, and what restrictions commercial licensing involves, is unclear in the available material. Third, the competitive-strategic: Ideogram 4.5 has been announced but not launched — a concrete comparison between the two editing models does not yet exist.

For now, the essentials stand: Flux 3 Image puts a structured, coordinate-based editing system into users' hands, with a discount period that expires October 8, and BFL has published its own numbers showing both the strength and the limits of the preservation claim.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.