Anyone who has finished a flat-pack bookcase only to find a spare screw rolling across the floor knows the quiet dread of an IKEA build gone wrong. Now there may be a machine that can settle the argument. According to Epoch AI's Furniture Assembly Benchmark (FAB), OpenAI's GPT-6 Astra can examine a photograph of furniture mid-build and identify exactly where the assembly went off the rails, scoring 80 percent accuracy on the task.

The result arrives roughly ten months after the benchmark's best performer, Claude Opus 4.5, managed just 28 percent. That jump, from near-total failure to catching the large majority of introduced errors, is the headline of the latest FAB leaderboard update.

How the Furniture Assembly Benchmark works

FAB is deliberately mundane, and that appears to be the point. The benchmark photographs three IKEA furniture pieces during assembly, with deliberate errors introduced along the way. A model is then shown those photos alongside the manufacturer's instructions, and must do two things at once: pinpoint the mistake and describe what went wrong in words.

That combination is what makes the task harder than a simple image-classification problem. The model cannot merely flag that something looks off. It has to localise an error inside a partially assembled object and then explain it clearly enough that a human builder could act on the description. The benchmark therefore folds visual reasoning, instruction-following and language generation into a single score.

Flat-pack furniture turns out to be a useful stress test precisely because it is boring. The pieces are rigid, the instructions are diagrammatic, and correctness is unambiguous. Either the shelf rail is mounted on the wrong face or it is not. There is very little room for the kind of subjective judgment that can muddy image-based evaluations.

From 28 percent to 80 percent in ten months

The scale of the improvement is the most striking element of the update. In November 2025, the strongest available model scored well under a third of the benchmark. Ten months later, the leader is closing in on the point where the tool would be right far more often than it is wrong.

  • November 2025: Claude Opus 4.5 leads the field at 28 percent.
  • September 2026: OpenAI's GPT-6 Astra reaches 80 percent, averaging about three minutes per photo.
  • Claude Fable 5.1 sits at 70 percent.
  • Claude Opus 5 lands at 61 percent.
  • Chinese open-weight models such as Kimi K3 trail the leaders by at least seven months.

Three of the five data points describe a field clustered in a meaningful band rather than a single runaway winner. Astra holds an advantage, but Fable 5.1 is close enough that the ordering could shift with the next round of releases. The more revealing gap is the one between the frontier and the open-weight challengers, where Kimi K3 and its peers remain roughly seven months behind.

Why the curve bent so sharply

The source material frames the gain as part of a broader shift rather than an isolated feat. Models were failing far simpler visual tasks not long ago, which makes an 80 percent result on a multi-step inspection problem notable. As vision-language systems have become better at grounding descriptions in what is actually visible in an image, benchmarks built around physical objects have become tractable in a way they previously were not.

Speed is the remaining bottleneck

There is an important caveat attached to Astra's score: three minutes per photo. That is far too slow for real-time assembly assistance, the scenario most people would imagine first. A helper that takes three minutes to render a verdict is not useful when your hands are full and you are halfway through step four.

OpenAI's GPT-6 Astra spots assembly errors in IKEA furniture photos with 80 percent accuracy, up from 28 percent just months ago. | Image: OpenAI
OpenAI's GPT-6 Astra spots assembly errors in IKEA furniture photos with 80 percent accuracy, up from 28 percent just months ago. | Image: OpenAI

Researchers nonetheless see the capability as a stepping stone. They suggest the underlying skill could eventually be pointed at tasks such as car repairs or appliance fixes, where a model would compare a photo of a malfunctioning or partially disassembled system against documentation and describe the fault. Those domains are messier than a bookshelf: components are hidden, lighting varies, and prior repairs can leave a machine in a non-standard state.

Speed is also, in principle, an easier problem to attack than accuracy. Inference costs have tended to fall and throughput has tended to rise across successive model generations, so a three-minute turnaround today says more about the current state of the art than about any hard ceiling.

Beyond flat-pack: visual robotics

The benchmark result does not stand alone. Astra also performs strongly on visual robotic tasks, according to the source material, which suggests the same underlying ability, reading a scene and reasoning about physical arrangement, transfers beyond static photographs of furniture.

That connection matters because robotics has always been limited less by actuators than by perception and common sense. A robot that can look at a half-finished assembly, compare it against a plan and say what is wrong is doing something structurally similar to manipulation: it is building an internal model of an object's correct state and comparing it against reality. Error detection of this kind is a logical precursor to error correction.

A benchmark built to stay honest

FAB's design also says something about how AI evaluation is maturing. Rather than testing abstract puzzle-solving, it borrows a task from ordinary life and introduces controlled faults, which makes scoring objective and reproducible. Deliberate errors can be catalogued, so a correct answer is verifiable and a wrong answer is unmistakable.

The price of that rigor is narrowness. Three IKEA pieces is a small sample, and success on them does not guarantee success on a kitchen cabinet or a bicycle. But as a directional signal of visual reasoning progress, the benchmark is unusually legible: it measures whether a model can look at a physical object and tell you what is wrong with it.

What to watch next

Three things will determine whether this result looks like a milestone or a footnote. The first is whether Astra's lead over Fable 5.1 and Opus 5 widens or narrows in the coming months. The second is whether inference times fall far enough to make real-time guidance plausible. The third is whether the error-detection skill generalises from flat-pack furniture to the messier repair tasks researchers have in mind.

For now, the practical takeaway is modest but real: the technology that could not reliably spot a reversed shelf bracket less than a year ago can now identify the mistake eight times out of ten and explain it. Anyone who has ever argued with a partner over which way a dowel goes might consider that progress worth having.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com