Skip to main content
GPT-6 Astra robot assembling IKEA furniture, demonstrating advanced AI capabilities in a domestic setting.

Editorial illustration for GPT-6 Astra scores above 28% on IKEA assembly test

GPT-6 Astra Tops IKEA Assembly Test at 28%

• 4 min read

Epoch AI runs a test called the Furniture Assembly Benchmark, or FAB, and it does exactly what the name suggests: it photographs three IKEA pieces mid-build, some with deliberate errors baked in, and asks AI models to compare the photos against the instructions and spot what went wrong. Last November, the best-performing model, Claude Opus 4.5, managed a score of 28%. That's not a typo. A quarter century of household frustration with Swedish flat-pack furniture and the leading AI model could barely tell you if you'd installed a shelf pin backwards.

Ten months is apparently long enough to fix that. OpenAI's GPT-6 Astra now scores 80% on the same benchmark, though it takes about three minutes to process each photo. Anthropic's Claude Fable 5.1 sits at 70%, with Claude Opus 5 at 61%. Chinese open-weight models, including Kimi K3, are running roughly seven months behind the frontier.

The jump matters less for shelf-building and more for what it signals: models that recently struggled with basic visual reasoning are now catching detailed assembly errors from a photo. Researchers see the same underlying capability eventually stretching toward car repairs or appliance troubleshooting, once the speed catches up to the accuracy.

AI can now spot when you've built your IKEA furniture wrong. Epoch AI's Furniture Assembly Benchmark (FAB) photographs three IKEA pieces during assembly with deliberate errors. Models compare photos against instructions, identify mistakes, and describe what went wrong.

Why this matters

An 80% score on FAB looks impressive until you remember the task: matching a photo to an instruction sheet and naming the error. That's a narrow, well-defined visual reasoning problem, not general-purpose robotics or real-world troubleshooting. The jump from 28% to 80% in ten months is real progress on spatial reasoning benchmarks, but three minutes per photo is not a number anyone building a product should ignore.

That's inference cost and latency that doesn't scale to, say, a customer support bot checking install photos in real time. For founders eyeing "AI that can debug physical objects" as a pitch, FAB is a useful signal that multimodal models are getting better at comparing structured references against messy real-world images, which matters for manufacturing QA, repair diagnostics, and accessibility tools. But the gap between Claude Opus 4.5's 28% and Astra's 80% also shows how volatile these leaderboards are: a benchmark that looked hard a year ago can get solved fast once labs target it directly.

Worth watching whether that speed generalizes past IKEA shelves.

Common Questions Answered

What is the Furniture Assembly Benchmark (FAB) and how does it test AI models?

The Furniture Assembly Benchmark, created by Epoch AI, photographs three IKEA furniture pieces mid-assembly with deliberate errors intentionally built in. AI models are then asked to compare the photos against the official instructions and identify what mistakes were made in the assembly process.

How much has AI performance improved on the FAB test from November to the present?

Claude Opus 4.5 achieved a score of 28% on the FAB test in November, while GPT-6 Astra has now scored above 80% on the same benchmark. This represents significant progress in spatial reasoning capabilities over approximately ten months.

What are the practical limitations of GPT-6 Astra's performance on the FAB test?

While GPT-6 Astra's 80% score is impressive, the model takes approximately three minutes per photo to analyze, which represents a substantial inference cost and latency issue. This processing time does not scale well for real-world applications requiring faster troubleshooting or product development scenarios.

Why is the jump from 28% to 80% on FAB important but not revolutionary?

The FAB test measures performance on a narrow, well-defined visual reasoning problem—matching photos to instruction sheets and naming errors—rather than testing general-purpose robotics or real-world troubleshooting capabilities. While the improvement demonstrates genuine progress in spatial reasoning benchmarks, it represents advancement in a specific, constrained task rather than broader AI competency.

LIVE12:44Nvidia’s SoL-Pi halves coding agent token use with evidence-preserving reducer