Skip to main content
AI model struggling with a visual puzzle, highlighting the gap in human-like problem-solving abilities.

Editorial illustration for AI Models Fail Simple Visual Puzzles Humans Find Easy

AI Models Struggle With Visual Puzzles Humans Solve

AI Models Fail Simple Visual Puzzles Humans Find Easy

4 min read

Arthur Samuel coined the term "machine learning" in a 1959 article about a checkers-playing program. Chess followed, then Go. For nearly seven decades, games and puzzles have been the yardstick researchers use to measure how far artificial intelligence has come, precisely because humans already know how to judge good play from bad.

That yardstick keeps moving fast. Columbia University researchers found in late 2024 that even top models solved only 18% of New York Times Connections puzzles. A few months later, some models were cracking them almost every time. Progress like that makes puzzles a strange kind of report card: useful for tracking capability gains, but also for exposing exactly where the technology still breaks down.

And it does break down, in specific and sometimes odd places. Small tweaks to familiar riddles can throw models off completely. Visual reasoning remains a soft spot across the board. What follows is a set of puzzles that have tripped up AI models at one point or another, a chance to see whether human intuition still holds an edge where machine pattern-matching doesn't.

Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time.

Why this matters

For anyone building products on top of these models, the puzzle results are a useful gut check against the marketing. Benchmarks built on math and coding tell us models can manipulate symbols; they say almost nothing about whether a system can rotate a shape in its head or track an object through occlusion. That gap matters if your roadmap includes robotics, AR, manufacturing QA, or anything else where a model has to reason about physical space rather than just describe it.

We'd push back on any pitch deck that leans on "world models" language without showing performance on tasks like these seven puzzles. It's also a reminder for researchers that visual input capability and visual reasoning are not the same claim, even though they get bundled together in product announcements. Until a model can beat a bored teenager at a spatial logic puzzle, treat claims about 3D understanding as aspirational.

Worth watching: whether any lab publishes puzzle-specific benchmarks instead of letting users discover the failures themselves.

Common Questions Answered

Why did Columbia University researchers use New York Times Connections puzzles to test AI models?

Connections puzzles serve as a meaningful benchmark for AI capabilities because they require semantic reasoning and pattern recognition beyond simple symbol manipulation. Unlike math and coding benchmarks, these puzzles test whether AI systems can understand relationships and group concepts meaningfully, similar to how humans approach visual and linguistic puzzles.

What was the dramatic improvement in AI model performance on Connections puzzles between late 2024 and early 2025?

In late 2024, even the best AI models could only solve 18% of New York Times Connections puzzles, but by early 2025, some models achieved near-perfect performance on these same puzzles. This rapid improvement demonstrates how quickly AI capabilities are advancing in puzzle-solving tasks.

How do puzzle benchmarks differ from traditional math and coding benchmarks for evaluating AI?

Math and coding benchmarks primarily measure whether models can manipulate symbols effectively, but they reveal little about spatial reasoning or physical understanding. Puzzle benchmarks like Connections test whether AI systems can rotate shapes mentally, track objects through occlusion, and reason about physical space—capabilities crucial for robotics, AR, and manufacturing applications.

Why should companies building AI products pay attention to puzzle-solving performance results?

Puzzle results provide a reality check against marketing claims by revealing gaps in AI reasoning capabilities that traditional benchmarks miss. For product roadmaps involving robotics, augmented reality, or manufacturing quality assurance, understanding whether models can reason about physical space rather than just describe it is essential for determining feasibility and performance.

How has the yardstick for measuring AI progress evolved since Arthur Samuel's 1959 checkers program?

Since Arthur Samuel coined the term "machine learning" in 1959, researchers have progressively moved from games like checkers to chess, then Go, and now to visual puzzles like Connections. This evolution reflects how AI benchmarks continue to advance as models master previous challenges, requiring increasingly sophisticated tests to measure genuine progress in artificial intelligence.

LIVE18:07OpenAI Researcher Warns AI Speed May Overwhelm Human Security Teams