Skip to main content
Inkling AI model, represented by a small brain icon, outperforms a larger model on coding tests.

Editorial illustration for Thinking Machines' Inkling Small Beats Larger Model on Key Coding Tests

Thinking Machines' Inkling Small Outperforms Larger Models

Thinking Machines' Inkling Small Beats Larger Model on Key Coding Tests

4 min read

Mira Murati's Thinking Machines has released its second model, Inkling Small, and the pitch is smaller footprint over raw scale. The open-weights reasoning model carries 276 billion total parameters with just 12 billion active, less than a third of what its predecessor Inkling uses. According to Artificial Analysis, Inkling Small scores 40 on the Intelligence Index, a single point behind Inkling's 41, and the firm says no open model of equal or smaller size posts a higher score.

The gap between the two models isn't uniform across tasks. Inkling Small edges ahead on some coding and reasoning benchmarks while trailing on others, a split that speaks to how Thinking Machines is trading breadth for efficiency. The model also handles text, image, and speech inputs, runs on a 256K-token context window, and ships under an Apache 2.0 license. Weights are available on Hugging Face, and the company built browser-based fine-tuning into Tinker Playground, reinforcing its broader bet that these models work best as a base users adapt with their own data rather than a finished product meant to be used as-is.

According to Artificial Analysis, the open-weights reasoning model scores 40 on the Intelligence Index, one point below Inkling (41), with less than a third of the parameters (276 billion total, 12 billion active). AA says no open model of equal or smaller size scores higher.

Why this matters

Inkling Small tells us something more interesting than another leaderboard flex: parameter count is losing its grip as the proxy for capability. A 12-billion-active-parameter model outscoring its own 41-point sibling on Humanity's Last Exam, while burning far fewer tokens per answer, is a compute bill founders should care about, not just a research curiosity. For teams running inference at scale, 24K average output tokens versus a bulkier reasoning model translates directly into latency and cost.

Thinking Machines shipping this as open weights also matters for researchers who want to poke at why a smaller model wins on coding and reasoning but loses on agent-based tasks and factual recall. That split is worth watching closely. It suggests Inkling Small was tuned for a particular kind of problem-solving rather than general competence, and buyers should read the Artificial Analysis numbers with that trade-off in mind, not as a blanket "smaller is better" verdict.

The real test comes when developers start swapping it into agent pipelines and see where the factual gaps actually bite.

Common Questions Answered

How does Inkling Small's parameter efficiency compare to its predecessor Inkling?

Inkling Small uses only 12 billion active parameters out of 276 billion total, which is less than a third of what its predecessor Inkling requires. Despite this significant reduction in active parameters, Inkling Small scores only one point lower on the Intelligence Index (40 versus 41), demonstrating that parameter count is no longer the primary determinant of model capability.

What makes Inkling Small competitive among open-weights reasoning models?

According to Artificial Analysis, Inkling Small scores 40 on the Intelligence Index, and Thinking Machines claims that no open model of equal or smaller size achieves a higher score. This positions Inkling Small as a leading option for developers seeking efficient reasoning capabilities without the computational overhead of larger models.

Why does Inkling Small's efficiency matter for teams running inference at scale?

Inkling Small generates an average of 24K output tokens while consuming significantly fewer computational resources than bulkier reasoning models, which translates directly into reduced compute costs for teams running inference at scale. This efficiency improvement represents a meaningful difference in operational expenses for organizations deploying large-scale AI systems, making it more than just a research achievement.

What does Inkling Small's performance reveal about the relationship between model size and capability?

Inkling Small's ability to outperform larger models on benchmarks like Humanity's Last Exam while using far fewer active parameters demonstrates that parameter count is losing its grip as the primary proxy for AI capability. This shift suggests that architectural efficiency and training methodology may be becoming more important factors than raw model scale in determining overall performance.

LIVE20:44Chinese AI Researchers Turn to X for Technical Audience