Thinking Machines is making the case for smaller reasoning models

Thinking Machines, the AI lab founded by former OpenAI CTO Mira Murati, has released Inkling Small, an open-weights reasoning model designed to compete on capability while using a much smaller active footprint than the company’s earlier Inkling model. The launch signals a familiar but increasingly important shift in the AI market: efficiency is becoming a product feature in its own right, not just a technical footnote.

According to the supplied source text, Artificial Analysis gives Inkling Small a score of 40 on its Intelligence Index, just one point below Inkling’s 41. That gap is narrow enough to make the headline clear. The new model is positioned as a smaller system that approaches the larger model’s overall standing while improving on it in selected coding and reasoning benchmarks.

The model’s reported size profile is central to the pitch. Inkling Small is described as having 276 billion total parameters with 12 billion active, which the source says is less than a third the size of the original Inkling. Artificial Analysis also says no open model of equal or smaller size scores higher on its index.

Benchmark results suggest targeted gains

The benchmark story is not simply that Inkling Small is nearly as strong overall as its predecessor. The source text says it outperforms the larger Inkling on several coding and reasoning tests. Two examples cited are Humanity’s Last Exam, where Inkling Small scores 32% compared with Inkling’s 30%, and GPQA Diamond, where it reaches 89% against Inkling’s 87%.

Those numbers matter because they suggest the model is not merely a compressed version that preserves most capability. In at least some tasks, it appears able to surpass the larger system. That kind of result is notable in a market where bigger models have often been treated as the default route to higher performance.

At the same time, the source does not present Inkling Small as a universal upgrade. It reportedly falls behind the original Inkling on agent-based tasks and factual knowledge. That limitation is important because it keeps the launch grounded in tradeoffs rather than hype. Users looking for stronger autonomous task execution or broader recall may still prefer a larger model, even if the smaller one is more attractive on efficiency.

Efficiency is the deeper competitive message

The stronger differentiator may be token efficiency. The source says Inkling Small averages 24,000 output tokens per task, compared with 45,000 for Deepseek V4 Flash and 78,000 for GPT-5.4 mini. On its face, that suggests a model that can do substantial reasoning work with less generated output than some competing systems.

That kind of efficiency matters for both economics and product design. Lower output token usage can reduce inference cost and latency, and it may simplify deployment in environments where throughput or budget matters as much as raw benchmark strength. For companies building AI features into products, a model that does more with fewer output tokens can become attractive even if it does not lead every leaderboard.

The launch therefore points to a broader shift in how AI vendors frame progress. Rather than only emphasizing absolute frontier performance, developers are increasingly competing on the quality-per-token equation. Inkling Small fits that pattern by offering a model that, based on the supplied figures, keeps close to its larger sibling on general standing while improving its efficiency profile.

Open weights and multimodal input broaden appeal

Thinking Machines is also trying to widen the model’s practical reach. The source text says Inkling Small supports text, image, and speech inputs and comes with a 256,000-token context window. It is released under the Apache 2.0 license, with weights hosted on Hugging Face.

Those details matter because they shape who can use the model and how quickly it can be adapted. An Apache 2.0 release lowers friction for experimentation and commercial use. Multimodal input support makes the system relevant to a broader set of applications than text-only reasoning. And a long context window increases its appeal for tasks that require handling large documents, long conversations, or extensive working memory.

The source also says users can fine-tune the model in the browser through Tinker Playground. That feature aligns with Thinking Machines’ positioning of its models as a foundation for customization with users’ own data. In other words, the company is not only selling a base model story. It is selling a workflow in which users adapt that base model to narrower goals.

Why the release matters in the current model market

Inkling Small arrives at a moment when AI developers are under pressure to show practical value, not just scale. Training ever-larger models remains expensive, and product teams deploying them have to weigh cost, speed, flexibility, and licensing alongside capability. That creates room for models that may not be the largest or most universally powerful, but are good enough on key tasks and more efficient to run.

The source frames Inkling Small as exactly that kind of offering: a smaller open-weights reasoning model that performs unusually well for its size. If that positioning holds up under wider use, it could strengthen the argument that model development is entering a more disciplined phase, where optimization and deployment value matter as much as raw expansion.

It also keeps attention on a growing segment of the market: open models that can be fine-tuned and embedded into custom workflows. Proprietary frontier systems still dominate many headline comparisons, but open-weight releases remain strategically important because they give organizations more direct control over how a model is adapted, hosted, and governed.

A measured step rather than a maximalist one

The most interesting thing about Inkling Small may be what it does not try to claim. Based on the supplied text, Thinking Machines is not presenting it as the single best model in every category. Instead, the release argues for a narrower, more pragmatic proposition: close to top-tier performance for its size, selected benchmark wins over a larger sibling, multimodal capability, long context, and a notably leaner output-token profile.

That makes the launch less about spectacle than about engineering priorities. Smaller active models that can reason well, stay open, and reduce token burn are likely to remain attractive as AI deployment matures. Inkling Small appears to be a direct bet on that future.

Whether it becomes a widely adopted foundation model will depend on real-world testing beyond the summary figures supplied here. But from the available information, the strategic message is already clear. Thinking Machines is trying to prove that in AI, efficiency is no longer a compromise category. It is becoming a serious competitive lane of its own.

This article is based on reporting by The Decoder. Read the original article.

Originally published on the-decoder.com