A Mixture-of-Experts Design
According to TestingCatalog AI News, Inkling-Small is a Mixture-of-Experts (MoE) transformer containing 276 billion total parameters, with 12 billion active during any given inference. That architecture — where only a small slice of the model "wakes up" for each token — is the same efficiency trick that has made Chinese open-weight models like DeepSeek so cheap to run. Thinking Machines confirmed the full weights are being released on Hugging Face.
The model also ships with a context window of up to one million tokens and what the lab calls "variable thinking effort," ranging from minimal to "xhigh," letting developers trade compute for performance depending on the task. Fine-tuning is available through the company's Tinker platform, and text, image, and audio chat are accessible in the Tinker Playground.
Trained on NVIDIA GB300 Systems
Thinking Machines said Inkling-Small was trained on NVIDIA GB300 NVL72 systems and uses "substantially less compute" than the original Inkling model, which launched on July 15, 2026. That first model — covered by WIRED and TechCrunch as Thinking Machines' debut release — established the lab as a serious competitor almost overnight.
The path to Inkling-Small reveals a deliberate efficiency strategy. The lab began work on the smaller model only after training Inkling, using the breathing room to revise its pre-training data mix and machine-learning recipe. An earlier preview checkpoint was post-trained partly through on-policy distillation, with Inkling serving as the teacher, followed by roughly two weeks of scaled agentic-coding reinforcement learning.
Where It Wins — and Where It Doesn't
The results of that recipe are notable. Thinking Machines claims Inkling-Small overtakes the original Inkling on reasoning and agentic-coding benchmarks, while Inkling remains stronger in knowledge coverage and factuality. On the text-only Humanity's Last Exam, a notoriously difficult evaluation, the smaller model reportedly scored 31.6%, compared with 29.7% for Inkling — a real improvement despite the dramatic reduction in active parameters.
The trade-off is clear: Inkling-Small is the sharper coder and reasoner, while its larger sibling retains the broader knowledge base. For developers, that is a feature, not a bug — it means choosing the right tool for the job rather than paying for a single giant model on every task.
A U.S. Player in an Open-Weight Field Led by China
Perhaps the biggest implication is geopolitical. As CNBC and others have noted, open-weight frontier models have become a U.S.–China flashpoint, with Chinese labs like DeepSeek and Alibaba's Qwen giving away powerful models and undercutting Western pricing. Forkast framed the Inkling-Small release bluntly: "Open-Weights Competition Now Has a US Entrant."
Until now, the leading open-weight models have come almost exclusively from China, raising concerns in Washington that American labs were ceding an entire category of the market. Thinking Machines — a well-funded U.S. startup led by one of the most prominent figures in AI — changes that calculus. By releasing full weights on Hugging Face and supporting open fine-tuning, the lab is making a direct play for the developer community that Chinese open-weight models have cultivated.
Built for Developers, Not Just Benchmarks
Thinking Machines has paired the release with tooling aimed squarely at developers. Beyond the open Hugging Face weights, Databricks confirmed the Inkling family is available on its platform, and the Tinker Playground supports multimodal text, image, and audio chat out of the box. The variable "thinking effort" setting — from minimal to "xhigh" — is designed for real engineering trade-offs: run cheap when a task is simple, spend more compute when a problem demands deeper reasoning. That flexibility matters for startups watching their inference bills, many of whom have already migrated to cheaper open-weight alternatives rather than paying premium API rates.
The Efficiency Trend Across the Industry
Inkling-Small also fits a broader pattern. Across the industry, the cutting edge is shifting from ever-larger monolithic models toward smaller, smarter systems that punch above their weight. From DeepSeek's price-war models to compressed architectures, the message is consistent: raw parameter count matters less than how efficiently a model uses its capacity.
For Thinking Machines, the bet on efficiency over size is now two models deep — and the early benchmarks suggest the strategy is working. With Inkling-Small available for download and fine-tuning today, developers can test that claim for themselves.
Stay Ahead of AI
New open-weight models are landing every week, and the competition is only intensifying. Keep up with every release through our AI industry coverage.
Read more AI news →