Thinking Machines Lab, the AI startup founded by former OpenAI Chief Technology Officer Mira Murati, released its first foundation model on July 15, 2026. The model, named Inkling, makes its full open weights available to developers who want to fine-tune it for their own use cases — a deliberate bet on customization over the one-size-fits-all approach that dominates the AI industry. The launch marks the first public proof point for a company that has spent over a year building AI infrastructure largely out of public view, and it arrives at a moment when Western enterprises are hungry for breaking AI news and credible alternatives to both gated proprietary models and low-cost Chinese open-source systems.
What Is Inkling?
Inkling is the first model fully trained from scratch by Thinking Machines. It is a mixture-of-experts (MoE) architecture featuring 975 billion parameters, although for the average prompt it draws on only about 41 billion of those parameters in order to process tasks faster and keep costs low. According to the company, the model was trained on approximately 45 trillion tokens of text, image, audio, and video, and can reason natively across all four input types. Its outputs, however, are limited to text — though that text can include code, styled artifacts, and structured data.
The model offers "thinking effort" controls that let developers make tradeoffs, such as sacrificing processing speed for accuracy. Unusually, Inkling also flags its own outputs for uncertainty rather than simply generating confident-sounding hallucinations. Developers can fine-tune the model directly on Tinker, the company's training API that launched in October 2025.
A Western Alternative to Chinese Open-Source Models
One of the most significant aspects of the launch is what it represents in the broader open-weight ecosystem. For the past year, that space has been dominated by Chinese AI firms such as DeepSeek and Alibaba, while Meta downplayed its Llama family of models in favor of a more proprietary approach. Inkling is positioned as the first credible Western alternative to those systems.
Futurum Group analyst Mitch Ashley told the Wall Street Journal that Inkling gives Western enterprises an alternative focused on customization economics — shifting spending from per-token API pricing to infrastructure that the enterprise itself controls. "Engineering teams should treat base-model selection as an architecture decision," Ashley said. "The model an organization fine-tunes becomes part of its software substrate, and switching costs compound with every downstream customization."
Honest About Its Limitations
Thinking Machines has been unusually transparent about the fact that Inkling is not the strongest model on the market. The company acknowledged that it does not match the most advanced proprietary AI systems available. Instead, it is betting that customizability will make up the difference: rather than releasing Inkling as a rigid chatbot-style app, the company positioned it as a base model that organizations should fine-tune and run on their own infrastructure.
The strategy already shows promise. In a collaboration with Bridgewater Associates, researchers used Tinker to fine-tune an open model with specialized financial data. The result was a low-cost, lightweight model that scored 84.7% on leading financial reasoning benchmarks — outperforming the most advanced proprietary alternatives at less than 10% of the cost.
Speed and Silicon
Thinking Machines said it was able to develop Inkling from scratch in less than nine months, a timeline that compares favorably to the multiyear development cycles seen at rivals like OpenAI and Anthropic. The model was trained on Nvidia's GB300 NVL72 system under a partnership between the two firms that was announced in March 2026.
In early benchmark results, Thinking Machines showed that Inkling achieved coding performance comparable to Nvidia's Nemotron 3 Ultra model while using two-thirds fewer tokens — a strong indicator of efficiency for cost-conscious organizations.
The Business Model: Tinker, Not Tokens
Rather than charging customers for access through a metered API, Thinking Machines plans to generate revenue primarily through Tinker, its paid service that makes it simple for developers to fine-tune open-weights models for specific tasks. This represents a direct challenge to the gated, per-token access model pioneered by Silicon Valley's largest AI firms.
For Murati, who left OpenAI in September 2024 and has long insisted that her new company is about accessibility, customization, and multimodal collaboration, Inkling is a statement of intent. The company is wagering that the future of AI lies not in a single dominant model that everyone pays to query, but in a world where organizations own and shape the models that power their most critical work.
Whether that bet pays off will depend on whether enough enterprises are willing to invest in fine-tuning and infrastructure rather than the convenience of a managed API. But with Inkling now available on Hugging Face and Tinker open for business, Thinking Machines has given the open-weight movement its most significant Western entry point yet.
Stay Ahead of AI
Want more AI industry coverage like this? AI Buzz Wire tracks every major model release, funding round, and policy shift so you don't have to.
Read more AI news →


