Chinese AI lab Moonshot AI unveiled Kimi K3 on July 16, 2026, releasing what it calls "the world's first open 3T-class model" — a 2.8-trillion-parameter system that narrows the gap with the most powerful proprietary models from Anthropic and OpenAI. The launch immediately became one of the most discussed AI releases of the year, reaching the top of Hacker News and drawing coverage from Reuters, Bloomberg, and The New York Times.
For the broader AI industry coverage, Kimi K3 matters because it pushes the frontier of open-weight models sharply higher while forcing a fresh debate over whether freely downloadable systems can match closed ones. Moonshot says the full model weights will be released by July 27, 2026.
A 2.8-Trillion-Parameter Open Model
According to Moonshot's technical announcement, Kimi K3 is a 2.8-trillion-parameter model built on two architectural updates the company calls Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). It uses a Mixture-of-Experts (MoE) design that activates only 16 of 896 experts for any given token, paired with what Moonshot describes as a Stable LatentMoE framework. The company claims these changes yield roughly a 2.5x improvement in overall scaling efficiency compared with its previous Kimi K2 model.
The model also ships with native vision capabilities and a one-million-token context window. Analyst Simon Willison, writing on his weblog, noted that Kimi K3 takes the open-size crown from DeepSeek's 1.6-trillion-parameter v4 Pro — "more than twice the size" of the 1-trillion-parameter Kimi K2.6 that preceded it.
Moonshot frames the release as the latest step in a sustained scaling push, saying that for nine of the past twelve months its models have set the upper bound of open-model sizes.
How It Compares on Benchmarks
Moonshot's own evaluation suite places Kimi K3's overall performance behind only two proprietary systems — Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol — while it consistently outperforms other tested models, including Claude Opus 4.8 and GPT-5.5. Third-party analysis broadly tracks those claims.
The independent evaluation service Artificial Analysis reported that on its private long-horizon knowledge-work benchmark, Kimi K3 reached an overall Elo of 1547 — a jump of 732 points over Kimi K2.6 and behind only Claude Fable 5. On Arena.ai's Frontend Code arena, Kimi K3 climbed to the number-one spot, surpassing even Claude Fable 5, according to a leaderboard post that circulated widely after launch.
Pricing is a notable part of the story. Artificial Analysis pegged the average cost per task at $0.94, comparable to GPT-5.6 Sol's $1.04 and roughly half of Claude Opus 4.8's $1.80, while noting that Kimi K3 used 21% fewer output tokens than its predecessor. On a per-token basis the model is priced at $3 per million input tokens and $15 per million output tokens — the same tier as Anthropic's Claude Sonnet series and, as Willison observed, the most expensive model released by a Chinese AI lab to date, up from Kimi K2.6's $0.95/$4.
Agentic Coding and a Compiler It Built Itself
Much of Moonshot's announcement focused not on chatbot benchmarks but on long-horizon agentic tasks. The company says Kimi K3 can sustain extended engineering sessions, navigate large code repositories, and orchestrate terminal tools with minimal human oversight.
In one demonstration, Moonshot said an early version of Kimi K3 handled the majority of the team's own GPU kernel optimization work during late-stage development. In another, the model built "MiniTriton," a compact Triton-like GPU compiler with its own tile-level intermediate representation, optimization passes, and PTX code-generation pipeline. Moonshot claims MiniTriton matched or beat the established Triton and torch.compile stacks on supported roofline benchmarks and sustained end-to-end nanoGPT training with stable convergence.
The company also showcased a striking proof of concept in chip design. In a single 48-hour autonomous run, Kimi K3 reportedly designed, optimized, and verified a chip built to serve a nano model on its own architecture, using open-source EDA tools on the Nangate 45nm library. Moonshot said the design closes timing at 100 MHz within 4 mm² and sustains over 8,700 tokens per second of decode throughput in simulation.
On the research side, Moonshot said Kimi K3 reproduced the I-Love-Q universal relations from computational astrophysics in roughly two hours — work the company estimated would typically take a researcher one to two weeks — by cross-validating more than 20 papers, implementing a full numerical pipeline across 300-plus equations of state, and writing over 3,000 lines of Python.
The Open-Weight Debate Returns
Kimi K3's arrival reignites a debate that has run through 2026 as Chinese labs have repeatedly closed — or appeared to close — the distance with U.S. frontier models. The pattern is now familiar: a Chinese release posts self-reported numbers near the top of leaderboards, Western labs emphasize their still-growing proprietary edge, and businesses weigh cost and openness against raw capability.
That tension is visible in the pricing. A model that undercuts Claude Opus 4.8 on cost-per-task while approaching its capability, yet charges more than any previous Chinese release, reflects Moonshot's confidence that it can command premium pricing even in an open ecosystem. Willison cautioned against reading too much into any single benchmark — his long-running "pelican on a bicycle" test, he wrote, no longer correlates reliably with real-world quality — but said the release still rewarded hands-on testing.
Kimi K3 is available now through Kimi.com, the Kimi Work desktop agent, the Kimi Code developer tool, and the Kimi API. Moonshot said it is working with inference partners and open-source maintainers to ensure a reliable rollout ahead of the full weights release.
Stay Ahead of AI
For ongoing coverage of model releases, benchmark battles, and the open-weight frontier, follow AI Buzz Wire for the developments shaping the global artificial intelligence race.
Read more AI model news →


