The identity of Ox Alpha, the unbranded AI model that has dominated developer conversations since it appeared on OpenRouter and OpenCode on August 20, is no longer a mystery. China's Z.AI confirmed to Bloomberg News on Wednesday that the model is a new iteration of its GLM series — and said it will release the model's weights tonight, making them available for anyone to download and run.
The confirmation ends one of the most frenzied guessing games the AI community has seen this year, and it puts a flagship-grade Chinese model on a collision course with Western labs that price comparable capabilities at a steep premium. For more context on this story, see our ongoing AI trends.
A stealth launch that worked exactly as designed
Ox Alpha arrived with a listing that described everything about what it could do and nothing about who built it. OpenRouter attributed it to "a third-party provider who has chosen to remain anonymous during this preview," and positioned it as a reasoning model built for coding, sustained agentic work, and production workloads.
The specifications alone were enough to turn heads: a context window of just over one million tokens, support for text, image, and video input, and completely free access through both OpenRouter and OpenCode. OpenCode advertised serving capacity of 100 trillion tokens per day for the model, and Nous Research claimed its own portal could push through a quadrillion tokens.
Running a frontier-adjacent model anonymously on neutral infrastructure let Z.AI collect real usage data and unfiltered benchmark chatter before attaching its brand — a playbook that sidesteps the discount narrative that follows every open-weight Chinese release, as OfficeChai noted in its analysis of the launch.
The benchmark run that lit the fuse
The mania traces back to a single data point. Within a day of release, developer Ben Davis ran an informal ten-task evaluation on the DeepSWE benchmark and reported 80 percent first-pass accuracy for Ox Alpha — ahead of Claude Fable 5 at 65 percent and GPT-5.6 Sol at 52 percent.
That result was enough to send the model viral and to end DeepSeek's 56-day streak at the top of the OpenCode leaderboard. It also drew immediate pushback. On Hacker News, where Bloomberg's confirmation racked up hundreds of comments, developers urged caution, noting that a ten-task run is far too small a sample to settle anything and that the model's results on other leaderboards were noticeably more mixed.
How the internet cracked the case first
Long before Wednesday's confirmation, community detective work had already settled on Z.AI — formerly known as Zhipu AI — as the leading candidate. A developer going by "dax" ran a tokenizer fingerprinting test across 25 prompts and found that Ox Alpha's raw token counts matched GLM's tokenizer on 11 out of 11 probes, while no other lab's model cleared four. On one digit probe, DeepSeek's tokenizer burned 98 tokens where GLM used 29.
Separate analysis found that Ox Alpha stumbled on the same "dirty token" that has historically tripped up Qwen and GLM-family models, narrowing the field to a Chinese lab early on. The other prominent theory — Xiaomi, whose MiMo team previously ran an identical stealth-launch playbook with MiMo-V2-Pro under the alias Hunter Alpha — now appears to have been a decoy.
The stakes of the reveal extend beyond bragging rights. The New York Times reported on Tuesday that a Chinese lab's model release could test the world's cybersecurity readiness, underscoring how closely Western governments are watching open-weight releases of this caliber.
Why open weights change the calculus
Z.AI's confirmation that weights are coming tonight is the part that matters for the market. The company's current GLM 5.3 is priced at $1.40 per million input tokens and $4.40 per million output tokens on its first-party API, with an 81 percent cache discount. The new model is expected to ship under the same MIT license Z.AI has used since GLM 5, according to OfficeChai.
Frontier-adjacent coding and agentic performance, released openly at a fraction of Western API pricing, is precisely the pressure point that premium labs have spent two years trying to avoid. It also keeps Z.AI competitive with DeepSeek on the open-weight side of the Chinese market, where freely downloadable models have become the default expectation for anything short of a lab's absolute frontier.
What to watch next
Key details remain unconfirmed — most importantly the parameter count and model size, which will determine how practical the release is for local deployment. Community discussion has already zeroed in on how the model responds to quantization and how fast it runs on consumer hardware rather than data-center GPUs.
What is certain is that a week of anonymous speculation has given Z.AI something no marketing budget could buy: impartial, brand-agnostic validation on neutral ground. Now that the model has a name, the benchmarks are about to get a lot more crowded — and a lot more rigorous.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →