Chinese AI lab MiniMax has taken an unusual approach to its latest release: it barely acknowledged it. On September 27, the company's official account confirmed what developers had already spotted in the product — a new model called M3.1-Flash-Preview had gone live inside MiniMax Code, the company's coding assistant. That was the extent of the launch: no model card, no benchmark table, no price per million tokens.

Chinese AI watchers on X had noticed the model appear in the product before the confirmation, according to Startup Fortune. The muted rollout is a striking contrast with how MiniMax shipped its last flagship. For more context on this story, see our ongoing AI news.

A Deliberately Quiet Debut

When MiniMax launched M3 in June, the company published architecture details, SWE-bench scores and a pricing sheet that undercut rivals by 90%. This time, the API endpoint for M3.1-Flash-Preview is gated, and the only way to try the model is inside MiniMax's own product, with reasoning traces and behavior as the main observable signals.

The low-key strategy suggests MiniMax may be treating the release less like a headline launch and more like an experiment: A/B testing a product feature and letting the internet find out. That could reflect confidence that the model will speak for itself once developers use it — or simply that this release is not the big one, but a faster, cheaper companion tier.

What the Base M3 Tells Us

The base M3 model offers a sense of what MiniMax can ship when it does want attention. It is a 428-billion-parameter mixture-of-experts model with roughly 23 billion active parameters per token, a one-million-token context window, and native image and video input. On SWE-bench Verified, MiniMax reported a score of 80.5%.

Flash is the version of the line built to sharpen that speed and price advantage for everyday coding — the quick bug fixes and small feature work developers run dozens of times a day, rather than the long-horizon agentic tasks reserved for the flagship. Splitting a model line this way, with a large frontier version and a fast, cheap companion, has become the standard playbook among Chinese labs competing for developer mindshare.

Shipping Into a Crowded Field

MiniMax is not releasing into a quiet market. Alibaba put out Qwen3.8-Max on August 3 — a 2.4-trillion-parameter flagship priced at $2 per million input tokens — then followed on August 12 with the first open-weight version of a Max-tier Qwen model. Zhipu's GLM-5.3 landed August 14 with a million-token context window of its own.

Developers have noticed the flood of capable, inexpensive options. According to CNBC's reporting on OpenRouter data, Chinese models accounted for 57% to 67% of total token usage on the platform for the week that included September 14 — up from just 6% to 13% back in February. Vercel saw a similar jump, with Chinese models' share of usage on its platform rising to 55% in August.

Washington Is Watching

The surge has drawn attention well beyond developer communities. Daniel Remler, a senior fellow in the technology and national security program at the Center for a New American Security, told CNBC that Chinese AI represents "real economic and security risks for the United States," warning against integrating Chinese models into critical workflows.

That scrutiny adds a layer of complexity to MiniMax's quiet release. A model that arrives without a model card, published evaluations or pricing gives independent developers less to work with when assessing quality and safety — even as the lab's products gain real usage share against American rivals.

Why the Quiet Launch May Be Strategic

Preview models gated inside a company's own product are a familiar pattern: they generate usage data and feedback before a wider API release, while limiting exposure if the model underperforms. For MiniMax, the muted debut of M3.1-Flash-Preview may also be a way to hold the spotlight for a larger launch — the company has more than doubled sales as it prepares a new flagship model, according to reporting cited by Startup Fortune.

For developers, the practical takeaway is simple: a new fast coding tier from one of China's leading labs is live now inside MiniMax Code, and the benchmarks and pricing that would normally frame such a release are still missing. Whether M3.1-Flash-Preview can extend the momentum that carried Chinese models to a majority share of OpenRouter traffic will depend on how it performs once developers start pushing real workloads through it.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →