Moonshot AI released Kimi K3 on July 16. The model was available immediately through Kimi’s web and mobile products, Kimi Work, Kimi Code and the API. Eleven days later, Moonshot published the full weights and technical report. The headline specifications are hard to miss: 2.8 trillion parameters, a one-million-token context window and native vision.
The capacity notice was just as revealing as the model card. On July 19, Moonshot said demand over the previous 48 hours had pushed close to its GPU limit, so it paused new subscriptions and protected access for existing members. New models draw crowds all the time. Rationing sign-ups three days after launch says something more useful: serving capacity is now part of model competition.
DeepSeek was pushing the V4 line at the same time, while Alibaba Cloud had started offering Qwen3.8-Max-Preview. “DeepSeek is cheap” no longer captures the Chinese AI market. Performance, price, licensing and cloud distribution are moving together.
That is why I put these releases side by side. The practical question is no longer whether a Chinese model can win one benchmark. It is whether leaving Chinese models off a serious procurement shortlist has become an obvious mistake.
Kimi K3’s numbers are impressive. They are still vendor numbers
Kimi K3 does not activate all 2.8 trillion parameters for every prompt. Its Mixture of Experts architecture routes each token through 16 of 896 experts. The point is to get the capacity of a very large model without paying the full compute cost on every pass.
Moonshot says K3 competes closely with leading models across coding, browsing, automation and knowledge work. The company also points to Frontend Code Arena, a public leaderboard built from human preference votes, where K3 ranked highly after release. Both are useful signals, but both still need independent replication.
That is where I would stop before repeating the ranking. The chart on this page comes from Moonshot. Different models were tested through Kimi Code, Claude Code or Codex, and reasoning settings and fallback behavior were not identical. Moonshot’s own release says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall.
That admission makes the document more credible. It tells us where K3 is competitive without pretending that one company-controlled table settles the market. Any headline that turns the same chart into “China now has the best model” is doing more work than the evidence can support.
K3’s stronger case is breadth. One system handles long documents, images, coding tools and a one-million-token context window, and the API was available on release day. The limitations deserve equal attention. Moonshot warns that output can become unstable if thinking history is not carried forward correctly or if a session switches models midway. It also warns that the model can act too proactively and recommends tighter boundaries in the system prompt or AGENTS.md.
Those details matter more to an operator than a two-point benchmark lead.
Open weights are not the same as free or unrestricted
Moonshot released the full K3 weights on July 27. Researchers can inspect them, and inference providers can build services around them. That does not make K3 an unconditional open-source model.
K3 uses a custom license. It allows research, development and a broad range of commercial uses, but large Model-as-a-Service providers cross into separate-contract territory. Attribution obligations also apply. “The weights are public” and “any company can do anything with them” are different statements.
Infrastructure remains a separate bill. Serving a 2.8-trillion-parameter model requires GPUs, an inference stack and people who can keep it running. Moonshot recommends supernode configurations with at least 64 accelerators. Downloading the files does not turn this into an inexpensive self-hosted option for an ordinary company.
The real leverage of open weights is therefore not that everyone runs K3 for free. It is that several inference providers can serve the same model and compete on price, region, privacy controls and availability.
DeepSeek changes the price conversation first
As checked on August 4, DeepSeek lists V4-Pro at $0.435 per million cache-miss input tokens and $0.87 per million output tokens. Kimi K3 lists cache-miss input at $3 and output at $15. Even within China’s model market, the spread is large.
Token price is not total cost. A model that consumes more tokens, responds slowly or creates more review work can erase its headline saving. But for high-volume text workloads, the difference is too large to dismiss.
OpenRouter’s own usage report shows the direction. Across more than 450 trillion tokens handled on its platform from January 1 through June 14, DeepSeek’s share rose from about 9% in January to roughly 18% in early June. This is not global market share. It is token volume on one routing platform. It is still useful evidence of where developers move traffic when several models are one API call away.
If I were writing a model business case, I would start here rather than with a three-point benchmark gap. Monthly input, output, retries, latency and human correction time can reverse the ranking on a real invoice.
Qwen3.8 needs observation before judgment
Qwen3.8-Max-Preview is available through Alibaba Cloud Model Studio’s Token Plan. The official page explains how customers can access the preview and where it fits in Alibaba’s tooling; it does not provide enough verified technical detail for a final comparison with K3 or DeepSeek.
The word “Preview” matters. On August 4, I could not find a final model card, released weights, a settled license and a broad set of independent evaluations on Alibaba’s official pages. Putting it in a definitive winner-and-loser table beside K3 and DeepSeek would be premature.
What stands out first is distribution. Qwen sits inside Alibaba Cloud and Model Studio, so existing customers can try it without rebuilding their stack. That does not prove adoption, but it lowers the cost of evaluation. Model quality and the route into daily use are separate competitive advantages.
The most immediate pressure on U.S. labs is price
The price sheets cited here support one clear claim: DeepSeek V4-Pro is markedly cheaper than Kimi K3. A broader comparison with U.S. models needs model-by-model pricing from the same date, including cache rules and service tiers. Even without that shortcut, a gap this large makes premium pricing harder to defend on benchmark performance alone.
I would not turn OpenRouter’s numbers into a global market-share claim. They cover activity on one platform. Even so, DeepSeek’s share there rose from roughly 9% to 18% in the period measured. That is enough to show that some users will reallocate traffic when price and performance line up.
Open weights widen the buying path as well. Instead of relying on one API, a team can compare inference providers by price, region and data policy, or host the model when the operating case justifies it. Competition therefore moves beyond benchmark scores to the total bill and the reliability of supply.
How I would shortlist them today
These models are not interchangeable. My current shortlist would look like this:
| Job | First candidate | Conditions to check |
|---|---|---|
| High-volume text at low cost | DeepSeek V4-Pro | Output length, latency, retry rate and data-processing location |
| Long context, vision and coding in one model | Kimi K3 | Kimi Code compatibility, no mid-session model switch, license and capacity |
| A quick trial inside Alibaba Cloud | Qwen3.8-Max-Preview | Preview changes, final pricing, model card and weight release |
| Customer or regulated data | Choose the provider before the model | Storage region, log retention, deletion policy and contractual responsibility |
For cheap, high-volume text, I would benchmark DeepSeek first. For work that combines long context, images and coding, I would test K3 and turn Moonshot’s own harness warnings into acceptance criteria. I would not standardize on Qwen3.8 until the preview ends and the final documentation is available.
I would also keep at least two providers viable. Pricing and access policies move quickly. I would regularly run the documents and requests our team actually sees through two or three models. Completion cost and human correction time tell me more than a launch-week leaderboard or one accuracy score.
So how far has China’s AI industry come?
It is too early to say China has won. Kimi’s own material acknowledges the gap with the strongest proprietary systems. DeepSeek’s price is formidable, but its quality and efficiency will not be identical across every job. Qwen3.8 remains a preview.
Something has changed nonetheless. Chinese models no longer fit neatly into the “cheap alternative” box. Capability is rising, pricing is lower, weights are increasingly available and distribution channels are mature. A shortlist containing only U.S. products can now miss a material part of the market.
My conclusion is operational rather than patriotic. Put Chinese models on the default shortlist, but separate company benchmarks from independent evidence, public weights from unrestricted use, and token price from the cost of completing the job. China’s biggest AI achievement this year is not a single number-one model. It is the arrival of several options that serious buyers can no longer ignore.

References and reporting
- Kimi K3: Open Frontier Intelligence Moonshot AI
- Kimi K3 technical report Kimi Team / arXiv
- Kimi K3 License Moonshot AI / Hugging Face
- DeepSeek V4 Preview Release DeepSeek
- DeepSeek Models and Pricing DeepSeek
- Qwen3.8-Max-Preview availability Alibaba Cloud
- DeepSeek V4 Is Earning Agentic Token Share OpenRouter
- Moonshot pauses Kimi subscriptions amid hot demand Reuters