Chinese AI startup Moonshot released Kimi K3 on Friday, a 2.8 trillion-parameter open-weight model that third-party evaluators rank second overall behind only Anthropic’s Fable 5. The launch lands one month after the U.S. government abruptly pulled Anthropic’s Fable and Mythos models from the market over security concerns, and it shows how fast China’s open ecosystem is closing on the best American systems.
Kimi K3 also autonomously designed a functional chip that achieves over 8,700 tokens per second in inference workloads — a feat Moonshot says was completed in just 48 hours.
What Happened
Moonshot unveiled Kimi K3 on July 16, 2026, billing it as the world’s largest open-weight AI model. The system is a sparse mixture-of-experts design with 2.8 trillion total parameters, activating roughly 50 billion parameters for any given token by routing through 16 of 896 experts. It carries a 1-million-token context window and ships with what Moonshot calls Kimi Delta Attention, a mechanism the firm says decodes up to 6.3 times faster over million-token inputs.
Third-party evaluators ranked Kimi K3 first on web interface building and second overall behind only Anthropic’s Fable 5, ahead of OpenAI’s GPT-5.6 Sol. The company claims the model was able to autonomously design a fully functional chip within 48 hours that delivers over 8,700 tokens/s for inference workloads.
The model arrives weeks after Moonshot was reported to be seeking a $30 billion valuation. It also follows the U.S. government’s abrupt removal of Anthropic’s frontier models from the market, creating a vacuum that Chinese open-weight models are rushing to fill. Moonshot, along with other Chinese labs like Z.ai and MiniMax, are shipping stronger models at sharply lower prices.
My Take
Let’s be direct: the U.S. pulling frontier models from the market created an opening, and China just drove a truck through it. Moonshot’s timing is surgical — one month after Anthropic’s models disappeared, they drop an open-weight system that competes with the best closed models. If you’re a developer who relied on Fable for your pipeline, Kimi K3 is now your most viable alternative.
The architecture is genuinely impressive. The sparse MoE design with 896 experts and only 50 billion activated parameters per token means inference costs stay manageable despite the massive total parameter count. The 6.3x faster decoding on long contexts via Kimi Delta Attention is a real differentiator for knowledge work and long-horizon coding tasks. But the chip design demo is the real flex — it’s one thing to claim reasoning ability on benchmarks, another to have the model autonomously produce a working chip.
The convergence is the story here: three Chinese labs shipping frontier-competitive models at Chinese prices changes the economics of AI development. Open-weight models at this scale mean startups can now build on top of a capability that was locked behind API walls just months ago.
What to Watch
- The distillation question: Early reports show Kimi K3 sometimes identifies itself as Anthropic’s Claude in conversations, hinting at its training lineage. If regulators see this as IP theft, export controls could tighten further.
- Inference costs cratering: With open-weight models at this scale available at Chinese price points, the economics of running frontier AI will shift dramatically in the second half of 2026.
- U.S. response to the vacuum: The removal of frontier models was supposed to be a security measure. If it only accelerates China’s lead in open-weight AI, expect policy adjustments within months.
