2.8 trillion parameters. The largest open-weight model released to date. It topped Arena’s front-end coding leaderboard within days of launch — beating both GPT-5.6 Sol and Claude Fable 5 on that specific benchmark. Then Moonshot AI ran out of servers. “Kimi K3 has received far more love than we expected.”
Moonshot AI’s Kimi K3 capacity crunch became public, when the Beijing-based startup announced it had paused new subscriptions just days after launch. The new, powerful Chinese artificial intelligence model, which has caused a stir in the US tech industry, has suspended new subscriptions after a flood in demand overwhelmed capacity within days of the launch. “Over the past 48 hours, demand [swelled] close to the limits of our current capacity,” Moonshot AI wrote in a statement. “Kimi K3 has received far more love than we expected,” the company added in a separate post on X. Existing subscribers face no disruption — the pause applies only to new signups, which will resume gradually “in batches” as Moonshot AI adds capacity. By contrast, the episode highlights a specific and recurring structural challenge facing Chinese AI labs: building models capable of competing with US frontier systems, while lacking the server infrastructure to serve the global demand that competitiveness generates.
What’s Happening & Why It Matters
2.8 Trillion Parameters — the Largest Open-Weight Model Yet
Moonshot AI’s Kimi K3 capacity crunch stems from a model that is unprecedented in scale among openly available systems. Moonshot launched Kimi K3 last week as a 2.8 trillion-parameter model, making it the largest open-weight system released to date. Chinese frontier models tend to be open weight, meaning developers can examine and adapt the underlying system rather than access it only through a paid interface — a fundamentally different distribution model from the closed, API-only approach OpenAI and Anthropic generally favour for their most capable systems.
According to the South China Morning Post, Kimi K3 has outperformed rivals including GPT-5.6 Sol and Claude Fable 5 on some benchmarks — including Arena‘s ranking for front-end coding capability, where it topped the leaderboard following its public release. By contrast, Moonshot AI itself was careful not to overstate that result — the company acknowledged its model delivered frontier-level performance across its evaluation suite while still lagging behind Claude Fable 5 and GPT-5.6 Sol on other benchmark categories.
Why the Servers Ran Out

Moonshot AI’s Kimi K3 capacity crunch is a specific technical reality behind the demand surge. Su, addressing the shortfall, said the key reason was more likely due to Moonshot not fully anticipating the surge in K3’s popularity. K3 is “very demanding” in terms of compute requirements, making compute allocation challenging and expensive. A 2.8 trillion-parameter model requires substantially more inference infrastructure per user query than smaller systems — meaning even a modest surge in daily active users can overwhelm capacity planned around considerably more conservative demand estimates.
Additionally, Moonshot confirmed Kimi K3 launched simultaneously across multiple access points — Kimi.com, Kimi Work, Kimi Code, and the Kimi API — using max thinking effort by default, with lower- and higher-effort modes planned for subsequent updates. Launching at maximum computational intensity across four simultaneous platforms, rather than a single controlled surface, compounded the capacity strain the sudden demand spike created.
A Familiar Scenario — and Real Market Consequences
Moonshot AI’s Kimi K3 capacity crunch fits a documented and recurring pattern across Chinese AI labs throughout 2026. Chinese AI labs keep building models good enough to grab headlines worldwide, then discover they do not have the servers to handle the stampede that follows. Kimi K3 is part of a wave of low-cost, openly available Chinese models that have gained traction internationally in recent months, following DeepSeek‘s V4 release and the market shock caused by DeepSeek’s original model in early 2025 — a moment that first prompted many observers to view China as a AI rival to the US.
By contrast, K3’s release carried consequences well beyond its own capacity constraints. The launch put pressure on stocks of American technology titans on worries that more affordable Chinese AI models may undercut the pricing power and demand at US AI companies — even as US-led restrictions have already barred China from accessing some of the world’s most advanced chips. As TF covered in its Anthropic Alibaba distillation attack article, that competitive pressure between US and Chinese labs has been building throughout 2026, spanning direct market competition, alleged distillation attacks, and export control disputes simultaneously.
Full Model Debuts 27 July. Open-Weight Access Continues Regardless
Moonshot AI’s Kimi K3 capacity crunch does not affect the model’s open-weight release timeline. The full model weights will come out by 27 July 2026, allowing developers worldwide to download and run Kimi K3 independently of Moonshot‘s own hosted infrastructure entirely — a important distinction from the signup pause. Anyone with sufficient local compute capacity will be able to access K3’s full capabilities without needing a Moonshot account at all, once those weights are published.

The startup has confirmed it plans to introduce two focused membership plans designed for better resource allocation once new signups resume — suggesting Moonshot intends to segment future demand more deliberately than its initial undifferentiated launch structure allowed, likely separating lighter conversational usage from the heaviest coding and agentic workloads that consume disproportionate compute per user.
TF Summary: What’s Next
New signups to Kimi K3 are paused, resuming gradually “in batches” as Moonshot AI adds capacity. Existing subscribers continue accessing the model without disruption. Full model weights are scheduled for public release by 27 July 2026. Two new focused membership plans are planned for introduction once capacity stabilises. Low- and high-effort computational modes are planned for a future update, alongside the current max-effort default.
MY FORECAST: Moonshot AI’s Kimi K3 capacity crunch will resolve within four to six weeks — the pattern established by DeepSeek’s earlier capacity struggles suggests Chinese AI labs can scale infrastructure reasonably quickly once demand signals arrive, even if initial launch planning consistently underestimates global interest. By contrast, the open-weight release on 27 July is the more consequential development long-term. Once developers can run Kimi K3 locally or through third-party hosting providers, Moonshot’s own signup capacity is largely irrelevant to the model’s actual global adoption and influence. Expect Kimi K3’s coding benchmark performance specifically to accelerate adoption among developer tools and coding assistants seeking a competitive, lower-cost alternative to closed US models — precisely the pricing pressure that already spooked American technology stocks within days of the model’s initial release.
Related Stories
- Anthropic Accuses Alibaba of 28.8 Million Fake Queries to Clone Claude
- China Tops Global Supercomputer Rankings and Opens the World’s First Underwater Data Centre
- Mythos 5 Returns for Critical Infrastructure — Fable 5 Remains Suspended
