Why Goldman Sachs Endorsed China's AI Pricing Power: The Kimi K3 Moment
On July 17, 2026, Beijing-based Moonshot AI unveiled Kimi K3, a 2.8-trillion-parameter open-weight model that instantly topped global coding leaderboards. Within 48 hours, demand overwhelmed its servers, Elon Musk called it "Impressive," and Goldman Sachs published a report that may mark a turning point for how the world understands China's AI industry.
But what made Wall Street's most influential bank pay attention wasn't just the benchmark scores. It was the price tag.
Goldman's Core Thesis: Who Can Charge More
On July 18, Goldman Sachs released "China AI Model Navigation." The headline was not that Kimi K3 had beaten GPT-5.6 Sol on Program Bench (77.8 vs. 77.6), or that it topped Arena.ai's coding rankings at 1,679. It was that Moonshot AI had set the blended API price at $2.30 per million tokens — the highest ever for a Chinese large language model.
The context is revealing. Alibaba's Qwen3.7 Max charges $1.40. Zhipu's GLM-5.2 costs $0.90. MiniMax's M3 is $0.22. DeepSeek V4 Pro, the budget developer's favorite, costs just $0.18. Kimi K3 is priced at roughly 13 times DeepSeek's rate and more than double any Chinese competitor.
| Model | Blended API Price | Relative to K3 |
|---|---|---|
| Kimi K3 | $2.30/M tokens | 1.0× |
| Qwen3.7 Max | $1.40/M tokens | 0.61× |
| GLM-5.2 | $0.90/M tokens | 0.39× |
| MiniMax M3 | $0.22/M tokens | 0.10× |
| DeepSeek V4 Pro | $0.18/M tokens | 0.08× |
| Claude Fable 5 | $15.00/M tokens | 6.5× |
| GPT-5.6 Sol | $15.00/M tokens | 6.5× |
Goldman's core argument was simple: China's AI industry is no longer competing on who makes the cheapest token. The battle has shifted to who can command a premium. The report noted that K3's blended rate still undercuts Claude Fable 5 and GPT-5.6 Sol, placing it in a unique middle ground: pricier than any Chinese rival, but cheaper than the American frontier.
The Numbers Behind the Hype
K3's technical specs justify the ambition. With 2.8 trillion parameters, it is the largest open-weight model ever released. It supports a 1-million-token context window with native multimodal input — text, images, and video. Its list pricing ($3.00/M input, $15.00/M output) mirrors Anthropic's Claude Sonnet tier, a deliberate signal that Moonshot AI sees itself competing at the frontier, not the discount rack.
On Artificial Analysis' Intelligence Index, K3 scored 57 — behind only Claude Fable 5 and GPT-5.6 Sol, ahead of Grok 4.5 and GLM-5.2. On Arena.ai's front-end coding leaderboard, it swept six of seven subcategories. UC Berkeley professor Ion Stoica estimated that the gap between Chinese open-source models and global leaders had narrowed from six-to-nine months to just two to three months.
Demand was immediate. Within 48 hours, Moonshot suspended new consumer subscriptions. "Kimi K3 has received far more love than we expected, and our GPUs are feeling it," the company posted.
Why Pricing Matters: From Price War to Value Pricing
For two years, the narrative around Chinese AI was dominated by a single word: cheap. DeepSeek's R1 shocked markets in January 2025 with near-frontier intelligence at a fraction of the cost. The subsequent race to the bottom drove per-token prices as low as $0.18 — great for developers, devastating for business models.
Goldman Sachs frames K3's pricing as evidence that this era is ending, outlining four structural shifts:
Pricing power signals a real business model. Moonshot's annual recurring revenue tells the story: $100 million in March 2026, $200 million in May, $300 million by June — tripling in three months, with API revenue accounting for over 70% of the total.
The market is bifurcating. Goldman expects high-end coding models to scale toward 2–5 trillion parameters, while low-end agent APIs stay at $0.10–$0.20 with razor-thin margins. The value lies at the top.
Application flywheels matter. Tools like Alibaba's Qoder, Tencent's Workbuddy, and Zhipu's ZCode collect real-world coding data that feeds back into training. Companies that close the "app → data → training → upgrade" loop will sustain their lead.
Infrastructure is the safer bet. Goldman's top pick remains cloud and data centers — Alibaba, GDS, VNET, and Kingsoft Cloud. Among model companies, it highlighted MiniMax as a Buy with an HK$860 target, betting its upcoming H3 video model and M3 coding update could restore competitiveness.
💡 The IPO Pipeline
Moonshot is preparing a Hong Kong IPO targeting up to $50 billion, with Goldman Sachs and CICC as sponsors. It has raised over $5.5 billion from investors including Alibaba, Tencent, Sequoia China, and Meituan. When a company faces public-market scrutiny, pricing power stops being vanity and becomes existential. K3's premium pricing is, in part, a dress rehearsal for life as a public company.
The Market Votes
The stock market's reaction was brutal. On July 17, Zhipu (Z.AI, HK:2513) plunged 28.49%. MiniMax (HK:0100) fell 15.63%. By July 20, the two-day toll was staggering: Zhipu had lost 42.47%, crashing through HK$1,000 and HK$900 to close at HK$890.50. MiniMax shed 24.57%, hitting an all-time low of HK$191.
The selloff wasn't purely about K3. Both stocks had been retreating from bubble-era highs for weeks. Zhipu peaked at HK$2,980 in late June — a HK$1.3 trillion valuation — before a month-long correction erased nearly 70%. MiniMax collapsed over 80% from its March high of HK$1,330. Lock-up expirations, discounted share placements, and a broader tech rout all played a role.
But K3 was the catalyst that crystallized an uncomfortable truth: when model leadership can shift with a single release, the "scarcity premium" that inflated valuations evaporates. As BofA Securities put it: "K3 raises the capability ceiling for China AI models, shifting the burden of proof to other independent AI labs."
China Explain: A Maturing Landscape
For readers outside China, three threads are worth pulling:
The distillation backstory. In February 2026, Anthropic accused Moonshot AI — alongside DeepSeek and MiniMax — of running "industrial-scale distillation attacks" against Claude, generating over 3.4 million exchanges through fraudulent accounts. The allegation underscored how Chinese labs had been absorbing capabilities from American frontier models. K3's benchmarks suggest that absorption phase may be giving way to genuine innovation. Its "Kimi Delta Attention" architecture — a hybrid linear attention mechanism — is a homegrown contribution, not a derivative.
The IPO pipeline reshapes incentives. When a company faces public-market scrutiny, pricing power stops being vanity and becomes existential. K3's premium pricing is, in part, a dress rehearsal for life as a public company.
Coding is the new battleground. Goldman expects a dense cluster of model releases in the second half of 2026, with coding as the primary arena. Alibaba has already countered with Qwen3.8-Max-Preview. MiniMax's M3 update is expected imminently. Coding agents consume massive token volumes and generate measurable business value, making them the most rational place for enterprises to pay a premium.
DeepSeek R1 shocks markets
Near-frontier intelligence at a fraction of the cost. Chinese AI companies race to the bottom on pricing.
Anthropic alleges industrial-scale distillation
Moonshot, DeepSeek, and MiniMax accused of running 3.4M+ fraudulent exchanges against Claude.
Moonshot ARR triples to $300M
API revenue surpasses 70% of total. Hong Kong IPO preparation begins.
K3 launches. Goldman publishes "China AI Model Navigation."
Zhipu drops 42.47% in two days. Wall Street stops treating Chinese AI as a curiosity.
What Comes Next
Goldman Sachs' report does not declare a winner. It declares that the rules have changed. The question is no longer "who can make the cheapest model?" but "who can build something worth paying for?"
Kimi K3 has provided an answer. Whether it holds will depend on the next six months of releases — and whether enterprises, not just benchmarks, agree the premium is justified.
For now, one thing is clear: Wall Street has stopped treating Chinese AI as a curiosity. It has started treating it as an investment.