On June 27, Brian Armstrong — CEO of Coinbase, the largest cryptocurrency exchange in the U.S. with a market cap approaching $40 billion — posted something on X that should make every American AI company nervous:

"We're experimenting with defaulting to open-weight models like GLM 5.2 and Kimi 2.7 through our own LLM gateway."

Both models are made by Chinese companies. Both are open-source. And together, they helped Coinbase cut its AI spending by nearly 50% — while token usage continued to grow exponentially.

This isn't a tech blog thought experiment. This is a publicly traded American company, with real compliance obligations and real security concerns, voluntarily making Chinese AI models the default choice for all of its engineers. Let's break down what happened, why it matters, and what it signals about the global AI landscape.

Key Takeaways

  • Coinbase made Chinese AI models the default for all 91% of engineers who never hit usage caps
  • AI spending dropped nearly 50% while token consumption kept climbing
  • GLM 5.2 costs 1/6th of Claude Opus for output tokens — $4.40 vs $25 per million
  • ~80% of American AI startups are now using Chinese open-source models
  • Chinese models' token share on OpenRouter surged from <2% to 40%+ in one year

What Coinbase Actually Did

Armstrong's post laid out a surprisingly detailed playbook. Here's the core idea:

91% of Coinbase's engineers never hit their existing usage caps. So instead of lowering limits or adding spending alerts, the company raised the caps — and changed what happens when you open a new chat. The default model is no longer Claude or GPT. It's now GLM 5.2 (from Zhipu AI) or Kimi K2.7 Code (from Moonshot AI), both Chinese open-weight models.

Engineers can still manually switch to frontier models for complex tasks. Code reviews use multiple models simultaneously, cross-checking each other's output. But for the daily grind — code reviews, documentation summaries, routine engineering tasks — the expensive American models are no longer the first choice.

The result? AI spending dropped nearly 50%, even as total token consumption kept climbing.

The Price Gap Is Staggering

Here's what makes the switch so financially compelling (per million tokens):

ModelInput ($/M tokens)Output ($/M tokens)vs Claude Opus
GLM 5.2 (Zhipu, China)$1.40$4.4082% cheaper (output)
Kimi K2.7 Code (Moonshot, China)~$1.50~$4.5082% cheaper (output)
GPT-5.5 (OpenAI)$2.50$10.0060% cheaper (output)
Claude Opus 4.8 (Anthropic)$5.00$25.00

Output pricing tells the real story: GLM 5.2 costs roughly 1/6th of Claude Opus 4.8 for output tokens. When a company is burning through billions of tokens daily across hundreds of engineers, that 6x difference compounds fast.

But cost alone doesn't explain why Coinbase went public with this. The other piece is performance.

How Good Are These Chinese Models, Really?

This is where the "cheap but bad" narrative falls apart.

GLM 5.2 (released June 12, open-sourced under MIT license) has 744 billion parameters and is currently the highest-scoring open-weight model on Artificial Analysis, a widely-cited third-party benchmark platform. Key results:

  • SWE-bench Pro (real-world coding tasks): GLM 5.2 scores 62.1, beating GPT-5.5's 58.6, and approaching Opus 4.8's 69.2
  • FrontierSWE (complex engineering tasks): GLM 5.2 scores 74.4, nearly tied with Opus 4.8's 75.1
  • Design Arena (human-preference coding): GLM 5.2 ranks #1, outperforming Claude Fable 5
  • Code security (Semgrep evaluation): GLM 5.2 produces 7% fewer vulnerabilities than Claude Code in raw, unprompted generation

Kimi K2.7 from Moonshot AI is also no slouch. Moonshot's earlier Kimi K2.5 model was quietly adopted by Cursor — the AI coding tool recently acquired by Elon Musk for $60 billion — as the basis for its Composer 2 self-developed model. Cursor's annual recurring revenue doubled from $100M to $200M+ after the switch, and Moonshot's valuation rocketed from $4.3 billion to $20 billion in six months.

The honest assessment: For routine, high-volume engineering tasks, these Chinese models are "good enough and half the price." For long-chain, multi-step agentic tasks requiring deep reasoning across large codebases, frontier models like Claude Opus still lead. This matches exactly what Coinbase is doing — cheap models for the 90% of tasks that don't need frontier intelligence, expensive models reserved for the hard stuff.

Coinbase Is Not Alone

This isn't a one-off experiment. It's part of a broader pattern:

  • Airbnb switched its customer service AI from GPT to Alibaba's Qwen last year
  • Lindy, an American AI company, migrated from Anthropic Claude to DeepSeek V4 — after its AI spending had exceeded its total employee payroll
  • Snowflake's CEO has publicly stated that GLM 5.2 can match Claude's performance at a fraction of the cost

The data backs this up. According to a March 2026 report from the U.S.-China Economic and Security Review Commission:

Approximately 80% of American AI startups are now using Chinese open-source models.

On OpenRouter, a major AI model aggregation platform, Chinese models' token share has surged from less than 2% a year ago to over 40% as of April 2026. Alibaba's Qwen series has accumulated over 700 million downloads on Hugging Face, surpassing Meta's Llama to become one of the most-downloaded open-source model families globally.

The OpenRouter leaderboard now reads like a roster of Chinese AI: DeepSeek, Xiaomi's MiMo, MiniMax, Tencent Hunyuan, and Zhipu GLM all occupy the top tier.

The Compliance Question: Is This Safe?

This is the obvious concern, and Coinbase addressed it directly.

The company has self-hosted the open-weight models on its own servers. The model weights are downloaded locally. Code, prompts, and data never leave Coinbase's infrastructure — nothing flows to APIs hosted in China. This is the key advantage of open-weight models over API-based services: you get the model's intelligence without giving the original developer access to your data.

For context, Zhipu AI was placed on the U.S. Entity List in January 2025, and Anthropic has publicly accused several Chinese AI companies (including Moonshot and Alibaba) of "distilling" Claude's capabilities through fake accounts. These geopolitical tensions make Coinbase's decision more, not less, significant — the company clearly weighed these risks and concluded that self-hosted open-weight models are both safe and economically rational.

Why This Matters Beyond Coinbase

1. The "AI Cost Crisis" Is Real

Goldman Sachs estimates global token consumption could grow 24x by 2030. If per-token costs don't drop, enterprise AI bills will become unsustainable. Several U.S. companies have already reported AI spending exceeding their total payroll. The current pricing of American frontier models — $10-25 per million output tokens — simply doesn't scale for mass enterprise adoption.

2. Open-Weight Models Are Rewriting the Rules

When Zhipu open-sourced GLM 5.2 under MIT license, it gave every company in the world a frontier-class model they could run, modify, and deploy independently. No vendor lock-in. No API dependency. No monthly bill that grows with usage. This is a fundamentally different economic model from what OpenAI and Anthropic offer.

3. The Pricing Pressure Is Now Existential for Western AI Companies

Anthropic filed a confidential IPO prospectus with the SEC on June 1, with a valuation approaching $1 trillion. That valuation depends on enterprise customers paying premium prices for premium models. If those customers can get 85-90% of the performance for 15% of the cost, the growth story gets a lot harder to tell.

4. U.S. Export Controls Are Hitting a Wall

The U.S. government has restricted Chinese AI companies from accessing American technology. But open-source models can't be restricted — you can ban API access, but you can't ban code that's already on GitHub. The more the U.S. restricts, the more Chinese companies open-source their work, and the more Western enterprises adopt those models.

What This Means for Developers and Enterprises

For developers and enterprise decision-makers, Coinbase's move is a signal worth paying attention to. The calculus has shifted. The question is no longer "Can Chinese models be trusted?" but rather "Can you afford not to consider them?"

The self-hosting model is key here. By downloading the weights and running the models on their own infrastructure, companies like Coinbase eliminate the data sovereignty concerns that have kept many enterprises away from Chinese AI APIs. You get the intelligence of a frontier-class model without sending a single byte of proprietary code or data to a Chinese server.

This is particularly relevant for industries with strict compliance requirements — finance, healthcare, legal — where data residency and audit trails are non-negotiable. Open-weight models offer a path to cutting-edge AI capabilities while maintaining full control over the data pipeline.

The practical implication: if you're still exclusively using American AI APIs without having evaluated Chinese open-weight alternatives, you're likely overpaying for your AI infrastructure by 3-6x. And in a world where AI costs are scaling linearly with usage, that overpayment compounds into a significant competitive disadvantage.

The Bottom Line

Brian Armstrong's post ended with a line that sums up the new reality:

"We don't cap usage. We make usage visible. The more you spend on AI, the more impact we expect from you."

Translation: use whatever you want, as long as it makes business sense. And increasingly, "business sense" means Chinese open-weight models for the heavy-lifting of daily work, with frontier models reserved for the hardest problems.

This isn't about politics. It's not about nationalism. It's about a $40 billion American company looking at its AI bill and making a pragmatic decision. When a model that costs $4.40 per million output tokens performs within striking distance of one that costs $25, the market will eventually choose on price.

Coinbase just showed that the future of enterprise AI might not be made in Silicon Valley. It might be downloaded from GitHub.