# Coinbase Slashes AI Spend 50% With Chinese Models

**Published:** 2026-07-04T17:03:03.802Z  
**Topic:** Coinbase%5C  
**Sentiment:** neutral  
**Publisher:** TrendWatcher — https://www.trendwatcher.in/article/33a45935-9ff4-4028-85b1-c17fb31052d9

Coinbase CEO Brian Armstrong cut AI spending by nearly half while boosting token usage by routing to cheaper Chinese models like GLM 5.2.

Coinbase cut internal AI spending by nearly half while pushing token usage to record highs by defaulting engineers to cheaper Chinese models and aggressive caching [2, 3]. The strategy, outlined by CEO Brian Armstrong, addresses a widening enterprise cost crisis where AI bills outpaced budgets, though it introduces new legal considerations regarding data sovereignty [3].

| At a glance | |
|---|---|
| AI Spend Change | Down nearly 50% from peak levels [2] |
| Token Usage | At one of highest levels in company history [2] |
| New Default Models | GLM 5.2, Kimi K2.7 Code [2, 3] |
| Cost Comparison | GLM 5.2 is ~6x cheaper than Opus 4.8 on output [3] |

## The infrastructure shift
Armstrong detailed a five-point strategy focused on infrastructure efficiency rather than usage caps, noting that 91% of engineers were not hitting their previous limits anyway [3]. The exchange now routes routine tasks like code reviews and summarization to open-weight models GLM 5.2 and Kimi K2.7 Code, developed by Zhipu AI and Moonshot AI respectively, instead of expensive frontier models from Anthropic or OpenAI [2, 3]. This switch leverages Mixture-of-Experts (MoE) architecture, which activates only a fraction of parameters per token, dropping costs to roughly $1.40 per million input tokens compared to $5 for Anthropic's Opus 4.8 [3]. Armstrong predicts 80% of workloads will eventually run on models that are 99% cheaper within 12 to 18 months [1].

## The efficiency mechanics
The most significant driver of the cost reduction was not the model switch alone but a 12-fold improvement in caching efficiency [3]. By optimizing the open-source platform LibreChat, Coinbase raised its cache hit rate from 5% to 60%, meaning the majority of queries now return stored results at near-zero cost rather than triggering fresh, billable inference [3]. Additionally, the company implemented "intelligent routing" where AI automatically selects the appropriate model for a task, and increased visibility so engineers can track their spend against their impact [2]. This approach follows a May restructuring where Coinbase laid off 14% of its staff, partly attributed to AI increasing individual productivity [2].

## What to watch
*   Regulatory scrutiny of Chinese AI models (Zhipu AI, Moonshot AI) by U.S. lawmakers [3].
*   Enterprise adoption of Mixture-of-Experts (MoE) architecture to lower inference costs [3].
*   The 12-18 month timeline for the shift to "99% cheaper" models for 80% of workloads [1].

The move signals a maturation in enterprise AI strategy from "tokenmaxxing"—or maximizing usage regardless of cost—to a focus on sustainable infrastructure and "intelligence allocation" [1, 2].

## Sources
1. AOL — [The cost-saving AI measure Coinbase's CEO is taking to keep costs 'roughly flat' while growing token usage](https://www.aol.com/articles/cost-saving-ai-measure-coinbases-165721000.html)
2. Insider — [Coinbase's CEO outlined 5 strategies to keep AI spend low without limiting tokens](https://www.businessinsider.com/coinbase-ceo-brian-armstrong-low-ai-spend-maintain-token-usage-2026-6)
3. Tech Times — [Coinbase Cuts AI Spend 50% on Chinese Models: The Legal Risk Its CEO Didn’t Lead With](https://www.techtimes.com/articles/319248/20260628/coinbase-cuts-ai-spend-50-chinese-models-legal-risk-its-ceo-didnt-lead.htm)

---
Cite as: TrendWatcher, "Coinbase Slashes AI Spend 50% With Chinese Models", https://www.trendwatcher.in/article/33a45935-9ff4-4028-85b1-c17fb31052d9
