At Google I/O 2026, Sundar Pichai said that Google is now processing more than 3.2 quadrillion tokens per month across its products and APIs.

That number is large enough to be meaningless without context. So here is the context:

PeriodMonthly tokensYoY growth
May 20249.7 trillion
May 2025~480 trillion~49×
May 20263.2 quadrillion~7×

A quadrillion is 10¹⁵. 3.2 quadrillion is 3,200 trillion. In two years, Google’s token throughput has grown approximately 330× — doubling, on average, faster than every 60 days.

This is not a product announcement. It is a measurement of something that is already happening, at a scale that changes how builders should think about the economics of building on AI infrastructure.


What Quadrillion-Scale Looks Like From the Inside

The headline figure is a platform total — every product that runs on Gemini, including Search, Google Workspace, YouTube summaries, Android AI features, and consumer Gemini apps. It is not exclusively API traffic.

The API-specific numbers are smaller but still striking — at I/O in May 2026:

Update, 2026-08-26: those were already stale two months later. At Alphabet’s Q2 2026 earnings call on July 22, 2026, Google reported Gemini API traffic at roughly 22 billion tokens per minute (up from 16 billion the prior quarter), “nearly 500” Cloud customers each processing more than 1 trillion tokens in the past year, and “more than 9 million” monthly developers — each higher than the I/O snapshot within a single quarter, underscoring this piece’s own point about the growth curve.

The (now ~500-strong) trillion-token-customer cohort is the number worth focusing on. These are organizations — not individuals — running production workloads at a trillion tokens or more per year. That is a new class of AI user that did not exist two years ago. Their workloads are not chatbots. They are pipelines: ingestion, classification, extraction, synthesis, code generation, agent loops running continuously.


Why the Growth Is Not Linear

The 7x year-over-year jump does not come primarily from more users asking more questions. It comes from the architecture of agentic workloads.

A human asking a chatbot a question might consume 500–2,000 tokens per exchange. A coding agent completing a feature branch might consume 500,000–2,000,000 tokens per task — context loading, tool calls, intermediate reasoning, output generation, verification loops. A background agent running for hours on a document corpus might consume more than that before producing a single output.

When developers shift from chat interfaces to agent pipelines, their token consumption does not increase by 10×. It increases by 100–1,000×. Google’s I/O 2026 numbers reflect an ecosystem in the early stages of that transition. The growth is structural, not behavioral.

The phenomenon even has a name now — coined not by an outside critic but by Google’s own CEO. Presenting the 3.2-quadrillion-token figure at I/O, Sundar Pichai told the audience, “Now some out there might call this tokenmaxxing and there’s probably some truth to it," before framing the growth as real developer and enterprise demand rather than an inflated metric. The term caught on in the following weeks, but tech coverage since I/O gives it a more skeptical spin than Pichai’s: it’s used to describe treating raw token volume itself as a stand-in for developer productivity or AI adoption — a leaderboard-style metric several companies have since pulled back from (Amazon reportedly shut down an internal token-consumption leaderboard in late May 2026). Separate from what the coined term critiques, the underlying economic pattern still holds: at current price points, it is often cheaper for developers to spend more tokens — large context windows, chain-of-thought prompting, multi-step reasoning — than to engineer around them.


The Pricing Consequence

High volume at the infrastructure layer creates downward pressure on per-token price. This is not speculation — it is already visible in the pricing announcements at I/O 2026.

Gemini 3.5 Flash launched on May 19, 2026 at $1.50 / $9.00 per million input/output tokens. It outperforms the prior-generation Gemini 3.1 Pro ($2.00 / $12.00) on agentic benchmarks — 83.6% on MCP Atlas versus 3.1 Pro’s lower score — at 25% lower cost.

This is the recurring pattern in commodity infrastructure markets: the dominant provider uses volume to fund price cuts that lock in the next generation of workloads before competitors can match on both capability and economics.

For builders with high-volume agent workloads, the effective cost reductions compound further:

  • Batch API: 50% discount for non-realtime jobs (Gemini 2.5 Flash-Lite falls to $0.05 / $0.20 per million tokens in batch mode)
  • Context caching: up to 90% reduction on cached prompt tokens for large repeated context windows
  • Committed use: enterprise contracts at scale trigger additional negotiated discounts

Google cited that if top companies shifted 80% of their frontier-model workloads to Gemini 3.5 Flash, they would save over $1 billion annually. That figure is an internal calculation, but it reflects the order-of-magnitude cost gap that now exists between frontier-class and mid-tier-class models — and mid-tier-class models in May 2026 outperform frontier-class models from eighteen months ago.


What This Means If You’re Building at Scale

The cost baseline is shifting faster than most roadmaps account for. A pricing assumption baked into a product in January 2026 may already be wrong. The Gemini 3.5 Flash launch cut costs 25% while improving benchmark performance; the direction of this trend is not ambiguous.

The trillion-token customer club is growing. 375 organizations each processed a trillion or more tokens with Google Cloud in the year to May 2026 — already nearly 500 by Google’s next earnings report two months later. This group will expand significantly over the next 12 months. If your application is approaching that tier, you have leverage for enterprise pricing conversations that did not exist two years ago — and you should be thinking about multi-cloud token routing to maintain negotiating position.

Agentic workload economics are different from chat economics. The standard “tokens per user per day” metric that made sense for chat interfaces does not port to agent pipelines. Capacity planning for agent workloads requires estimating task-level token consumption — context window utilization per agent invocation, loop depth, tool call overhead — not per-user averages. The builders running trillion-token organizations are almost certainly doing this per-pipeline, not per-seat.

Google’s scale commitment is a signal about infrastructure stability. Google guided to $180–190 billion in annual capex heading into I/O — then raised that guidance again to $195–205 billion (w.media, BigGo Finance) at its Q2 2026 earnings call on July 22, 2026, with CFO Anat Ashkenazi citing accelerating demand for AI infrastructure capacity. That, plus TPU 8th generation and 3.2 quadrillion tokens of monthly throughput, indicates that Gemini’s infrastructure is not going to experience the capacity constraints that plagued early GPT-4 deployments. For workloads where predictable latency and availability matter more than lowest-possible cost, this matters as much as pricing.


The Uncomfortable Reading

Google’s 3.2 quadrillion tokens figure is also a proxy for competitive moat. The company is not just processing tokens — it is accumulating usage patterns, feedback loops, and implicit preference data at a scale that no competitor processes from a single product surface.

OpenAI has disclosed an API-level run rate — 15 billion tokens per minute by the end of March 2026, up from 6 billion in October 2025, per Wall Street Journal reporting on CFO Sarah Friar’s remarks — but that is roughly 650 trillion tokens/month if sustained, still short of quadrillion scale, and it is an API figure, not a platform-wide total spanning every consumer surface the way Google’s is. Anthropic’s revenue trajectory tells a similar volume story without a token number: $4.73B in Q1 2026 grew to a preliminary $11.5B+ in Q2 2026 — more than double quarter-over-quarter, per Bloomberg reporting — but Anthropic’s token throughput itself remains undisclosed. The competitive picture is not “Google processes the most tokens.” It is “Google is the only company that has disclosed a platform-level figure in the quadrillions, at the moment when that scale begins to produce qualitatively different infrastructure economics.”

Whether that matters depends on what you’re building. For most developers, the practical implication is simpler: the token economy is real, it is large, and the pricing direction is down. Build accordingly.


Coverage note: This builders-log focuses on the economic and infrastructure implications of Google I/O 2026’s token scale announcement. For the technical architecture Google revealed at I/O — Gemini 3.5 Flash, Antigravity 2.0, ADK 2.0, Managed Agents API, Gemini Spark, WebMCP — see our Google I/O 2026 system analysis.