Sundar Pichai's Google I/O 2026 keynote opened with a landmark figure: the Gemini app crossed 900 million monthly active users, more than double the 400 million reported a year earlier, while AI Mode in Search separately surpassed one billion monthly active users, according to Pichai's keynote blog post. The session also unveiled two new TPU chip architectures, a restructured AI Ultra pricing plan, and Google's formal declaration of the "agentic Gemini era."
Gemini 3.5 Flash and the pivot from chat to action
Pichai opened with a statement that functioned as a mission update: "It's clear we're firmly in our agentic Gemini era." The phrase reframes what Google is asking consumers to expect from its AI products. Gemini 3.5 Flash, the first model released that day across the Gemini app, Google Search, and the Gemini API, is described by Google as "our first in a series of models combining frontier intelligence with action," a deliberate shift away from question-answering toward task completion, per the official keynote post.
The model surpasses Gemini 3.1 Pro across coding, agentic, and multimodal benchmarks while running at four times the output-token speed of other frontier models, CNBC reported. Gemini 3.5 Pro, the heavier sibling, is in testing and due in June 2026. Two additional releases complete the I/O slate: Gemini Omni, an any-to-any model that accepts and emits image, audio, video, and text; and Gemini Spark, an agentic assistant initially exclusive to AI Ultra subscribers, per 9to5Google.
The user-growth data gives these releases urgency. Daily request volume on Gemini grew sevenfold year-on-year, per Pichai's keynote. AI Mode in Search, one year since launch, hit one billion MAU with queries "more than doubling every quarter." The scale of that distribution is what separates Google's position from every pure-play AI lab. For additional context on how Google builds agent systems, see our coverage of DeepMind's agent infrastructure.
Silicon as strategy: the TPU 8t and 8i split
The most structurally significant announcement at I/O was not a model but a chip architecture decision first disclosed at Google Cloud Next in April and reinforced on stage. Google's eighth-generation TPU family divides into two purpose-built designs: the TPU 8t for training and the TPU 8i for inference. The split is a public signal that Google now regards training compute and inference compute as different economic problems requiring different silicon, per the Google infrastructure blog.
