Tensor G6's TPU Upgrade Is Google's Biggest Bet Yet on On-Device AI Agents

Google put most of the Tensor G6's upgrades into the TPU rather than the CPU because a personal AI agent only works if your data can stay on the phone

Arjun Sha profile pictureby Arjun Sha
Link Copied
copy link iconcopy link icon
illustration showing the Tensor G6 chipset inside the Pixel 11 phone schematic

Image Credit: Google

When Google announced the Pixel 11 series earlier this month, most of the stage time went to Gemini AI features like the multi-step task automation, the Rambler voice input and the camera features. But the most important announcement for me was Tensor G6's TPU improvement.

The Tensor G6 chipset packs an in-house TPU with 50% more compute and twice the memory bandwidth, and Google improved it far more than the CPU or GPU this year. I think this was a deliberate choice, and it shows that Google is slowly building the Tensor chipset to power on-device AI agents.

What Runs On Your Pixel Phone?

Not all of the AI features on a Pixel run on the device, but some important and private features run locally. Here are the current set of Pixel AI features powered by on-device AI.

Pixel AI FeaturesWhere it runs
Gemini NanoOn-device
Now Playing, Live CaptionOn-device
Call Screen, Hold for Me, Direct My Call, Call NotesOn-device
Scam Detection, Clear CallingOn-device
Live Translate, Voice TranslateOn-device
Assistant Voice Typing, Smart ReplyOn-device
Pixel Screenshots, Recorder TranscriptionOn-device
Magic Eraser, Best Take, Add MeOn-device
Night Sight, Astrophotography, Super Res Zoom, Pro Res Zoom, Real ToneOn-device
Proactive Assistance (formerly Magic Cue)On-Device

As you can notice, the private and everyday tasks run on the phone using the Gemini Nano model and the TPU. The heavier reasoning, task automation, Video Boost, image generation (Nano Banana), etc. still go to Google's servers.

By the way, this is not new on Google phones. Long before Gemini existed, Pixel devices had run AI on the device for years, starting with Now Playing in 2017 and Live Caption in 2019. The call features, including Scam Detection and Live Translate, work the same way. What is new is that Google now wants to push further with AI by bringing a local-cloud hybrid model.

Rambler is a good example of that approach. The rewritten output is generated by Gemini's cloud models, but when you are offline, the phone falls back to a basic local version. Similarly, Proactive Assistance mostly runs on-device, but for heavier suggestions, it uses cloud models.

Tensor G6's Biggest Change Is the TPU, Not the CPU

Now, look at the numbers Google shared on stage. On the CPU side, Google claims 25% faster web browsing, 15% quicker app opening and up to 20% better efficiency. These are decent gains, as we have noted in our Tensor G6 benchmarks guide.

tensor g6 infographics showing performance improvements in tpu, cpu, isp
Image Credit: Google
tensor g6 infographics showing performance improvements in tpu, cpu, isp
Image Credit: Google

That said, the TPU got a much bigger upgrade. It has 50% more computing power and 2x the memory bandwidth, and Google tuned that bandwidth to power its most advanced Gemini Nano v3 on-device AI model. The company also says some on-device AI tasks now run up to 3.5 times faster and use less power.

Again, Google is not trying to beat the Snapdragon 8 Elite Gen 5 processor or Apple's chips on CPU speed because it knows it would lose. What Google cares about most is inference, which simply means running AI models on the phone quickly and without draining the battery. The TPU is the part of the chip that does this work to provide the best AI experience, and that is where Google is putting its effort.

Google Is Following Its Own Data Centre Strategy

I think there is a clear parallel between Google's TPU servers and the much-smaller TPUs Google is building for its Tensor chipset. Google has been making TPUs for its data centres for over a decade, and its latest server chips show where the company is focusing. The seventh-generation TPU (called Ironwood) was described by Google as its first TPU built for the age of inference.

tensor chipset layout showing tpu, cpu, gpu, isp and other hardware units
Image Credit: Google
tensor chipset layout showing tpu, cpu, gpu, isp and other hardware units
Image Credit: Google

Next, at Cloud Next 2026, Google separated its eighth-generation TPU into two chips: the TPU 8t for training and the TPU 8i for inference. Now, Google described them as chips for the agentic era. Google says the TPU 8i is built to serve large numbers of AI agents while keeping the cost low.

Now, look at the latest Pixel 11 lineup. The Tensor G6 has a TPU designed mainly for running AI tasks and agents on the device and feeding the model efficiently. This is the same approach as the server strategy, just smaller in scale that can fit in your pocket. Google has decided that the future of AI is inference for personal and private agents, and it's building its chips around that idea.

As we move forward, we will see more and more AI and agentic features powered by on-device AI and TPU.

Personal Intelligence Only Works If Your Data Stays Private

You may ask, why does an AI agent need to run on the device at all? The straightforward answer is trust. Google's plan for the next few years is what it calls personal intelligence, which is a part of Gemini Intelligence. This is an agentic experience that can read your messages, your calendar, your photos, your location and your daily habits and then perform actions for you across apps while remembering your preferences.

For this to work well, it needs access to your most private data. And for you to allow that, you need to believe your data is not being collected, kept unencrypted, sent to the cloud or exposed to any third-party services.

personal intelligence setting open on pixel 11 pro fold
Image Credit: Beebom Gadgets
personal intelligence setting open on pixel 11 pro fold
Image Credit: Beebom Gadgets

This is why on-device inference matters so much. If the AI model that reads your personal data runs on the device using the TPU, your data does not need to go to a server at all. The phone can also handle the orchestration, which is the agent basically deciding what to do, which app to open next and what to show you.

Your personal information does not go to the cloud, there are no server logs, and your messages are not used for training by AI companies. This is the main reason on-device AI matters, especially on a device as personal as smartphones. And it's not possible without a capable chip and a dedicated AI accelerator to run the model locally.

Major AI Features Still Run in the Cloud

I am not going to say that Google has solved AI processing with its local TPU. The most capable features still run on the cloud today because a phone TPU can't match a data centre. So Google has also built a private way to use the cloud, and it's called Private AI Compute.

The agent that performs actions across your apps, Gemini Live, Circle to Search, the Magic Editor photo tools, and image generation like Nano Banana and Video Boost all run on Google's servers. These are the heavier tasks that need large models and a lot of compute, and a phone simply can't do them yet.

a user holding pixel 11 pro fold and using the rambler feature
Image Credit: Beebom Gadgets
a user holding pixel 11 pro fold and using the rambler feature
Image Credit: Beebom Gadgets

So, whenever a task is too heavy for the phone but still involves your personal data, it goes to a secure part of Google's cloud that runs on Google's own TPUs and uses something called Titanium Intelligence Enclaves. It uses remote attestation and encryption, and Google says it can't see the data you send there.

Basically, Google is moving ahead with a local-cloud hybrid model. Google runs the tasks locally that fit on the phone and sends the personal tasks it can't handle to a secured part of the cloud.

This Could Follow the Path of End-to-End Encryption

To describe the development around AI agents and on-device AI experiences, I have an analogy to make. About ten years ago, end-to-end encryption went from a niche feature to something people simply expect from a messaging app. You would not use a messaging app today that doesn't have end-to-end encryption.

I think private and personal AI is heading in the same direction. You would not use personal intelligence if the company didn't run your private data locally. In the coming days, it might become the bare minimum.

proactive assistance turned enabled on pixel 11 pro fold
Image Credit: Beebom Gadgets
proactive assistance turned enabled on pixel 11 pro fold
Image Credit: Beebom Gadgets

That said, we are also seeing the emergence of verifiable AI processing via cloud models, as I explained above. Companies are building private cloud infrastructure that can run personal data in a secure part of the cloud that you can actually verify.

And in this effort, Google is not alone. Apple's Private Cloud Compute powers Apple Intelligence and the company allows external researchers to verify that even Apple can't read the data it processes on its cloud models. Meta has brought the same thing to WhatsApp with a feature called Private Processing. You can summarise chats or get writing help without Meta or WhatsApp seeing your messages.

All in all, we will continue to see the rise of on-device AI, as the success of personal intelligence relies heavily on local AI processing, and Google understands that pretty well.

Recommended For You

Popular Mobile List