News · 2026-10-03
Gboard now trains its next-word models on encrypted uploads inside auditable sealed hardware
Google says its Gboard keyboard now trains its English and Japanese next-word prediction models with a new system in which phones upload encrypted training data that can only be decrypted inside sealed server hardware running publicly listed code. Outside auditors can check which programs are allowed to touch the data. Google announced the system on October 2 alongside a technical paper, and says it cut training time for the English model from about two months to three weeks.
Key facts
- Speed: the English model trained in three weeks, against two months for the previous production model, with neutral results on key typing measures.
- When: announced October 2, 2026; paper arXiv:2609.31494 dated September 30.
- Who: Google Research, written up by software engineer Katharine Daly and research director Daniel Ramage.
- Primary source: Google Research's announcement.
Google introduced federated learning in 2017 to improve phone features such as Gboard's next-word suggestions without collecting what people type. In the earlier design, each phone computed a small model update locally and sent only that update, protected by cryptographic secure aggregation and differential privacy. The weakness was verification: users had to trust that Google's servers handled those updates the way Google said.
What changed
In the new design, phones encrypt raw training examples and upload them. They can only be decrypted inside trusted execution environments, or TEEs: regions of a server chip designed so that even the machine's operator cannot inspect what runs inside, and which can produce a cryptographic attestation of exactly what code is running. A key management service, itself running in TEEs, hands out decryption keys only to workloads that match a pre-published access policy. That policy names the exact Python training program allowed to process the data, and the program releases only differentially private model weights and metrics.
Before uploading, each phone checks that the policy has been published to Rekor, a public transparency log, and binds its upload to that policy. A good analogy is a sealed ballot box that can only be opened by a counting machine whose blueprint is posted on the courthouse door, and whose serial number the box checks before it unlocks.
"In earlier FL systems, device data was uploaded for the purpose of immediate aggregation, but there was no way for external observers to verify that the data was never logged or inspected," the Google team writes. The new system, they say, is "a step toward rigorous proof that server side processing preserves individual privacy."
What auditors can check
The Confidential Federated Compute repository documents how an outsider can follow Rekor entries to see which workloads phones will accept, inspect the published training program and its privacy accounting, and instrument an Android device to see which policy it actually verified. Google says the server binaries "can be reproducibly built from open source code."
Moving the work to servers also speeds things up. Because uploads are collected first and processed later, training no longer depends on which phones happen to be idle and charging at a given hour. For a Japanese model, the paper reports collecting 17.8 million uploads in about six days, where the old workflow reached 8.5 million devices over 38 days.
Why it matters
This is the AI-security story of how a large company can make a privacy promise checkable rather than asking for trust. It extends the same sealed-hardware approach Ground Truth covered for Google's private AI inference from answering questions to training models. Our lesson on encrypted inference explains how sealed hardware compares with fully homomorphic encryption.
The caveat
Sealed hardware is not encrypted math. Data is decrypted inside the chip, so privacy rests on the hardware, its attestation chain, and the correctness of the privacy code. The authors name side-channel attacks as an open concern and describe full proofs of the software's correctness as future work. Google's repository describes full source-to-running-binary transparency as an eventual goal, and no independent auditor has publicly rebuilt the exact binary running for Gboard. The server operator can still see which encrypted uploads were picked for each training round, though not their contents. Models so far are small, up to 10 million parameters. The A/B test showed faster training with neutral typing performance, not an accuracy gain. No independent expert response had appeared by October 3.
Key questions
Does Google see what I type on Gboard under the new system?
Can an outsider actually check Google's claims?
What do you still have to trust?
Cite this
APA
Ground Truth. (2026, October 3). Gboard now trains its next-word models on encrypted uploads inside auditable sealed hardware. Ground Truth. https://groundtruth.day/news/gboard-now-trains-on-encrypted-typing-data-inside-auditable-sealed-hardware.html
BibTeX
@misc{groundtruth:gboard-now-trains-on-encrypted-typing-data-inside-auditable-sealed-hardware,
title = {Gboard now trains its next-word models on encrypted uploads inside auditable sealed hardware},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/news/gboard-now-trains-on-encrypted-typing-data-inside-auditable-sealed-hardware.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.