Dev News Daily ENDE

Google moves federated learning into attested enclaves, with Gboard as the first user

Google has described the next generation of the federated learning system it uses to train models such as Gboard's next-word prediction on data that stays private to users' devices. A post on the Google Research blog on 2 October, by Katharine Daly and Daniel Ramage, says the new system moves the server side into Trusted Execution Environments (TEEs) so that its privacy guarantees can be checked from outside rather than taken on trust.

What was missing before. In Google's earlier federated systems, device updates were uploaded for immediate aggregation, but nobody outside Google could verify that the data was never logged or inspected. Secure Aggregation later protected uploads cryptographically, but the post says it was not compatible with the strongest central differential-privacy guarantees, and the server logic that adds the privacy noise could not be verified by devices or auditors.

How the new system works. Devices encrypt training examples locally and upload them with a pre-authorised access policy: the list of server-side computations allowed to process that data, each of which may release only anonymised results. Those policies must be published to Rekor, a public transparency log. A key management service, itself a cluster of TEEs using the Raft consensus protocol, hands decryption keys only to TEE workloads that match the published policy. A root TEE runs a Python training loop and delegates parallel work to worker TEEs, using Federated Language, an open-source orchestration language derived from TensorFlow Federated. Only metrics and differentially private model weights leave the enclaves, and the encrypted uploads can be processed only for a limited time. The key management and processing binaries can be rebuilt reproducibly from the open-source Confidential Federated Compute repository.

Room for proprietary code. To keep model architectures and preprocessing secret while staying auditable, the processing TEEs can load serialised components at run time, as long as all privacy-relevant logic stays hard-coded in the published Python program. The post points to its whitepaper for side-channel limits of current TEEs.

In production. Gboard already uses the system for English and Japanese next-word prediction models, which Google says have stronger privacy guarantees, better accuracy and much faster compute times than under the previous system.

Google moves federated learning into attested enclaves, with Gboard as the first user
Google moves federated learning into attested enclaves, with Gboard as the first user — Dev News Daily

What it means

The shift is from "trust our server" to "check our server". Remote attestation plus a public log of allowed computations gives auditors something concrete to verify. The guarantee is only as strong as the TEE hardware and the reviewers who actually read the published policies, which is why the reproducible builds matter as much as the enclaves.