Google Says Its Cloud AI Memory Will Be Unreadable Even to Google. Is That Proven?
Google says Private AI Compute can add persistent cloud memory that even Google cannot read. The architecture and audit offer evidence, limits, and two open findings.
Research cutoff: September 29, 2026. This article analyzes Google's announced architecture for Private AI Compute with secure server-side memory, its September 2026 technical brief, the commissioned Trail of Bits assessment, and current reporting. Google describes a capability it plans to bring to Private AI Compute; it has not announced broad product availability.
Google has published credible technical evidence for a cloud AI memory designed to stay unreadable to Google outside protected processing, but that claim is not proven in the absolute sense. The design keeps a key-derived secret on the user's devices, stores each user's memory as encrypted data, and releases plaintext only inside attested confidential-computing environments. An external audit found no reviewed mechanism by which a Google employee acting alone could read user data.
The same audit found ten issues. Eight were marked resolved, while a low-severity deletion rollback risk and a high-severity response-hash exposure remained unresolved in the final report. Google also has not named a general launch date, supported products, recovery rules, or complete user controls. Parts of the service remain closed source, and some verification improvements are still future work.
So the defensible answer is: Google has shown a serious architecture and meaningful audit evidence, not an unconditional guarantee that no Google system, privileged coalition, compromised device, future software change, or implementation error could ever expose the data.
Key Takeaways
- Google announced a planned persistent memory layer for Private AI Compute on September 23, 2026.
- The design stores a separate encrypted memory database for each user while device-derived secrets protect the key needed to open it.
- Memory becomes plaintext temporarily inside protected cloud environments so the service can retrieve context, run inference, and save new context.
- This is confidential cloud processing, not on-device-only processing and not a fully offline AI.
- Persistent memory requires a stable user identifier. Google therefore does not claim the same network-level non-targetability that its stateless Private AI Compute path can provide.
- Trail of Bits reviewed the memory feature and found no reviewed mechanism for a Google employee acting alone to access user data.
- The audit recorded ten findings: eight resolved, one low-severity unresolved deletion rollback issue, and one high-severity unresolved response-hash exposure.
- The assessment was commissioned by Google, covered a point in time, and did not re-audit every closed-source or infrastructure component.
- Google has not said when the feature will be generally available, which products will use it, or exactly how recovery, deletion, exports, account compromise, and enterprise controls will work.
- Users comparing AI memory should separate where history is stored, who holds keys, when plaintext exists, how software is verified, what deletion means, and what happens if a device or account is compromised.
What Is Confirmed
Google Is Designing Persistent AI Memory For The Cloud
Google DeepMind's September 23 announcement describes an extension to Private AI Compute that can retain context across sessions and devices.
The earlier Private AI Compute architecture was designed around stateless requests. A protected cloud environment handled a task and discarded its temporary context when the task ended. That approach is easier to reason about because the service does not need to find and reopen a user's long-term memory later.
Persistent assistance changes the problem. An assistant that remembers a project from a phone, continues it on a laptop, and later refers to it through smart glasses needs durable state. Google proposes keeping that state in the cloud as an encrypted per-user database rather than repeatedly sending an entire history from one device.
The announcement repeatedly uses future-facing language: the new layer "will" be able to function as memory and is "designed" to enable cross-device experiences. Help Net Security likewise describes the examples as potential uses rather than features that are already broadly available.
This is an announced architecture with reviewed code, not evidence that every Gemini account now uses it.
The Device Protects The Key Needed To Open The Memory
Google's updated Private AI Compute Technical Brief describes two key layers:
- a per-user data-encryption key protects the contents of the memory database; and
- a key-encryption key derived from a secret on the user's device protects that data key.
The cloud stores the memory and the wrapped data key. The device supplies the key material needed to unlock the database only through an authenticated encrypted session with approved confidential-computing software.
That design narrows who can turn stored ciphertext into readable memory. A database administrator looking only at the storage layer should see encrypted records, not a readable transcript or profile. A Google employee acting alone should not be able to fetch a database, retrieve a server-held master key, and open it.
This is stronger than ordinary server-side encryption where the service provider controls both the encrypted database and the production keys that decrypt it. It is still not the same as saying the information never enters a Google-operated computer.
Plaintext Still Exists During Authorized Cloud Processing
When the assistant needs memory, the device establishes an encrypted connection to an attested Oak server running inside an AMD SEV-SNP confidential virtual machine. The protected environment receives the user's key-encryption key, unwraps the database key, and decrypts relevant records in volatile memory.
An orchestration environment can then combine those records with the active request and send the protected workload to Google's hardened TPU inference environment. New memory can be returned to the memory server, encrypted again, and committed to persistent storage.
Google says volatile prompt context, tokens, and intermediate activations are erased after the response. The encrypted long-term memory remains because persistence is the feature.
The privacy claim is therefore about where plaintext may exist and which approved software can access it, not about eliminating plaintext from computation. A model cannot use a memory it never receives in usable form.
Attestation Is Supposed To Block Unapproved Server Software
Remote attestation lets a client check cryptographic evidence about the software and protected environment on the other end of a connection. In Google's design, the device should release sensitive data only to a server with an approved identity and configuration.
Google has also published the memory server through Project Oak. Reproducible builds are intended to connect inspectable source code to the binary digest recorded for production. An append-only transparency ledger is intended to preserve a public record of endorsed server software.
These mechanisms address a hard question: even if an enclave is isolated, how does a user know it is running the reviewed program rather than a modified one that logs plaintext?
Attestation is meaningful evidence, but it is not magic. It must cover the right components, validate the right measurements, reject obsolete or malicious configurations, and be checked by the client before keys are released. Our explainer on remote attestation and prompt logging covers why proof about an environment is not automatically proof about every surrounding data path.
Persistent Memory Gives Up Network-Level Non-Targetability
Google's brief is unusually direct about one lost property. A stateless Private AI Compute request can be routed so the inference service cannot associate the request with a particular user. Persistent memory cannot work that way at every layer because the service has to locate the correct user's database.
The stateful system therefore uses a stable per-user identifier and does not claim network-level non-targetability. Google argues that targeting a particular user's store still yields ciphertext unless the user's keys are available inside an approved protected environment.
That is an important tradeoff, not a footnote. The architecture aims to hide content from infrastructure operators while accepting that the surrounding system can identify which encrypted store belongs to which account during operation.
The External Audit Found No Single-Employee Access Path In Reviewed Code
Google commissioned Trail of Bits to threat-model and review the secure server-side memory feature. The public September 21 security assessment says four consultants worked from June 4 through July 17 for a total of five engineer-weeks, with a fix review completed on August 7.
Trail of Bits says it did not find a mechanism in the codebase under review by which Google employees could access user data. It also praised the choice to keep at-rest memory keys on user devices so keys and plaintext memories are available to trusted server hardware only during active use.
That is real supporting evidence for Google's central claim. It is narrower than "Google can never read it": the conclusion applies to the reviewed code and threat model at that point in time.
Eight Of Ten Audit Findings Were Marked Resolved
The assessment recorded two high-severity, one medium-severity, two low-severity, and five informational findings. Google's fixes resolved eight before the final report.
Resolved findings included weaknesses involving memory-blob binding, uncleared confidential-computing state, attestation event-log validation, build transparency, multi-operator deployment controls, audit logging, prompt separation, and a dangling memory view.
The fixes matter because an audit that finds problems and triggers concrete changes is more useful than a ceremonial review that reports only reassurance. But two findings remained open, and they directly limit how broadly the privacy claim should be interpreted.
The Two Unresolved Audit Findings
Deleted memories may be restored into the active database — Low severity, unresolved; risk accepted. A privileged storage operator could restore an older encrypted database, and a later authorized session could make old content active again. The scenario requires persistent privileged storage access and a subsequent user session. The audit notes there is no standard cryptographic solution for provable server-side deletion.
Hashed response chunks leave the trusted computing base — High severity, unresolved. Deterministic truncated hashes derived from model responses are sent to an external recitation service. An attacker with privileged internal access and the right seeds and service access could try to reconstruct response text. Trail of Bits calls exploitation technically difficult. Google was developing an in-enclave Bloom filter to reduce which hashes leave, but the final report still marked the issue unresolved.
Deletion Does Not Yet Have A Cryptographic Freshness Guarantee
Encrypting stored data does not by itself prove that deletion is final. If a privileged storage layer can restore an older encrypted snapshot, the enclave may accept historical records during a later valid session unless the client or server can prove the database is the newest authorized state.
Trail of Bits says Google fixed the ability to substitute one memory blob for another by binding identifiers to ciphertext. The residual rollback issue remained accepted risk while Google investigated client-assisted state verification.
This does not mean a random attacker can undelete a user's memory. It means the architecture did not provide a cryptographic guarantee against a privileged rollback of the whole encrypted state. Product-level deletion language should be judged with that distinction in mind.
Response-Derived Hashes Can Leave The Protected Boundary
Google uses a recitation check intended to stop models from reproducing protected training content. The reviewed system divided responses into text segments, hashed them deterministically, truncated the hashes, and sent them to a service outside the confidential boundary.
Trail of Bits found that an attacker with privileged internal access to the recitation service and knowledge of the hashing setup could precompute likely text, match hashes, and iteratively recover portions of a response. The report rated the issue high severity and high difficulty.
Google planned to move a Bloom filter inside the trusted boundary so most irrelevant hashes would never leave. Trail of Bits said that could substantially reduce exposure, but would not eliminate all information leakage for text already represented in the recitation index.
The careful conclusion is not "Google sends every prompt out in plaintext." The audit explicitly says raw prompts and responses are not directly exposed by this path. It is that deterministic data derived from some responses crossed the boundary and created a technically difficult but meaningful reconstruction risk.
What Is Still Unclear
Google's current public material does not establish:
- when secure server-side memory will become generally available;
- which Gemini, Android, Pixel, Workspace, web, smart-glasses, or other products will use it;
- whether the final product will preserve every key, attestation, and transparency property in the reviewed design;
- how a user will enable, disable, inspect, correct, export, or delete individual memories;
- what retention rules will apply to encrypted databases, backups, snapshots, logs, and account deletion;
- how device replacement, account recovery, multi-device enrollment, lost devices, and revoked devices will affect keys;
- whether Google or a user can recover memory after every authorized device is lost;
- which metadata remains visible outside enclaves beyond the known stable user identifier;
- how enterprise administrators, legal process, abuse prevention, or support workflows interact with the design;
- whether the unresolved hash exposure will be fully remediated before launch;
- whether client-side attestation checking and independently witnessed transparency logs will be complete at launch;
- how future server changes will be reviewed and whether every production binary will remain reproducible from public source; or
- how the design responds to a compromised phone, browser, account, operating system, firmware layer, or authorized client.
These are not reasons to ignore the architecture. They are the questions that turn a design into a product users can evaluate.
Does "Even Google Cannot Read It" Mean End-To-End Encrypted AI?
It depends on what the phrase is meant to cover.
If it means Google infrastructure outside approved confidential-computing environments should not be able to open stored memory, the architecture is designed to support that claim. The device-derived key, encrypted per-user database, enclave boundary, attestation, and transparency record all point in that direction.
If it means no Google-operated hardware ever sees plaintext, the answer is no. Approved protected server workloads must decrypt relevant memory and process it. The security goal is that the plaintext remains inaccessible to human operators and unapproved systems even while the authorized workload uses it.
If it means no compromise anywhere could expose the content, the answer is also no. Device compromise, account takeover, malicious authorized clients, implementation defects, side channels, future configuration changes, multi-party insider actions, or flaws outside the review scope remain distinct risks.
Before trusting any broad label, ask who holds the encryption keys for the AI chat, when those keys are released, which components can see plaintext, and how a user verifies the software receiving it.
Cloud Memory, On-Device Memory, And Browser-Local History Are Different Choices
Privacy comparisons often collapse three architectures into one word: local.
On-device-only memory
The durable memory and processing stay on hardware the user controls. This can minimize cloud exposure, but it may limit model size, cross-device continuity, availability, battery life, and performance. Syncing across devices introduces another data path that must be secured.
Confidential cloud memory
The memory is stored in the cloud as ciphertext and processed inside protected cloud hardware. This can support frontier models and cross-device continuity. It shifts trust into device keys, enclave hardware, attestation, production software, transparency logs, recovery flows, and operational controls.
Browser-local chat history
The user's normal conversation archive stays in browser storage rather than a provider's normal server-side history database. Active prompts still travel to a hosted service for processing, and any files, search requests, voice, image, or video tools have their own processing paths. Browser data can also be lost by clearing site data, changing browsers, or losing the device.
None is universally best. The right choice depends on whether a user values seamless cross-device memory, cloud-scale models, local control, easy recovery, or reducing the amount of durable personal context held by the provider.
Our guide to AI chatbots with no server chat history explains why ongoing history storage and active-request processing must be evaluated separately.
What Should Users Check Before Enabling AI Memory?
Use a concrete checklist rather than a single privacy label:
- Availability: Is the feature actually active for your account, or only announced?
- Scope: Which chats, files, searches, voice interactions, apps, and devices can write to memory?
- Storage: Is memory on-device, browser-local, conventionally server-side, or encrypted for confidential computing?
- Key control: Who can release the key, and what happens during device enrollment and recovery?
- Plaintext boundary: Where does readable data exist during inference, retrieval, indexing, safety checks, and support?
- Identity metadata: Can the system associate the encrypted store, request timing, device, or account with a person?
- Verification: Does the client independently check attestation, or does the provider merely state that a protected workload is running?
- Source-to-binary proof: Can outsiders connect inspectable code to the actual production binary?
- Deletion: Does deletion remove the active record only, or also snapshots, backups, embeddings, logs, and derived data?
- Recovery: Can the provider recover memory without an existing trusted device? Convenience here can change the key-control claim.
- Audit scope: Which components, commits, configurations, and infrastructure were actually reviewed?
- Future changes: How will users know that a production update still matches the reviewed security properties?
Where OpenVeil Fits - And Where It Does Not
OpenVeil offers a different privacy choice for adults who want hosted AI tools without a normal server-side chat-history archive. Normal chat history and custom personas are stored in the user's browser, and OpenVeil does not maintain a normal server-side chat-history record for private chat sessions. OpenVeil also does not use documented prompts, uploaded files, images, audio, selected local-history context, or AI outputs to train foundation models.
That can be a useful fit when the goal is to avoid building a long-lived cloud memory profile in the first place. It does not provide the same seamless cross-device persistent-memory experience Google is designing.
OpenVeil is not fully offline, anonymous, a confidential-computing enclave, a cryptographic audit, or proof that a provider never processes data. Active requests still require processing by OpenVeil and necessary AI, search, upload-processing, media, hosting, routing, security, billing, and infrastructure providers. Browser-local history does not protect data sent during an active request, secure a compromised browser or device, or eliminate separate operational and security records.
If browser-local normal history better matches your tradeoffs, review what browser-local chat history means, read the current privacy policy, and try OpenVeil.
Frequently Asked Questions
Is Google Private AI Compute Memory Available Now?
Google has announced and documented the architecture, but its September 23 post does not provide a broad launch date, complete product list, or account-eligibility details. Treat it as a planned capability unless a specific Google product tells you the feature is active and links its controls and policies.
Can Google Employees Read The Stored AI Memory?
Google's design is intended to prevent that, and Trail of Bits says it found no mechanism in the reviewed code by which Google employees could access user data. That is meaningful evidence, but it is scoped to the reviewed implementation and threat model. It is not proof against every future change, multi-party insider scenario, compromised client, or component outside the review.
Does The AI Memory Stay On My Device?
No. The durable memory is designed to be stored server-side as encrypted data. A secret derived from the user's device protects the key needed to open it. Relevant memory is decrypted temporarily inside approved confidential-computing environments during authorized processing.
Is This Fully Offline AI?
No. The feature is specifically a cloud architecture for using powerful hosted models while trying to preserve privacy properties associated with on-device processing.
Did The Audit Find A Backdoor For Google?
No reviewed source reports a backdoor. Trail of Bits found no single-employee access mechanism in the code it reviewed. It did find ten issues, including two that remained unresolved in the final report.
Can Deleted Memories Come Back?
The audit found a low-severity rollback risk: a privileged storage operator could restore an older encrypted database and a later authorized session could make previously deleted data active again. Google accepted that residual risk while investigating client-assisted state verification. The report does not say ordinary users or outside attackers can casually restore another person's memory.
Do Response Hashes Mean Google Receives Plaintext Answers?
Not through the disclosed recitation path. The audit says truncated deterministic hashes derived from response text leave the trusted boundary, not raw plaintext. It nevertheless rated the issue high severity because a privileged internal attacker with specific access and knowledge could attempt to reconstruct text from those hashes.
Is Confidential Computing The Same As Browser-Local History?
No. Confidential computing protects processing inside isolated cloud hardware and can support encrypted server-side persistence. Browser-local history keeps the user's ongoing normal chat archive in the browser while active requests still use hosted processing. They solve different parts of the privacy problem.
Is An Independent Audit A Guarantee?
No. It is evidence about a defined scope at a point in time. The strength of that evidence depends on auditor access, components reviewed, fixes verified, unresolved findings, production equivalence, and how future changes are controlled.
Does OpenVeil Use Secure Enclaves For Chat History?
OpenVeil's documented privacy model is different. Normal private-chat history is browser-local rather than kept as a normal server-side chat-history record. OpenVeil does not claim that its active processing is fully offline, that all providers use Google's Private AI Compute design, or that browser-local storage replaces enclave security.
Bottom Line
Google's secure server-side memory proposal is more credible than a vague "encrypted AI" promise. It describes device-held secrets, per-user encrypted storage, protected execution, attestation, reproducible code, a transparency record, and an external assessment that found and prompted fixes for real issues.
It is also not the same as keeping memory only on a device. Plaintext must exist inside approved Google-operated confidential environments during use. The memory path needs a stable user identifier. Two audit findings remained open. Significant parts of the wider service remain outside public inspection, and several verification mechanisms are still developing.
The best answer to "Can Google read this AI memory?" is therefore precise: the reviewed design is intended to stop Google employees and ordinary infrastructure from reading stored memory, and the audit found supporting evidence for that boundary. The public record does not justify an absolute claim that Google systems can never expose the data or that every future product implementation will match the reviewed architecture.
For users who want cross-device continuity, that may be a worthwhile and technically sophisticated tradeoff. For users who would rather avoid durable server-side AI memory, browser-local history is a different option with its own limits.
Sources
- Google DeepMind: Advancing Private AI Compute with secure, server-side memory
- Google: Private AI Compute Technical Brief, updated September 2026
- Trail of Bits: Google Private AI Compute Secure Server-Side Memory Security Assessment
- Project Oak: Oak Private Memory
- Help Net Security: Google plans to give Private AI Compute a memory that follows users across devices