AI Safety & Security
Google HEIR Makes Encrypted Inference More Buildable
Google open-sources HEIR for encrypted inference across four demos; its README still marks the compiler as experimental.
Google has open-sourced HEIR, a compiler toolchain that can translate models for inference over encrypted inputs. The edition’s GPU-kernel correctness analysis makes the parallel explicit: a toolchain milestone is not a production guarantee until its contracts are measured. Google demonstrated HEIR across 4 applications: recommendations, fraud detection, network-intrusion detection, and hotword recognition. The useful operator figure is 50%: the project README lists eight supported backend-and-scheme cells out of a 16-cell matrix—8 ÷ 16—which makes HEIR easier to explore but not yet a universal private-inference layer.
Teams handling sensitive data should put HEIR in a research sandbox, not in the critical path by default. The opportunity is to keep a cloud service from seeing the input while preserving access to a model’s capability. The cost is cryptographic overhead, unfamiliar compiler work, backend qualification, and a still-evolving toolchain. Switch only when the data exposure problem is more expensive than the performance and engineering tax.
Privacy moves below the model
Google’s HEIR announcement frames fully homomorphic encryption as a way for a server to compute on ciphertext and return an encrypted result without seeing the underlying data. That is a different boundary from ordinary transport encryption: the service can process a user’s features, packet data, or audio without receiving those values in plaintext. Google explicitly says the technique has a nontrivial cost overhead, so the privacy gain is not free.
HEIR—the Homomorphic Encryption Intermediate Representation—is Google’s answer to the usability problem. The HEIR project site and getting-started documentation describe it as a compiler toolchain for fully homomorphic encryption, with a goal of compiling high-level programs into encrypted equivalents and supporting hardware accelerators as they mature. The Google HEIR repository makes the project tangible for engineers: it is an MLIR-based compiler effort with multiple backends and schemes, not a hosted API that silently encrypts any model.
The demonstration portfolio is deliberately practical. Google compiled a deep-learning recommendation model, a credit-card fraud detector, the Kitsune anomaly detector for encrypted network traffic, and a hotword detector. The company says the examples present latency numbers on a single-threaded CPU, but the announcement’s accessible text does not provide a comparable plaintext baseline, a production throughput target, or an end-to-end cost. The correct conclusion is that the compiler can express several privacy-sensitive workloads—not that those workloads are ready for general deployment.
The project has a growing ecosystem. Google names 4 accelerator partners—Belfort, Niobium, Cornami, and Optalysys—and says it plans to demonstrate their latency benefits in the future. That future tense matters. An accelerator partnership is evidence of a direction, not evidence that an encrypted inference workload has reached a buyer’s latency or cost target.
The archive’s Claude provenance analysis made a related point at the output boundary: a control is useful only when a team can preserve the signal through a real workflow. HEIR moves the control to the input and computation boundary. The integration question is therefore not “does encryption exist?” but “can the team keep the encrypted path intact across model conversion, serving, observability, debugging, and incident response?”
Half a matrix is not a platform
The README’s support table is a useful reality check. It lists OpenFHE with BGV, BFV, and CKKS; Lattigo with BGV, BFV, and CKKS; tfhe-rs with CGGI; and Jaxite with CGGI. That is eight listed backend-scheme cells. Against the four backends and four scheme columns visible in the matrix—16 possible cells—the direct arithmetic is 8 ÷ 16 = 50%. This is not “50% of private AI” or a quality score. It is a compact measure of how much of the displayed compatibility surface is explicitly populated.
The project documentation supplies more friction than the announcement’s “one-click” aspiration. HEIR’s getting-started guide describes compiler binaries, backend installation, Bazel builds, and workflows that are still being assembled. It labels nightly binaries as intended for testing compiler passes rather than production use. The README also states that HEIR is not an officially supported Google product. Those caveats should sit in every internal architecture review.
The strongest use case is a workload where plaintext access is itself unacceptable: a fraud or recommendation service that needs to compute on sensitive features, or an organization that must collaborate without pooling raw records. The weakest use case is an ordinary low-latency chatbot whose data can already be protected with access controls, enclaves, or local inference. FHE’s extra computation and debugging complexity need a risk reduction large enough to pay for them.
The thesis breaks if encrypted latency remains too high, the supported schemes cannot express a production model, or the operational team cannot diagnose failures without decrypting the evidence it needs. It also breaks if a backend’s security parameters, accuracy drift, or accelerator path are not stable enough for a regulated service. Evidence that would change the verdict is a matched plaintext-versus-encrypted benchmark, documented security parameters, a reproducible model conversion, and a real deployment with an availability and incident process.
The HEIR repository’s research references record academic work built on the project, while Google says four peer-reviewed publications have used HEIR. The two counts should not be silently treated as the same list; they show activity, not production maturity. A serious pilot should pin a commit, record the backend and scheme, measure setup and inference separately, and keep a non-encrypted fallback for degraded service.
- Privacy-sensitive platform teams should prototype one narrow detector first. Fraud or anomaly detection has a bounded input and output, making the encrypted overhead easier to measure than an open-ended assistant.
- Security and infrastructure owners should demand a plaintext control. Report p50 and p95 latency, throughput, accuracy, key-management work, and cost against the same model without FHE.
- Product teams should keep the compiler out of the SLA until the toolchain stabilizes. The 50% README coverage is a map of current support, not a promise of compatibility with the next model.
Google has made encrypted inference more approachable by turning part of the cryptography problem into a compiler problem. That is meaningful progress. It is not permission to forget that compilers, backends, and operational controls still have to earn the right to carry production data.