Google DeepMind Secures Gemini AI Benchmarks with Crypto Walls
Google DeepMind is piloting a new approach to AI benchmark testing built around cryptographic walls that shield confidential test material from both sides of an evaluation. The company says the goal is protecting the integrity of AI model scores at a moment when benchmark results increasingly drive real business and regulatory decisions.

How the Double-Blind Setup Works

Google DeepMind calls this the first double-blind evaluation built for proprietary frontier AI models. The system places both the model and the evaluation prompts behind a cryptographic wall, keeping each side hidden from the other during testing. Google DeepMind said the approach protects confidential benchmarks on one side and proprietary model weights on the other.

The pilot tested Gemini 2.5 Flash Lite in partnership with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.

  • AVERI evaluated the model using reserved prompts from MLCommons’ AILuminate safety benchmark, which covers areas like cyberattacks and hate speech
  • The Singapore AISI tested the model separately with confidential prompts focused on harmful content relevant to Singapore

The system runs on Google Cloud’s Confidential Computing technology, placing the model and evaluation data inside a protected environment where the evaluator cannot access Google’s model weights and Google cannot access the evaluator’s test prompts. The pilot ran on a Google Cloud A3 Confidential VM, using Intel TDX host memory encryption alongside an NVIDIA H100 Confidential GPU.

The Contamination Problem This Targets

Benchmark contamination has become a persistent headache in AI testing. If a model or its developer sees test questions beforehand, scores get inflated in ways that reflect familiarity rather than actual capability. Zero-logging policies have offered some protection, but cryptographic safeguards add a layer that doesn’t rely on trust alone.

The setup targets a specific trade-off evaluators have long faced. They either shared sensitive test material directly with developers, or asked companies to expose proprietary model weights for inspection. Isolating both sides at once means neither party has to give up what it’s protecting.

What Google Didn’t Publish

DeepMind’s announcement detailed the architecture but stopped short of publishing Gemini 2.5 Flash Lite’s actual scores. The technical report also flagged several limitations worth noting.

  • Some proprietary inference code could not be fully inspected during the process
  • Individual Confidential Space builds were not independently reproducible
  • Google services were used to sign attestation reports, which keeps Google inside the verification path rather than fully external to it

MLCommons cautioned that technical secrecy alone isn’t enough. Legal protections and careful benchmark stewardship still matter, and the process needs independent reproducibility and broader transparency to scale across different models and benchmarks.

Hashlytics Take

The cryptography here is real and the engineering is genuinely clever, but it solves a narrower problem than the framing suggests. This proves a model and a benchmark can stay mutually hidden during a test. It does not prove the results are trustworthy, since Google still controls the infrastructure, signs the attestation reports, and chose not to publish the actual scores from its own pilot. Buyers evaluating AI vendor claims should keep asking who supplied the benchmark, who ran the evaluation, what got disclosed afterward, and which parts of the process still required taking the model provider’s word for it. Secure plumbing is not the same as an independent audit.

Follow Hashlytics on Bluesky, Facebook, LinkedIn , Telegram and X to Get Instant Updates