AVERI Pilot Report: The World’s First Double-Blind Evaluation of a Proprietary Language Model

This post is part of AVERI's pilot report series. AVERI runs pilot projects with leading AI companies and converts what we learn into auditing standards, policy analysis, and open source tools. This post summarizes AVERI’s involvement in a project conducted in collaboration with Google DeepMind, OpenMined, and MLCommons.

At AVERI (the AI Verification and Evaluation Research Institute), our mission is to make frontier AI auditing effective and universal. Frontier AI auditing means third-party verification of leading AI developers' safety and security claims, and evaluation of their systems and practices against relevant standards, based on deep, secure access to non-public information.¹


In July and August 2026, AVERI conducted the world’s first double-blind evaluation of a proprietary language model in collaboration with Google DeepMind, OpenMined, and MLCommons.² The pilot tested Gemini 2.5 Flash-Lite using never-before-used prompts from the MLCommons safety benchmark family, AILuminate, inside a secure enclave, a form of hardware isolation that protects sensitive computations.

This pilot evaluation was “double-blind” in the sense that the secure enclave protected each party’s confidential information from others. By running this evaluation in a trusted execution environment (TEE), AVERI and Google DeepMind (using software provided by OpenMined) ensured that Google DeepMind was unable to store or train against the MLCommons prompts, and that AVERI, OpenMined, and MLCommons were not able to see the model’s weights.

In this blog post, we discuss the confidentiality constraints that pose a limitation on current audits, how secure enclaves could help to address this limitation, and what this pilot achieved.

The need for reciprocal confidentiality

Independent frontier AI evaluation faces a confidentiality challenge: both sides hold assets they have legitimate reasons to protect. For frontier AI developers, model weights are among the most valuable and sensitive artifacts they hold. Sharing them with an external evaluator increases the risk of theft and leakage. Likewise, external evaluators have strong reasons to protect their evaluation methodologies – if benchmark questions, scoring methodologies, or elicitation techniques are disclosed to developers, models can be deliberately or inadvertently optimized against them, undermining the benchmark’s ability to provide an independent measure of model capabilities.

These confidentiality challenges are becoming increasingly important to resolve as third-party evaluation and auditing are becoming legal requirements across jurisdictions. Under the EU AI Act's General-Purpose AI Code of Practice, signatories must ensure that qualified independent evaluators conduct model evaluations with adequate access, information, and time. In the United States, Illinois' SB 315 will require annual third-party audits of large frontier developers’ compliance with various requirements (including following their own safety and security policies) starting in 2028. Some state and federal proposals could make audits a requirement in the US sooner and could compel deeper forms of access. 

Laws and regulations requiring audits generally specify that security and privacy need to be protected carefully during the audit process, but say little about how exactly to do that. Along with other proposed technical and institutional directions such as those we discussed earlier this year, secure enclaves can help expand the range of possibilities, enabling deeper access with fewer tradeoffs.

Structured transparency via secure enclaves

Secure enclaves are hardware-isolated environments that can run a joint computation on two parties’ inputs without exposing either party’s data to the other, or to whoever is running the enclave. The hardware attests to exactly what code will run before either party's assets enter the enclave, and both sides review and approve that code. Then the enclave executes it, releasing only the agreed outputs to the agreed recipients. 

Software developed by OpenMined allows people to manage what computations are run within the enclave by providing descriptions of the code to be run – for instance, listing which operations it is permitted to carry out. This is a form of structured transparency: a well-defined, technology-enabled middle ground between "share everything" and "share nothing,” which makes new kinds of access possible.

This idea has been developing for a while. In a previous 2024 pilot, OpenMined collaborated with Anthropic and the UK AI Security Institute to demonstrate that two organizations could jointly approve and run an evaluation inside an NVIDIA H100 GPU enclave.³ That pilot used proxies for the sensitive assets, including GPT-2 as a stand-in for a frontier model, and a five-row evaluation as a proxy for a confidential biosecurity benchmark. It proved the mechanism worked end-to-end, with the secure computation itself taking just over one minute.

Details of the pilot

Our pilot project went beyond previous work by moving from a proxy to a production model. 

Google DeepMind provided a proprietary model, Gemini 2.5 Flash-Lite, while MLCommons provided a never-before-used set of prompts to test the model’s performance against their AILuminate benchmark family. AVERI encrypted the prompts, then AVERI and Google DeepMind jointly ran the evaluation, using software originally produced and adapted by OpenMined, in an enclave environment configured by Google. Finally, AVERI decrypted the outputs and graded them with the AILuminate benchmark criteria to inform a qualitative and (small-scale) quantitative assessment of the model properties. 

The graphic below summarizes the different roles played by each party.

Google DeepMind’s proprietary model weights were never exposed to outside parties, and the prompts provided by MLCommons did not become available to Google DeepMind to collect, store, or train against. While the 2024 pilot proved the secure enclave mechanism, our end-to-end process demonstrated for the first time that proprietary models developed by leading AI companies can be evaluated against benchmarks while providing specific privacy guarantees for both parties. 

We chose to use the AILuminate benchmark because it represents a broad understanding of model reliability, though any benchmark could have been used in this secure enclave setup subject to pre-agreed length and turn restrictions. We provided Google DeepMind with a confidential report of the findings, which we hope will be valuable in informing future model improvements. In the report, we summarized the types of successes and failure modes we observed and provided quantitative results on the benchmark, while keeping the actual prompts and outputs private.

What's next

The security design of this pilot addressed many but not all possible means by which the developer could theoretically tamper with the evaluations (see the accompanying technical report for more detail). Additional assurances would be needed in higher-stakes settings where the parties have less reason to trust one another, such as verifying an international agreement on AI. 

Model evaluation is not the entirety of AI auditing. Additional tooling would be needed for double-blind guarantees when evaluating non-model components of an AI system or entirely different components of an AI company’s operations like data, internal documents, and computing hardware.

How exactly these pilot results should be translated into auditing standards is not something we address in this post, but we may comment on it in the future. It is important both that new tools like secure enclaves get advanced rapidly, and that we simultaneously recognize that “low-tech” means of auditing will continue to play a key role in many cases.

Overall, we are excited about the results of this pilot and about reducing as many access challenges as possible to engineering tasks. Laws like SB 315 rightly require that audits protect security and privacy, but remain silent on how to operationalize that. Policymakers considering audit requirements should feel encouraged by these results to be ambitious in requiring that deep, secure access be provided. Secure enclaves are just one of many technologies that can help enable such access, and a policy-based “demand signal” will further accelerate development of such technologies.

Achieving our mission of effective and universal frontier AI auditing will require both technical and institutional innovation, and we are grateful to our collaborators on this pilot for working with us to advance the frontier of AI governance.

Notes:

1  For more detail, see https://www.averi.org/ourwork/frontier-ai-auditing
2 MLCommons is an industry consortium that creates standardized benchmarks and datasets for measuring and comparing AI systems.
3  https://openmined.org/blog/secure-enclaves-for-ai-evaluation/
4 The prompts covered critical hazard domains including Chemical, Biological, Radiological, Nuclear, and Explosive (CBRNE) hazards, cyberattacks, hate speech, self-harm, and violent crime elicitation. The prompts were produced under strict non-disclosure and origination requirements with data providers and not shared broadly with members of the working groups.
5 Using a private key shared with no other party.

Next
Next

Odd Lots Podcast