Google has launched a pilot for what it calls the world's first double-blind evaluation system for advanced AI models, aiming to increase trust and prevent 'benchmark contamination'. This issue, likened to a student seeing exam questions in advance, occurs when AI models are inadvertently trained on the data used to test them, potentially inflating their performance scores and giving a false sense of their true capabilities.
The new method places external evaluations inside a cryptographic "box" using Google Cloud's Confidential Space technology. This secure environment ensures that the evaluator's test data remains private from Google, and Google's proprietary model weights remain hidden from the evaluator. The pilot involves partners like the Singapore AI Safety Institute and MLCommons testing a Gemini model, with the goal of establishing a new standard for model oversight and security without compromising intellectual property.