What Gets Measured Gets Governed: Benchmarking Privacy in Frontier AI Development and Deployment
As frontier AI systems grow more capable and autonomous, questions about how they handle personal data are no longer theoretical. Standardized benchmarks have long served as one of the primary methods of measuring and comparing frontier AI performance, capabilities, and limitations. However, there is currently a lack of standardized ways to evaluate privacy risks and outcomes across AI systems, despite the challenges these systems raise regarding anonymization, data minimization, data sharing, and other key privacy practices. A set of standardized privacy benchmarks would give developers’ engineering and product teams a concrete score to aim for during AI development, with measurable targets that can be tracked over time, while giving third-party evaluators a consistent, standardized tool for assessment. Deployers, meanwhile, would gain a more reliable method for comparing potential vendors and conducting risk assessments.
This piece explores:
- What benchmarks are and why they matter: How benchmarks are used to measure and compare models, and why they’ve become an important tool for evaluating frontier AI.
- The privacy risks raised by frontier AI: From data memorization and regurgitation to the heightened risks posed by increasingly autonomous AI agents.
- Why we need benchmarks for measuring privacy in AI: A shared, measurable way for developers, deployers, lawmakers, regulators, and researchers to evaluate and compare models on privacy.
- The limitations of benchmarks: Why quantitative benchmarks may not be suited to assess all privacy harms.
- Emerging privacy benchmarking efforts: Standardized, cross-industry benchmarks are currently in development.
1. Background: what are AI benchmarks and why are they important?
Benchmarks are a common method of evaluating AI models and have existed since well before the current wave of generative AI. They traditionally consisted of a fixed set of questions and a key containing the answers to these questions, in order to “grade” a model’s performance. Because they are standardized, benchmarks allow for comparison between models, and tracking progress over time.
Over the past several years, benchmarking has grown more sophisticated: many benchmarks now test agentic tasks involving multiple steps, use model-based grading instead of one fixed correct answer, and simulate realistic scenarios instead of static prompts. A wide range of benchmarks already measure many different aspects of AI model performance. GPQA Diamond, for example, tests frontier models’ general reasoning and knowledge abilities in biology, chemistry, and physics. SWE-bench Verified, meanwhile, measures the ability of coding agents to solve software issues. Benchmarks have even been created to evaluate models’ potential impacts on safety, security, or wellbeing, such as the KORA AI Child Safety benchmark, the AILuminate Benchmark suite of safety and security benchmarks, and the Weapons of Mass Destruction Proxy benchmark, which tests a model’s knowledge in biosecurity, cybersecurity, and chemical security.
Benchmarks are important because they do more than just describe models after the fact. They can also help improve research by allowing for more standardized and rigorous comparisons between models, shaping the direction of research and model development. They also often become de facto governance instruments, as deployers rely on them to evaluate potential vendors or set performance requirements for third parties, or policymakers make policy decisions based on their outputs. At the same time, benchmarks are just one method for evaluating AI. They measure a model’s characteristics in isolation, rather than how it might perform in real-world or adversarial scenarios, something methods like red-teaming are better equipped to test.
2. Frontier AI privacy risks
Frontier AI models are built using large amounts of training data that has been scraped from the internet and other readily available public sources, such as books, academic papers, and other media. This includes personal data, much of which is already publicly available online on the internet, or collected directly from users’ interactions. As a result, personal data is regularly collected for training LLMs in ways that have drawn scrutiny from data protection authorities worldwide with respect to the implications for privacy concepts such as notice, consent, and purpose limitation.
Once personal data is ingested for training a model, there is a risk that the model could memorize and then “regurgitate,” disclose, or reproduce personal data it has ingested for training purposes. Nor is the problem easily reversed after the fact: “machine unlearning,” in which a machine learning model is retrained in order to remove specific “problematic” data, faces significant technical and practical limitations. This makes it challenging to operationalize privacy-preserving mechanisms like data deletion requests and the right to be forgotten, an issue regulators have already confronted when requiring companies to delete models trained on improperly collected data.
As AI companies deploy new agentic capabilities, the risks stem less from what information a system retains, and more from what it can do. Because AI systems can aggregate and analyze formerly disparate data, they may be able to infer sensitive data about individuals, such as health conditions, political views, or religious beliefs, even based on non-sensitive data. Relatedly, models can now re-identify pseudonymized data based on an analysis of unstructured natural language datasets, at a scale and accuracy that traditional de-identification techniques, largely designed for structured datasets, cannot easily mitigate.
Beyond privacy risks from models and agentic systems, the use of AI-generated media can enable a range of harmful behaviors involving appropriation of another person’s likeness, such as nonconsensual deepfakes, impersonation scams, and false or defamatory “hallucinations.” AI may also classify people based on their personal data in order to automate or inform consequential decisions involving employment, healthcare, legal determinations, or other significant outcomes. These decisions may threaten human agency by judging individuals based on actions they have not yet taken, and often reflect or reinforce historical—and potentially biased—data. And AI agents’ increasing autonomy and privileged access to a user’s data and systems, granted so they can complete tasks on a user’s behalf, raises the risk that agents will collect, retain, or share more personal data than a given task requires, undermining the principle of data minimization.
3. Why we need benchmarks for measuring privacy in AI
Benchmarks for measuring privacy in AI models could help bake privacy into the technical development of, and controls placed on, models. Rather than relying solely on legal instruments like DPIAs and contractual promises, which govern how a system is deployed rather than how a model behaves, a benchmark can provide engineers and product teams with a concrete score to aim for and compare across models over time. They also offer more flexibility compared to other forms of evaluation, since their ability to be built or updated relatively quickly helps them keep pace as AI capabilities evolve.
Being able to understand and evaluate AI systems in regards to privacy is particularly important given the increasing adoption of agentic systems, which operate more autonomously and typically have access to more of a user’s data. AI agents are useful precisely because of their privileged access to a user’s internal databases and systems, as well as their ability to interact with external tools on their behalf. At the same time, greater access raises other privacy and data protection considerations, such as over-collection and over-retention of personal data, susceptibility to threats like prompt injection, and the potential for disparate data to be combined in ways that enable sensitive inferences. Addressing these risks isn’t just about harm mitigation. Building more privacy-preserving agents may help build consumer trust and support broader adoption of these tools.
Given AI’s privacy risks, there is a critical need for standardized, industry-vetted benchmarks to measure and compare privacy and data protection outcomes across models. A number of benchmarks already probe specific privacy dimensions, but none has yet emerged as a common, agreed-upon standard that developers, deployers, and regulators can point to. Some existing privacy benchmarks include:
- ConfAIde: Tests contextual privacy leakage.
- PrivacyLens: Extends privacy scenarios into agent trajectories to measure privacy leakage in agents’ actions.
- CI-Bench: Evaluates whether AI assistants protect personal information during inference across roles, information types, and transmission principles.
- AgentDAM: Tests data minimization and privacy leakage in autonomous web agents completing multi-step tasks.
- SAPA-Bench: Evaluates privacy awareness in smartphone agents.
- PrivacyBench: Tests conversational secret-keeping in personalized AI assistants.
Standardized privacy benchmarks would be valuable to a wide range of actors, including:
- Developers: When “privacy” can be measured, it can become a routine, repeatable check built into the model development and release process, rather than a one-time policy review.
- Deployers: When comparing vendors, selecting models to use, or conducting risk assessments, deployers can use a privacy benchmark to inform their decision-making.
- Lawmakers and Regulators: When evaluating claims about model privacy or considering new regulations, standards, or oversight approaches, lawmakers and regulators can use privacy benchmarks as a technical evidence base and shared, measurable reference point, including concrete use cases and incentives for developers.
- Independent researchers and third-party evaluators: When evaluating models and their potential impacts on privacy, independent or public interest researchers and third-party evaluators will have a standardized measure for comparison and a way to detect potential privacy violations.
4. Limitations of benchmarks
Benchmarks are just one method of evaluating AI models, and their results should be understood in light of their limitations. Generally, benchmarks can only capture what can be easily quantified, measured, and compared, and there are many behaviors and impacts that don’t lend themselves well to being benchmarked. In the context of privacy, for example, community or societal-level harms like surveillance may be difficult to capture in a benchmark, which typically measures how a model behaves in a given scenario rather than the aggregate systemic effects of its deployment. More specifically, there are a number of limitations that researchers, policymakers, regulators, and practitioners should keep in mind when analyzing frontier model performance on benchmarks:
- Benchmark datasets may contain errors, and it can be hard to trace their lineage to discover this.
- Benchmarks may not effectively measure what they claim to, either because of bad data or because what’s actually being measured is a poor proxy for the desired target of measurement.
- Many benchmarks are not adequately diverse and representative, with many based on datasets that are disproportionately text-based and in English.
- Benchmarks may favor those industry labs that have the resources to most readily compete in benchmarks.
- Developers may “game” benchmarks by optimizing their models to the benchmark, akin to “teaching to the test,’ which can limit their usefulness or generalizability.
- Benchmarks are bounded by their creators’ knowledge, limiting their ability to assess capabilities or failures that have not been anticipated.
Privacy’s complex and multi-faceted nature further complicates the measurement problem. Privacy harms can be difficult to quantify because they may be intangible, resulting in non-physical or non-financial setbacks such as anxiety or embarrassment; may lead to future downstream consequences that are not immediately visible, such as unauthorized data sharing with a third party; or may be the result of an aggregation of small or dispersed harms that are difficult to quantify on an individual level. While a benchmark cannot measure these harms directly, it can measure specific model behaviors that precede or contribute to them. The aforementioned existing privacy benchmarks for AI models touch on these aspects, such as the likelihood that a model shares sensitive personal information beyond the intended context or the likelihood that an agent leaks personal data while completing realistic tasks. Beyond these measurement challenges, building common, standardized benchmarks is further complicated by the diversity of privacy laws and norms across jurisdictions, each with its own definitions and obligations. A benchmark, while potentially useful for measuring compliance with key common elements across laws, is unlikely to be able to capture the complexity of global regulations.
Despite their limitations, benchmarks are still helpful as long as those shortcomings are recognized. They can also be created or altered relatively quickly in response to issues or feedback, making them flexible in comparison to other forms of evaluation. To ensure benchmarks are useful measures of AI model capabilities, researchers and developers should design them to test reliability, or performance over time and in varying conditions.
5. Emerging Efforts
Benchmarks are one of the key ways to evaluate frontier AI model capabilities, and are critical for pushing AI research forward. While benchmarks for testing AI models on the dimension of privacy exist, they are generally not standardized or vetted across different organizations. A reliable method of evaluating AI models for privacy would help developers, deployers, lawmakers and regulators, and independent researchers and third-party evaluators better understand and compare models’ performance on certain privacy metrics.
A new working group convened by MLCommons—an engineering consortium developing performance benchmarks for AI—is one example of an effort aiming to fill this gap. FPF is pleased to be participating in the Privacy and Confidentiality Working Group alongside leading AI companies and external experts. Part of MLCommons’ AI Risk & Reliability Program, the Working Group is developing standardized benchmarks to evaluate privacy risks and mitigation tools, with current efforts focused on measuring sensitive information disclosure and agentic data minimization. The working group also developed a privacy risk taxonomy to establish shared understandings of privacy risks, with the goal of providing frontier AI developers with reliable ways to measure how their models may result in privacy harms. Efforts like this also reflect a broader shift toward independent, third-party evaluation as a way to assess system risk, rather than relying solely on developer self-assessments. As this evaluation ecosystem matures, standardized privacy benchmarks could become a key part of how frontier AI systems are assessed consistently and credibly across the industry.
Our group of cross-industry, academic, and civil society experts aims to provide AI developers, deployers, regulators, and policymakers with the practical, standardized tools to objectively measure privacy risks at the frontier. Success means moving beyond high-level policy guidelines to equip the engineering community with concrete, measurable targets for designing privacy into the products that serve people worldwide. We are pleased to announce the publication of our Agent Privacy Risk Taxonomy, and look forward to publishing our benchmarks. We welcome industry and civil society collaboration to help shape these technical standards.
– Kristie Chon Flynn and Vinh Nguyen, Co-Chairs of the MLCommons Privacy and Confidentiality Working Group