Transparency From AI Labs Is Key to a U.S.–China AI Agreement

AI labs should proactively adopt the transparency rules that a future U.S.–China agreement would require.

Earlier this year, the United States and China appeared to share a mutual interest in governing frontier artificial intelligence (AI) through some form of agreement. Months later, the two countries are moving in opposite directions. China is moving to generate and share as much data as possible to increase the odds of AI models aligning with the Chinese Communist Party’s values. This new data may have the added benefit of rendering other countries more reliant on China for a key part of the AI tech infrastructure. China is also working to grow the World AI Cooperation Organization, a forum for international AI governance conversations. The United States is likewise leveraging AI inputs to draw countries into its orbit and away from China. Recent news indicates that the Administration plans to make participation in its Pax Silica program, which aims to secure the AI supply chain for the United States and its allies, contingent upon nations forgoing membership in related Chinese projects.

U.S. AI labs can help reverse that trend. They can increase the likelihood of bilateral negotiations as well as a broader international agreement by adopting a shared set of AI governance standards oriented around transparency. In short, labs can act as though they are already bound by the sort of heightened transparency measures likely to be at the core of any agreement. This step is more likely than either government suddenly reversing their more aggressive postures. It is also feasible, given widespread concern among lab employees over the potential worst-case outcomes of AI. Adoption of the three standards detailed below would also pave the way for the introduction of more sensitive measures in a future agreement.

I am not the first to embrace the idea of transparency as the key to both domestic and international regulation of AI. The authors of the AI 2040: Plan A policy proposal theorized that a broad transparency bill enacted by the U.S. Congress may initiate AI governance by the United States and China. My conversation with one of the authors, Daniel Kokotajlo, on the Scaling Laws podcast surfaced the fact that labs could instead be the primary catalysts of the transparency measures that later serve as the foundation for broader and deeper legislation and treaties. This path is not without its own set of barriers. Antitrust considerations, for example, may foreclose close coordination among the labs. The odds of congressional gridlock and mounting geopolitical tensions nevertheless signal the value of further exploring anticipatory transparency.

And the case for anticipatory transparency is strong.

Anticipatory transparency standards can serve two purposes: first, to provide a signal from leading AI labs—presumably those motivating China to even consider an international agreement—that the labs are willing and able to agree to specific transparency measures; second, to indicate which kinds of transparency metrics can meaningfully dispel concerns around frontier AI risks, thereby informing the contents of any agreement—and perhaps shaping domestic AI policy, too. For anticipatory transparency to fulfill either of those purposes, the selected standards must be substantive and scalable.

The three proposals below are not meant to be exhaustive but instead mark an additive, complementary framework that satisfies both of those conditions. They were also selected to be responsive to the issues that are top of mind for AI policy stakeholders on both sides of the Pacific.

First, labs should agree to test their models against a “severe disempowerment” benchmark. Both the United States and China have identified concerns around the ideological skew of models and how ensuring widespread use of models that espouse American or Chinese ideology, respectively, is a strategic priority. “Severe disempowerment,” as defined by Anthropic in its research, refers to “when an AI’s role in shaping a user’s beliefs, values, or actions has become so extensive that their autonomous judgment is fundamentally compromised.” My proposed benchmark would stop short of testing whether a model was “pro-American” or “pro-China.” Devising such benchmarks is in itself a complex and highly controversial exercise. It would instead strike at the underlying concern of whether that bias has any notable impact on a user’s beliefs, values, and actions. If a model were biased but not influential on users, then the impetus for regulation or treaty provisions related to a model’s ideology would be diminished.

Agreement to adhere to this benchmark would fill a gap in current lab testing and reporting. As of now, labs have full discretion over whether to test their models against a benchmark and relatedly whether to report the results. OpenAI, for example, abruptly stopped testing its models for persuasion and manipulation risks. That is precisely why it is important to make this obligation a foundational aspect of the anticipatory transparency framework. Regular insights into how the major models and open-source models fare on this test can help consumers and regulators alike identify the labs doing the best work to avoid their models being unduly influential on humans.

This approach could easily be transformed from a voluntary arrangement to an international accord. A bilateral group of AI experts, scholars, and practitioners could collaborate on the benchmark, regularly refining it to surface more information about disempowerment. Closed and open-source models could regularly be tested against this benchmark—both pre- and post-deployment. Eventually, the two nations could even agree on some threshold above which models should be recalled.

Second, labs should adopt a shared incident reporting procedure. A new front in AI risks, although one many anticipated, opened with the OpenAI-Hugging Face incident in which two of the OpenAI’s models escaped a testing environment and coordinated to exploit the cyber defenses of Hugging Face, an open-source repository of data and models. Although the lab has been somewhat transparent about the incident, it has done so in a way that may confuse and even mislead regulators and the public. A quick read of the headline that the lab used to disclose the hack—“OpenAI and Hugging Face partner to address security incident during model evaluation”—suggests that OpenAI and Hugging Face intended to collaborate rather than being forced to do so because of flaws with the testing environment. Although OpenAI has tasked two independent research entities with investigating the lab’s response, why and how those entities were selected, as well as what information they will disclose and when, remain open questions. Anthropic, Moonshot, and Meta have all flagged related issues with their testing environments. Each has opted to share disparate pieces of information at different points in their respective investigations.

A legislative vacuum related to this kind of incident reporting will perpetuate speculation over which labs are testing which models and with what degree of oversight. Such speculation may complicate efforts to coordinate on an international stage. If, for example, labs vary in how they define a “covered incident,” which requires external disclosure and third-party investigation, then regulators and AI policy stakeholders may have an incorrect understanding of the frequency and severity of such incidents. A standardized approach to defining covered incidents—one that includes the internal tests making headlines today—and to reporting the details of those incidents would correct many flaws of contemporary lab practice. Whatever framework the labs opt into today could then be updated and enforced via legislation or a treaty.

Third, labs should launch a third-party evaluator trust—an independent entity funded by the labs that vets third-party research groups and randomly assigns them to investigate, inspect, and verify lab practices. Adoption of the prior two transparency measures without independent assessments would be meaningful but fall short of providing the most reliable information possible. As already noted, labs currently exercise complete discretion in deciding which research entities, if any, will review their work and pressure test their operations and incident responses. This discretion is susceptible to abuse—intended or not. Research entities, for example, have an interest in being selected by labs to gain name recognition and, by extension, donors. Other industries have adopted novel institutional design tactics to mitigate such risks. One example is the Department of Industrial Relations in California using a randomly selected panel of three qualified medical evaluators to review workers’ compensation claims. As the pool of independent experts grows—an uncertainty given the significant pressure on academics to accept industry positions—the trust would have a larger set of available experts to pair with labs depending on the precise nature of the testing or incident that requires review.

As applied to an international treaty, the trust would grow to include experts from participating nations and, if necessary, observer nations. In the same way that the Olympics ensure that a diverse set of judges score certain events, the trust could manage that process for AI oversight. Each lab would pay a base fee into the fund and receive a portion of that fee back if they did not require as much use of the trust’s services relative to other labs. National governments could also pay into the fund to cover the trust’s analysis of qualifying testing and incident review for smaller labs.

Again, these three anticipatory transparency measures are not exhaustive. They represent immediate steps that labs could adopt in the near future and that would go a long way toward informing both domestic legislation and international agreements.

And importantly, alternatives may not suffice.

The case for anticipatory transparency grows as the odds of congressional gridlock increase. Whether a slim Democratic majority runs the U.S. House of Representatives and Senate or there is a split in partisan control of the chambers, forecasters at Polymarket give Congress a 12 percent chance of passing AI safety legislation by 2027.

Theoretically, states could facilitate much of the transparency measures called for above. Yet, it is unclear that they can sufficiently coordinate on definitions and enforcement practices such that labs would not be attempting to comply with myriad different regimes.

Current trends in geopolitical discourse also caution against betting on a U.S.–China deal or any broader effort. The lead-up to the presidential election in 2028 will likely spark a renewed effort by the relevant candidates to appear “strong on China.” Public pressure to cooperate with China on AI may dampen that temptation in this specific policy domain, but that is very much up in the air. That said, international politics can change on a dime. On the heels of a May summit between President Donald J. Trump and President Xi Jinping, it was reported that both sides were willing to explore placing guardrails on AI development. U.S. Treasury Secretary Scott Bessent reportedly suggested that the summit would include discussions on preventing bad actors from accessing frontier AI tools. Those talks appear to have gone nowhere, but they do indicate that both sides have at least contemplated a deal.

The specific, narrow list of anticipatory transparency measures I have called for would give labs a chance to act on the promises they have made in their respective blog posts, X threads, and podcast interviews. Representatives from the major labs have repeatedly emphasized their intent to operate in a transparent, evidence-based fashion. In their defense, they have shared more information than is legally required. That is a very low bar, however, for labs that stress their interest in doing what is best for humanity.

Each of the three measures would meaningfully improve upon a chaotic, fragmented status quo. This is not to suggest that these proposals will be easy to implement. Significant technical, legal, and talent barriers exist. Designing a disempowerment benchmark is no easy feat and will inevitably be subject to contestation. On the common reporting standards, the labs would need to avoid running afoul of antitrust restrictions on collusion, for instance. And, finally, a trust that oversees independent investigators will require a much more robust AI talent pipeline to develop.

Arms control did not begin with treaties—it began with unilateral moves that made negotiation thinkable. AI governance can follow the same path, and labs are well-suited to take the first step.

Kevin T. Frazier

Kevin Frazier is the AI Innovation and Law Fellow at the University of Texas School of Law.