← News·MarketsMarkets

AI evaluation race hits a conflict-of-interest wall as Washington waits

With binding federal AI legislation still absent, a competition has opened up over which organizations will earn the right to evaluate the safety and capabilities of frontier AI models. Washington is not expected to…

NM
NewsMV Markets Desk
3 min read
18 September 2026Markets desk
Share this dispatch

Key takeaways

  • With no binding federal AI law in place, organizations are competing for the authority to evaluate the safety and capabilities of frontier AI models, while the White House favors industry self-policing.
  • The safety organization METR is at the center of a credibility fight, facing claims it is too close to the companies it assesses and too tied to the effective altruism movement.
  • Booz Allen's AI head Eric Syphard says many AI evaluations focus on performance and lack measures for how often models produce software vulnerabilities or how much outputs vary across similar prompts.
  • Proposed fixes—including US-China labs reviewing each other's models, naming John Carmack, and OpenAI's incident reporting playbook—are all voluntary, which critics say is insufficient.
  • Treasury Secretary Scott Bessent is set to lead talks this weekend with Chinese Vice Premier He Lifeng and says there is an opening for AI safety discussions with Beijing.

With binding federal AI legislation still absent, a competition has opened up over which organizations will earn the right to evaluate the safety and capabilities of frontier AI models. Washington is not expected to move quickly. The White House favors an arrangement where the industry polices itself, and that posture has placed evaluator credibility at the center of a growing fight.

The tension is sharpest around METR, the safety organization that investigated the OpenAI-Hugging Face incident. Some White House officials and AI executives see the existing pool of safety and benchmarking organizations as too close to the companies they would assess and too tied to the effective altruism movement. One of the lead outside investigators in the Hugging Face episode is married to Paul Christiano, a seasoned AI safety and technology official who recently joined the board of OpenAI's non-profit foundation. METR also recently received a hire from Anthropic. The organization has said it does not take funding from frontier labs, and a spokesperson noted that Christiano joined the OpenAI board after the investigation concluded. AI companies and safety officials have countered that the research community is small and that people from frontier labs are well positioned to assess model capabilities.

A White House official offered a pointed read on the stalemate. "The top frontier companies need to come to a consensus on what they want because they haven't agreed on anything," the official said, adding that companies have "every right, reason, and ability to throttle their models" if they see the situation as dire.

Eric Syphard, Booz Allen's head of AI, identified a separate gap. Some facets of existing AI evaluation systems grew out of efforts to market model capabilities, including performance on coding benchmarks and math tests. Syphard said many evaluations still lack components for measuring how often models generate software vulnerabilities or how much their outputs vary across similar prompts. "They all lean in to that performance-only view of the world," he said. Booz Allen is among the defense and tech contractors that already evaluate AI systems for clients.

Proposals to fill the credibility gap have run wide. Elon Musk suggested this week that labs in the United States and China could review each other's models, a prospect some see as unlikely given the competition between the two countries' AI developers. A former Trump adviser floated programmer John Carmack for the role. OpenAI released a public incident reporting playbook after identifying six new incidents, with the framework calling for third-party auditors in more complex cases. That playbook, like the White House framework and Musk's peer review idea, is voluntary. Henry Papadatos, executive director of SaferAI, said voluntary action falls short. "Public transparency is really important here. What these companies are doing is good, but it's only based on their goodwill and I don't think that's sufficient," he said.

What to watch

Treasury Secretary Scott Bessent is set to lead talks this weekend with Chinese Vice Premier He Lifeng and has said there is an opening for AI safety discussions with Beijing. That conversation is the nearest concrete test of whether any cross-border evaluation framework moves beyond the proposal stage.

Categoryworld

Filed via axios.com

Keep reading

More from the markets desk

Frequently asked

Why is METR facing scrutiny as an AI evaluator?

Some White House officials and AI executives view METR as too close to the companies it would assess and too tied to the effective altruism movement, partly because a lead investigator is married to Paul Christiano, who joined OpenAI's non-profit foundation board, and METR recently hired from Anthropic. METR says it does not take funding from frontier labs and that Christiano joined the OpenAI board after the investigation concluded.

What gap in AI evaluations did Booz Allen's Eric Syphard identify?

Syphard said many evaluations grew out of marketing model capabilities and lean on a performance-only view, lacking components to measure how often models generate software vulnerabilities or how much outputs vary across similar prompts.

What solutions have been proposed to address the evaluator credibility gap?

Proposals include Elon Musk's idea for US and Chinese labs to review each other's models, a former Trump adviser floating John Carmack for the role, and OpenAI's public incident reporting playbook calling for third-party auditors in complex cases. All of these, along with the White House framework, are voluntary.

Why do critics say voluntary measures are not enough?

SaferAI executive director Henry Papadatos said public transparency is important but the companies' actions are based only on their goodwill, which he does not consider sufficient.

What upcoming event could test a cross-border evaluation framework?

Treasury Secretary Scott Bessent is set to lead talks this weekend with Chinese Vice Premier He Lifeng, which is described as the nearest concrete test of whether any cross-border AI evaluation framework moves beyond the proposal stage.