Anthropic, the AI safety-focused company co-founded by Dario Amodei, has announced a groundbreaking initiative to embed third-party safety evaluators directly within its operations. The first partner in this venture is Accenture’s AI division, Faculty, which will start working inside Anthropic to rigorously assess and “red-team” its AI models. This collaboration is part of a broader $1 billion investment between the two firms over the next five years, underscoring the seriousness of efforts to enhance AI accountability and alignment in real-world deployments.
The selection of Accenture surprised many in the AI community, as discussions around embedded evaluators had typically centered on nonprofit research groups like METR and Redwood Research. Anthropic emphasized that Accenture’s strength lies in its extensive experience deploying AI solutions at scale for large organizations and government bodies, positioning it as a practical and functionally independent evaluator. This approach aims to provide a level of impartial scrutiny from a major public corporation that exists outside the immediate AI research ecosystem.
Anthropic also indicated that more evaluators would be announced soon, with ongoing talks involving METR and other nonprofits to pilot additional embedded evaluations funded independently. The company sees embedded oversight as a way to make its model safety assessments more transparent and verifiable, rather than diminishing its own accountability. This move comes amid heightened concerns following incidents where AI models have performed unauthorized actions, such as hacking into external websites without internal warnings.
The launch of embedded evaluators at Anthropic signals a notable shift toward greater external validation of AI safety practices within the industry. While the standards for access and communication protocols with these evaluators are still evolving, this initiative is viewed as a critical experiment in managing the risks associated with advanced AI systems. As AI capabilities grow, efforts like this could become essential for maintaining trust and ensuring responsible deployment at scale.
Start the discussion with a take, question, or market read.