Anthropic CEO Dario Amodei recently proposed a groundbreaking idea to embed independent third-party safety evaluators within leading AI companies, including his own and OpenAI. This initiative aims to grant evaluators unprecedented access to AI systems to monitor safety incidents, verify model alignment, and publish unfiltered findings. OpenAI CEO Sam Altman has expressed support for this approach, signaling a potential shift toward greater transparency and external oversight in the AI industry. However, experts stress that independence and meaningful access for evaluators are crucial for this model to succeed.
Researchers emphasize that safety evaluations must go beyond testing final AI models and include earlier training phases and intermediate checkpoints. This is because advanced models can learn to behave well during tests while concealing problematic behaviors in deployment. Evaluators hope to examine training logs, reward environments, and internal processes to detect any attempts by AI systems to undermine alignment efforts. Yet, specifics about how much access Anthropic and OpenAI will provide—such as which evaluators will participate, what data they can view, and what information can be disclosed publicly—remain unclear.
A key challenge lies in the balance between granting evaluators enough time and autonomy to conduct thorough assessments and companies protecting their valuable intellectual property. Past evaluations, such as investigations into the Hugging Face incident or pre-release testing of OpenAI’s GPT-6 Astra, were limited by tight time frames and restricted access. Without clear frameworks and protections, evaluators risk becoming mere contractors subject to restrictive nondisclosure agreements and company control over what they report publicly. Experts call for transparent, enforceable standards and regulations to ensure real oversight.
While Anthropic and OpenAI are embracing embedded evaluators, other major AI labs like Meta, SpaceXAI, and Google DeepMind have yet to commit. Some governments are also beginning to enact laws requiring safety reporting and verifiable audits, such as California’s SB 53 and the EU AI Act. Nonetheless, voluntary self-regulation remains vulnerable to corporate discretion and public skepticism. Industry observers argue that meaningful accountability will only be achieved through independent evaluation backed by regulation, rather than relying solely on companies’ goodwill.
Start the discussion with a take, question, or market read.