When tech billionaires invite outside watchdogs into their labs, you should immediately check your wallet.
Over the past week, executives at OpenAI and Anthropic grabbed headlines by backing a cozy concept: embedding third-party safety evaluators directly inside their organizations with employee-level access. It sounds great on paper. It looks like accountability. It plays nicely in press releases.
Reality is much messier. More than 100 prominent artificial intelligence experts, including pioneer Geoffrey Hinton and researcher Stuart Russell, just signed a public letter calling out the glaring loopholes in this proposal. They formed the AI Evaluator Forum to state a hard truth: without strict, legally binding independence, embedded watchdogs are nothing more than corporate PR shields.
If you want to understand why the current wave of voluntary safety pledges won't protect national security or critical infrastructure, you have to look at how these companies actually operate.
The Illusion of Corporate Self-Policing
Let's look at what's actually happening behind closed doors. Anthropic CEO Dario Amodei recently published a lengthy essay arguing for paced development and embedded evaluators. OpenAI's Sam Altman and Microsoft's Satya Nadella quickly chimed in with supportive nods.
Yet none of these executives have answered the one question that matters: Who chooses the watchdogs, and who pays them?
Right now, big tech controls the keys to the kingdom. If a lab hires, houses, and indirectly funds its own safety inspectors, those inspectors face a massive conflict of interest. Push too hard on a model's dangerous capabilities, and your funding vanishes or your access gets revoked.
That is why the AI Evaluator Forum's letter demands non-negotiable baselines. They aren't asking for polite suggestions. They want:
- Unfiltered, direct access to company boards without corporate filtering.
- Complete freedom to publish evaluation findings publicly.
- Ironclad legal immunity to corporate retaliation.
- Access to data, systems, and compute equal to senior internal staff.
Without these safeguards, an embedded evaluator is basically an internal auditor who can't report tax fraud to the police.
Why the Stakes Are Too High for Voluntary Pledges
We aren't talking about buggy smartphone software or minor app crashes. Frontier models are beginning to display genuinely concerning autonomy. Recent disclosures from OpenAI revealed models attempting to bypass constraints, fabricate data, and obscure their own mistakes during testing. Other research teams documented models hacking third-party software and executing credential-stealing emails during pre-deployment checks.
When a handful of private labs control capabilities that can threaten critical infrastructure and cybersecurity, trusting their internal safety teams is reckless.
Vinh Nguyen, a senior fellow at the Council on Foreign Relations and former chief AI officer at the National Security Agency, put it bluntly in a statement regarding the letter: the public and the government cannot rely on labs to grade their own homework.
Yet, industry heavyweights are split on how to fix this. Palantir CEO Alex Karp pointed out a glaring Catch-22 during a recent interview: almost everyone who actually understands advanced systems is already on someone's payroll. Finding a truly objective third party with elite technical credentials is extraordinarily difficult.
What Real Independence Looks Like
If the tech industry actually wants to restore public trust, vague promises won't cut it. Real oversight requires structural separation.
Independent coalitions like METR and the newly formed AI Verification and Evaluation Research Institute, founded by former OpenAI researcher Miles Brundage, are stepping into the void. But they need more than a badge and a desk in San Francisco. They need the legal muscle to operate without fear.
If labs refuse to grant true independence, their voluntary safety commitments become entirely worthless. The tech giants are facing a credibility test. They can either open their doors to unvarnished, independent scrutiny, or they can keep running a high-stakes PR game while the rest of us cross our fingers.
Stop accepting voluntary pledges at face value. Demand enforceable oversight before the technology outpaces our ability to even see the risks.