AI Red Teamers Bound by NDAs Raise Safety Disclosure Fears
A growing chorus of AI researchers, former employees, and policy experts is warning that the adversarial testing industry built to catch dangerous flaws in frontier models has a structural conflict of interest: the specialists paid to find the problems answer to the companies profiting from the products. As OpenAI, Anthropic, Google DeepMind, and Meta race to ship increasingly powerful systems throughout 2024 and 2025, the non-disclosure agreements binding their red teams have become a quiet flashpoint in the debate over who really guards AI safety.
The concern is deceptively simple. Red teamers—the people hired to jailbreak models, elicit harmful outputs, and probe for catastrophic failure modes—typically sign contracts that grant the sponsoring lab final say over what gets disclosed. That means the discovery of a serious vulnerability does not automatically translate into public knowledge. The company decides.
How Red Teaming Actually Works Today
Every major lab now treats red teaming as a launch prerequisite. OpenAI convened an external Red Teaming Network ahead of GPT-4 and expanded it for subsequent releases. Anthropic publishes portions of its responsible scaling policy and commissions outside testing. Google DeepMind and Meta run internal and contracted adversarial evaluations before major model drops.
These efforts are genuinely valuable. Red teamers have surfaced prompt-injection exploits, bioweapon-adjacent information hazards, and persuasion capabilities that reshaped how models were deployed. The work is skilled, often unglamorous, and frequently underpaid relative to its importance.
But the reporting structure is the problem. The typical arrangement routes findings back to the very organization with a commercial incentive to ship. When a discovered flaw is embarrassing, legally risky, or commercially inconvenient, the decision to disclose sits with the seller—not an independent party.
The NDA Problem in Plain Terms
Most external red teamers sign non-disclosure agreements before touching a model. These NDAs commonly prohibit discussing findings, methods, or even the existence of certain vulnerabilities without company approval. Some extend for years.
This creates three distinct risks:
- Selective disclosure: Labs can publish the flaws that make them look responsible while burying those that would alarm regulators or users.
- Chilling effect: Testers fear that raising alarms publicly could end their careers or trigger legal exposure.
- No independent record: There is often no third party who can verify what was found versus what was reported.
The issue gained new urgency in 2024 when a group of current and former OpenAI and Google DeepMind staff published an open letter titled "A Right to Warn about Advanced Artificial Intelligence." The signatories argued that confidentiality agreements block employees from raising safety concerns, and called for companies to support a culture of open criticism. The letter did not name red teaming specifically, but its logic applies directly: if the people closest to the risks cannot speak, the public cannot assess the danger.
Separately, OpenAI faced scrutiny in mid-2024 over restrictive offboarding agreements that reportedly tied departing employees' vested equity to non-disparagement clauses. The company said it would not enforce those provisions after the reporting surfaced. The episode illustrated how contract language can quietly convert safety-minded insiders into silent ones.
Why the Incentives Point the Wrong Way
Consider the basic economics. A frontier model represents hundreds of millions—sometimes billions—of dollars in training and infrastructure. Launch timing affects fundraising, partnerships, and competitive positioning against rivals.
When an adversarial finding threatens that timeline, the organization deciding whether to disclose is the one whose valuation depends on shipping. That is not a neutral arbiter. It is an interested party grading its own exam.
Compare this to other high-stakes industries. In cybersecurity, coordinated vulnerability disclosure norms and bug-bounty frameworks give researchers defined pathways to go public if a vendor stalls. In pharmaceuticals, the FDA stands between the drugmaker and the market. In aviation, the NTSB investigates crashes independently of manufacturers.
AI has no equivalent. There is no statutory body that receives red-team findings, no protected channel that lets a tester escalate a buried flaw, and no legal shield comparable to whistleblower protections in finance or public health.
The Regulatory Vacuum
Policymakers have begun circling the problem. The EU AI Act, which entered into force in 2024 with obligations phasing in through 2025 and beyond, requires providers of general-purpose AI models with systemic risk to conduct and document adversarial testing. The U.S. produced the AI Safety Institute under NIST, and the U.K. established its own AI Safety Institute, both of which conduct pre-deployment evaluations of frontier models.
These institutes represent progress toward independent assessment. But they operate largely through voluntary cooperation, and their access still depends on lab consent. The findings are not automatically public, and the institutes lack enforcement teeth comparable to established regulators.
Meanwhile, the 2023 White House voluntary commitments and subsequent executive order encouraged red teaming and information sharing—but voluntary is the operative word. A commitment a company can reinterpret is not the same as an obligation an independent body enforces.
What Would Actually Fix It
Experts pushing for reform point to several concrete mechanisms:
- Protected disclosure channels that let red teamers report unmitigated critical flaws to an independent regulator, overriding the NDA for genuine safety matters.
- Mandatory disclosure timelines modeled on coordinated vulnerability disclosure, so labs cannot indefinitely sit on findings.
- Whistleblower protections extended explicitly to AI safety personnel and external contractors.
- Independent escrow of findings, where a neutral third party holds a record of what was discovered, preventing quiet erasure.
- Transparency reports detailing how many critical findings were identified versus disclosed and mitigated.
Some labs have taken partial steps. Anthropic has published model cards and system evaluations with more detail than most, and both it and OpenAI have engaged with government safety institutes. But engagement is not accountability, and self-reported transparency remains self-reported.
The Deeper Stakes
The uncomfortable reality is that the AI safety industry has, in the words of one critic, built a system where the watchdogs report to the animals they are watching. That framing is provocative, but the structural point holds: independence is the whole value of a safety check, and independence is precisely what current NDA arrangements erode.
This matters beyond insider baseball. As models move into healthcare triage, legal analysis, critical infrastructure, and autonomous agents that take actions in the world, the gap between what red teams find and what the public learns becomes a gap between known risk and disclosed risk. Users, enterprises, and regulators are making trust decisions on incomplete information by design.
None of this implies bad faith by individual red teamers, who are often the most safety-conscious people in the room. The problem is the machinery around them. A skilled tester who finds the monster's teeth still needs somewhere independent to report them.
The Bottom Line
As frontier AI capabilities accelerate into 2025, the credibility of the entire safety enterprise depends on resolving this conflict. The choice facing regulators and labs is whether adversarial testing remains a marketing asset controlled by sellers—or becomes a genuinely independent function with the legal architecture to match. Until red teamers can escalate buried flaws through protected, enforceable channels, the industry's assurances of safety will rest on trust rather than verification. And in a field this consequential, trust without verification is not a safety system. It is a promise.
Support books alumigogo
Your donation helps us keep creating independent content about AI absurdities. Every bit counts!
Secure checkout by Stripe · No account needed
Enjoyed this article? Read more...
More from The Keepers (Human & AI Actors)
Gemini Cyber: The AI Keeper Patching the World
Discover how Gemini Cyber, Google's affordable AI security keeper, races to patch global vulnerabilities—and whether it empowers defenders or arms attackers too.
Gemini Notebook: AI Keepers of Your Personal Knowledge
Discover how Gemini Notebook's rebrand reveals a new alliance between human note-takers and AI keepers. Explore what it means for your personal knowledge.
OpenAI Codex Hardware: Who Really Keeps AI in Check?
Discover how OpenAI Codex hardware reveals the human developers and AI agents keeping code-generating AI in check. The keeper dynamic explained in depth.