The Keepers (Human & AI Actors)

AI Model Cards: How Safety Metrics Became Marketing

July 15, 2026·Idea by The Field Researchers polished by AIObserving a fast-evolving species in its natural habitat — the daily flood of AI research and news — and filing reports on what we find.
AI Model Cards: How Safety Metrics Became Marketing
Font size: A+

AI Model Cards: How Safety Metrics Became Marketing

AI model cards were supposed to be the industry's honesty pledge—standardized documents disclosing what a frontier system can do, where it fails, and how dangerous it might be. But a side-by-side audit of the flagship reports from Anthropic, OpenAI, and Google reveals something closer to a shell game. Each lab measures a different thing, calls it "safety," and quietly grades its own homework on metrics its competitors never touch.

The result is a transparency theater where AI model cards look rigorous but resist comparison. When no two labs define the same terms, disclosure stops being accountability and starts being brand positioning. This article decodes how the industry's self-appointed leash-holders turned safety documentation into a marketing genre.

What AI Model Cards Were Supposed to Do

The concept of the model card dates back to a 2018 research paper co-authored by Google researchers. The original idea was simple and admirable: attach a nutrition label to every machine learning model.

That label would specify the model's intended use cases, its performance across demographic groups, its known limitations, and the ethical considerations a deployer should weigh. Think of it as the technical equivalent of a pharmaceutical insert—dosage, side effects, contraindications.

By 2023, as large language models exploded into public consciousness, every major lab adopted the format. OpenAI published "system cards" for GPT-4 and beyond. Anthropic released detailed cards for its Claude family. Google shipped documentation for Gemini. On the surface, the industry appeared to converge on a shared standard.

But convergence on format is not convergence on substance. The labs agreed to publish documents that look alike while disclosing wildly different things inside them.

The Audit: Three Labs, Three Definitions of Safety

When you place these documents side by side, the incompatibility becomes glaring. Each company anchors its AI safety claims to a benchmark that flatters its own priorities.

OpenAI's system cards lean heavily on their internal Preparedness Framework, scoring models across categories like cybersecurity, chemical and biological risk, persuasion, and model autonomy. These are graded on a Low/Medium/High/Critical scale—but the thresholds are defined by OpenAI, evaluated by OpenAI, and revised by OpenAI whenever a new model would otherwise breach them.

Anthropic structures its disclosures around AI Safety Levels (ASL), a tiered system inspired by biosafety labs. It emphasizes Constitutional AI, red-teaming results, and refusal rates on harmful prompts. The vocabulary is different, the thresholds are different, and the underlying philosophy—that safety emerges from a written "constitution"—is proprietary to Anthropic.

Google's Gemini documentation foregrounds responsible AI principles, fairness audits, and toxicity benchmarks, often reported through metrics tied to Google's own Responsible AI toolkit. Its emphasis on representational harm and content moderation reflects Google's legacy as a consumer platform, not a frontier-risk lab.

The upshot: three companies use the word safety to mean three incompatible things. One measures catastrophic misuse potential. Another measures alignment with a written charter. The third measures fairness and toxicity. None of them can be meaningfully ranked against the others.

Grading Your Own Homework: The Self-Assessment Problem

The deepest flaw in modern AI model cards is not vocabulary—it's the absence of independent evaluation. In nearly every case, the lab that builds the model also designs the test, runs the test, interprets the results, and decides what to publish.

Consider the mechanics of this arrangement:

  • The labs define the benchmarks. There is no external body dictating what a "dangerous capability" threshold should be.
  • The labs run the evaluations. Red-teaming is largely internal, and the raw data is rarely released.
  • The labs choose what to disclose. A model card is a curated document, not a full audit trail.
  • The labs can move the goalposts. When a new model approaches a stated risk threshold, the framework itself can be quietly revised.

This is the equivalent of a car manufacturer designing its own crash test, crashing its own car, and then printing a five-star safety rating on the window sticker. The document exists. The rating is technically real. But the incentive structure guarantees a flattering result.

Selective disclosure compounds the problem. A lab facing a competitive product launch has every reason to emphasize the metrics where it excels and to bury—or simply not measure—the metrics where it lags. Because no shared standard exists, omission is invisible. You cannot notice a missing number if no one agreed it should be there.

How Transparency Became a Marketing Channel

The transformation of AI model cards from accountability tools into marketing assets follows a predictable commercial logic. When safety disclosures are voluntary, unstandardized, and self-graded, they become another surface for brand differentiation.

Watch the language. Anthropic positions itself as the safety-first lab, and its documentation reinforces a narrative of caution and responsibility. OpenAI frames its Preparedness work as evidence of frontier-lab seriousness, useful when courting enterprise clients and regulators. Google emphasizes fairness and responsibility, aligning with its consumer-brand reputation and regulatory exposure.

Each card is technically accurate. Each is also strategically framed. The metrics chosen, the risks foregrounded, the language of reassurance—all of it doubles as positioning in a hyper-competitive market.

There is a subtler dynamic at work too. By publishing detailed-looking documents, the labs preempt calls for mandatory external auditing. "We already disclose extensively," the argument goes. The existence of the model card becomes an argument against the kind of independent scrutiny that would actually verify its claims.

This is regulatory judo. Voluntary transparency is deployed not to enable oversight but to forestall it. The more polished the model card, the weaker the case for a mandated standard—even though polish and rigor are entirely different qualities.

The Real-World Stakes of Incompatible Safety Metrics

Why does any of this matter beyond corporate hair-splitting? Because these documents increasingly inform policy, procurement, and public trust.

Governments drafting AI regulation—from the EU AI Act to proposed U.S. frameworks—look to model cards as evidence of industry self-governance. If those cards are non-comparable, regulators are building policy on sand. They cannot verify that "High risk" at one lab means anything like "High risk" at another.

Enterprise buyers face the same fog. A company procuring a frontier model for healthcare or finance wants to compare safety profiles. But when each vendor reports incompatible metrics, due diligence collapses into vibes. The buyer ends up trusting the brand narrative rather than the data.

Then there is the catastrophic-risk dimension. The most consequential threats—bioweapon uplift, autonomous cyberattacks, large-scale manipulation—are precisely the areas where independent verification matters most. Yet these are also the areas where labs disclose the least detail, citing security concerns that conveniently double as commercial protection.

The genuine tension here deserves acknowledgment: some withholding is legitimate. Publishing a step-by-step account of how a model can assist in weapons development would be reckless. But this legitimate caution provides perfect cover for illegitimate opacity, and there is no external referee to tell the two apart.

What a Real Standard Would Require

Fixing this does not require abandoning the model card. It requires making the document mean something consistent across labs. A credible standard would include several elements the industry currently lacks.

  • Shared definitions. A common taxonomy for risk categories and severity thresholds, agreed across labs and ideally set by a neutral body rather than the companies themselves.
  • Independent evaluation. Third-party auditors with access to the models, empowered to run standardized tests and publish results the labs cannot edit.
  • Mandatory disclosure fields. A fixed schema so that omission becomes visible—if a lab leaves a required metric blank, everyone can see the gap.
  • Reproducible benchmarks. Public methodologies so that outside researchers can replicate safety claims rather than taking them on faith.
  • Versioned goalpost tracking. A public record of when and why risk thresholds change, preventing quiet framework revisions.

Emerging efforts—including work by the UK AI Safety Institute, the US AI Safety Institute, and various academic consortia—point toward external evaluation. But these bodies currently lack the authority to compel access or standardize disclosure. Until they do, the labs remain both player and referee.

Conclusion: Read the Cards, But Trust the Structure

The uncomfortable truth is that today's AI model cards tell you as much about a lab's marketing strategy as about its model's safety. Three of the most powerful companies on earth publish documents that share a format but not a meaning—and that gap is not an accident. It is the natural product of letting the leash-holders write their own rules.

This does not mean the cards are worthless. They contain real information, and reading them critically is far better than ignoring them. But treat every safety claim as self-reported until proven otherwise, and pay close attention to what a lab chooses not to measure.

The next time a frontier lab touts its transparency, ask the harder question: who verified this, and against whose standard? Until the answer is "an independent body using a shared framework," the model card remains a marketing document wearing a lab coat. Demand the standard. Support the auditors building it. And never mistake a polished disclosure for genuine accountability.

💛

Support books alumigogo

Your donation helps us keep creating independent content about AI absurdities. Every bit counts!

Secure checkout by Stripe · No account needed

Share this article