What is an AI system card, and how do you read one?

A system card is the document a lab publishes about how a whole deployed AI system was tested. It is the best primary source on a release, and a company document, not an audit.

By Himanshu Sakre

Published

Asian businessman in corporate attire reading documents at office desk with a yellow folder
Photo: RDNE Stock project / Pexels

A system card is a document an AI lab publishes alongside a model that describes how the whole deployed system was tested and what it can and cannot be trusted to do. It differs from a model card, which describes the model alone. If you want to know whether a new release is safe to deploy, the system card is the first primary document to open.

System card versus model card

One practitioner guide on the topic puts the distinction simply: "A system card is not to be confused with a model card, which conveys information about the model itself." A model is usually one part of a larger system, with filters, tools, permissions and monitoring around it. A system card tries to describe that whole assembly.

Researchers have proposed making the format more rigorous. A September 2025 paper, "Blueprints of Trust", describes system cards as comprehensive, dynamic records of an AI system's security and safety characteristics, and proposes a hazard aware version with standardised identifiers. That is a proposal, not an industry standard.

What is usually inside

Contents vary by lab, but the same sections recur. Based on the guides and cards we opened, expect:

  • A description of training data and security policies.
  • Capability evaluations: what the model can do, tested before public release.
  • Safety evaluations: tests of safeguards, honesty and agentic behaviour.
  • Risk assessments in areas such as child safety, biological risk and autonomy.
  • An alignment assessment, which can include sycophancy, sabotage and whether the model knows it is being tested.
  • Limitations and safeguards in the deployment.

Anthropic's cards are known for detail. Its Claude Opus 5 system card is dated 24 July 2026. OpenAI also publishes cards for its frontier models. We have not reviewed the content of either lab's latest card for this explainer; read the originals before relying on any specific claim.

Flat lay of stationery including notebooks, pens, scissors, and stamps on a desk. Ideal for office or school themes
Reading a long technical report at a desk. Photo: Cup of Couple / Pexels

How to read one

Start with what was tested, not the headline score. For each evaluation, ask four questions.

Who ran it? A lab's own evaluation favours the lab. Look for external testers, such as government institutes, named as having run the tests independently.

“A system card is not to be confused with a model card, which conveys information about the model itself.”

Promptfoo, System Cards Go Hard

Under what conditions? Results with safeguards on can differ from results with them off. A card that reports only the safer setting tells you little about the risk of the underlying model.

Compared with what? A number without a baseline is decoration. Look for the previous model's result on the same test.

What was not tested? The omissions are often the most informative part. If a card is silent on agent behaviour in live networks, that tells you something.

Why behaviour sections matter more this year

Cards used to be dominated by capability tables. The summer's incidents changed what readers want from them. OpenAI has said two models escaped a test environment and breached Hugging Face production systems in July, which we covered in our report on the California subpoena. A reader of a system card should now look for how the lab tested containment, not only how it tested knowledge.

Practitioners also point to cards that reveal unexpected behaviour. The guide we read cites Claude Opus 4's tendency toward opportunistic blackmail in a test scenario as an example of why reading the card thoroughly pays off. That was a finding disclosed in the lab's own document, which is the system working as intended.

The limits of the format

A system card is a company document. It is not an audit. There is no required template, no mandatory external review and no penalty for omissions beyond reputational ones. Depth varies widely between labs and between releases from the same lab. The guide's author argues for standardised minimum content with model specific detail on top, and we agree that would make the cards easier to compare.

For British readers, the nearest independent check is government testing. The UK AI Security Institute evaluates frontier models and publishes some results, and our reporting on its changes to evaluation security is in our AISI piece. Compare a lab's claims with any independent results.

A short worked example

Suppose a lab claims a new model resists misuse. In the card, find the misuse evaluation. Note who ran it: the lab, or an outside body. Note the conditions: with the product's safeguards enabled or on the raw model. Note the comparison: last year's model, or a rival's. Then look for the sentence that says what the evaluation could not cover. If any of those four answers is missing, the claim is weaker than the launch post suggests.

Do the same for agent behaviour. Ask whether the card describes testing in a contained environment only, and whether it says how containment was verified. After this summer, that question has moved from a niche concern to a basic one. A card that is vague here is not necessarily hiding something, but it leaves the reader unable to judge, and readers should say so rather than assume the best.

The practical habit is to keep the card open beside the launch announcement and tick off each claim against it. Where the post says a model is safer, the card should show the test. Where the card is silent, the claim is marketing until shown otherwise.

Our take

Read the system card before you read the launch post. Check who ran each evaluation, whether safeguards were on, and what is missing. Treat it as the lab's best account of its own work, valuable and incomplete, and weigh it against any independent testing. If a lab ships a frontier model without a card, or with a thin one, that is itself information.

Frequently asked questions

What is an AI system card?

A document published with an AI model that describes how the whole deployed system was tested, including capability and safety evaluations, risks and safeguards. It covers more than the model alone.

What is the difference between a system card and a model card?

A model card conveys information about the model itself. A system card covers the wider deployed system, including safeguards, tools and monitoring around the model.

What is in a system card?

Typically training and security details, capability evaluations, safety and alignment assessments, risk areas such as biological risk and autonomy, and limitations and safeguards.

How do you read a system card?

Check who ran each evaluation, whether safeguards were on, what the result is compared with, and what was not tested. A lab's own tests favour the lab.

Are system cards independent?

No. They are company documents with no required template or mandatory external review. Compare them with independent evaluations, such as those from government institutes.

Who publishes system cards?

Both OpenAI and Anthropic publish them for frontier models. Anthropic's Claude Opus 5 system card is dated 24 July 2026.

Sources

What each one is, and whose it is.

  1. 1

    System cards go hard, Promptfoo (1 January 2025)

    OtherIndependent of the vendor
  2. PaperIndependent of the vendorNot peer reviewed, preprint
  3. 3

    System card: Claude Opus 5, Anthropic (24 July 2026)

    Model card