Model Cards: From Hugging Face Template to AI Act Evidence

A model card is a short, structured document that travels with a trained AI model: what the model is for, how it was trained and tested, where it fails and who to call. Google researchers proposed the format in 2018, Hugging Face turned it into the README.md of every model repository, and most teams now meet it as the first page of a vendor’s documentation. Since 2 August 2025 the EU AI Act asks providers of general-purpose AI models for documentation that looks like a model card but is longer and legally demandable, and from 2 December 2027 providers of high-risk systems owe a technical file the card only begins. This guide reads the model card the way an auditor or a market surveillance authority would: what it contains, what it misses, which law can demand it, and how to turn it into evidence.

Model card illustrated by a single index card standing in a wooden card-catalogue drawer with a brass label holder

Key takeaways

  • A model card is a voluntary documentation format with nine canonical sections, defined by Mitchell et al. in 2019. No law requires the format itself.
  • The Hugging Face template is the de facto standard, yet it covers roughly half of what EU AI Act Annex XI and the GPAI Model Documentation Form ask for.
  • Since 2 August 2025 providers of general-purpose AI models must keep model documentation (Annex XI), share it with downstream providers (Annex XII) and publish a training-content summary. A high-risk technical file (Annex IV) follows from 2 December 2027.
  • Transparency is falling: the December 2025 Foundation Model Transparency Index averages 41 out of 100, 17 points below 2024, and the least-filled sections on Hugging Face are environmental impact, limitations and evaluation.
  • A model card becomes evidence when it is versioned, signed by three roles, written in the regulator’s field names and linked to the AI system record.

What is a model card?

The term comes from a paper presented at the FAT\* conference in Atlanta in January 2019, Model Cards for Model Reporting, by a team of Google researchers including Margaret Mitchell and Timnit Gebru. Their definition still holds: a model card is a short document that accompanies a trained machine learning model and reports benchmarked evaluation under a variety of conditions, including across cultural, demographic or phenotypic groups. The paper worked through two examples, a smiling-face classifier trained on the CelebA dataset and the Perspective API toxicity model, and named an audience wider than engineers: policymakers, organisations deciding whether to adopt a model, and the people a model’s decisions affect. Hugging Face gave the format its mass distribution. On the Hub, the model card is the README.md of the model repository: a Markdown document with a YAML metadata block that records the licence, languages, training datasets, base model, task, library and, where the author filled them in, structured evaluation results and CO2 emissions. The Hugging Face documentation treats the card as the unit of discoverability and reproducibility, and its Model Card Guidebook (Ozoani, Gerchick and Mitchell, 2022) makes a point governance teams should keep: filling a card properly takes three roles. The developer writes the training and technical sections, the “sociotechnic” (a lawyer, ethicist or rights advocate) writes bias, risks and out-of-scope uses, and the project organiser owns the model details, the intended uses and the contact point. That division of labour is the first clue that a model card is not a developer artefact alone. It is the place where an organisation states what it knows about a model, in a form a third party can read without the code. For the wider set of documents the Act expects, see our guide to AI system documentation requirements.

Model card, system card, data card: three artefacts, three objects

Three names circulate and they do not describe the same thing. A model card describes the weights: how a model was trained, what it was evaluated on, where it fails. A system card describes a deployed product built around a model: the prompts, the guardrails, the tools it can call, the safety evaluations run on the whole. OpenAI set that convention with the GPT-4 System Card in March 2023, and Anthropic uses the same name for its Claude releases. A telling detail: when OpenAI released its open-weight gpt-oss models in August 2025 it published a model card, not a system card, because what left the building was the weights, not a product. A data card, or datasheet, describes the dataset (Gebru et al., Datasheets for Datasets, 2018). The rule is one artefact per object. When a vendor hands you a single document called a model card for a product that has prompts, retrieval and tools, you are missing at least one card. Our note on what an AI model is draws the line between the model and the system that wraps it.

The nine sections of a model card and the question each one answers

The 2019 paper proposed nine sections. Two things matter for an audit reader: the question each section is supposed to answer, and the way each one usually goes wrong. <table header-row=”true”> <tr> <td>Section</td> <td>Question it answers</td> <td>Typical weakness</td> </tr> <tr> <td>Model Details</td> <td>Who built it, which version, when, what type, under which licence, who to contact</td> <td>Version and date missing, so nobody knows which weights the card describes</td> </tr> <tr> <td>Intended Use</td> <td>Primary uses, primary users, out-of-scope uses</td> <td>Written as marketing; out-of-scope uses left empty</td> </tr> <tr> <td>Factors</td> <td>Which groups, instruments or environments change performance</td> <td>Never listed, so nothing is disaggregated later</td> </tr> <tr> <td>Metrics</td> <td>Which measures, which decision thresholds, how variation is reported</td> <td>Accuracy without thresholds or confidence intervals</td> </tr> <tr> <td>Evaluation Data</td> <td>Which datasets, why they were chosen, how they were preprocessed</td> <td>Evaluation set identical or related to the training set</td> </tr> <tr> <td>Training Data</td> <td>What the model learned from, or why it cannot be disclosed</td> <td>”Proprietary” with no provenance categories</td> </tr> <tr> <td>Quantitative Analyses</td> <td>Unitary results per factor and intersectional results</td> <td>Aggregate score only</td> </tr> <tr> <td>Ethical Considerations</td> <td>Sensitive data, risks to human life, mitigations</td> <td>A disclaimer rather than an analysis</td> </tr> <tr> <td>Caveats and Recommendations</td> <td>What the evaluation did not cover, how to use the model despite it</td> <td>Missing, or a repeat of the licence terms</td> </tr> </table> The Hugging Face template reorganises these nine into Model Details, Uses (direct, downstream and out-of-scope), Bias Risks and Limitations with Recommendations, Training Details, Evaluation, Environmental Impact, Technical Specifications, Citation, Glossary, Authors, Contact and How to Get Started. The guidebook adds three definitions worth copying into your own template: a bias is a skew in performance for some subpopulations, a risk is a socially relevant issue the model might cause, and a limitation is a likely failure mode that the recommendations address. Most cards blur the three into one paragraph of caution. Keeping them separate is what lets an auditor match each limitation to a mitigation and each risk to an owner. The scores themselves deserve the treatment described in our guide to AI benchmarking: a number without the dataset, the date and the threshold is not evidence.

What a model card is not: voluntary, self-reported, uneven

No law prescribes the model card format. Nobody reviews a card before it is published. The author decides what to measure, which groups to disaggregate and which failures to mention. Two measurements show what that produces at scale. The first is Stanford’s 2025 Foundation Model Transparency Index, published in December 2025, the third edition of the series. Thirteen companies were scored against 100 indicators covering the upstream resources, the model itself and its downstream use. The average score was 41 out of 100, 17 points below the 2024 edition. IBM scored 95; xAI and Midjourney scored 14. The companies with the most capable models are, as a group, telling the public less about them than they did a year earlier. The second is a study of the cards themselves. Liang et al. (February 2024) analysed 32,111 model cards on Hugging Face. The training section is the one most consistently filled in. Environmental impact, limitations and evaluation are the least filled. In other words, the sections a buyer or a regulator needs most, what the model cannot do and how well it does what it claims, are the ones authors skip. The same study added detailed cards to 42 popular models that had none or sparse ones and observed a moderate correlation with higher weekly downloads: documentation is rewarded, and still not written. For a governance team the consequence is simple. A model card you receive from a vendor is a claim, not a finding; it tells you what the vendor chose to say on a given date. A model card you publish is a statement you will be held to, by customers under contract and by authorities under the Act. Both readings call for the same discipline: a card is tied to one version, and a model changes after release. Our guide to model drift explains why a card that was true at launch stops being true without anyone editing it.

Where the model card meets the EU AI Act

The EU AI Act never uses the words “model card”. It uses “technical documentation”, “information and documentation for downstream providers” and “instructions for use”, and it lists their minimum content in annexes. Reading those lists against a card shows that the card is where the legal documentation starts, not where it stops.

General-purpose AI models: Article 53, Annex XI and the Model Documentation Form

Article 53 has applied to providers of general-purpose AI models since 2 August 2025. Paragraph 1 imposes four duties: (a) draw up and keep up to date technical documentation with at least the elements of Annex XI, available to the AI Office and national authorities on request; (b) draw up documentation for providers who integrate the model into their systems, with at least the elements of Annex XII, so that they understand its capabilities and limitations; (c) put in place a copyright policy; (d) publish a sufficiently detailed summary of the training content on the template the AI Office published on 24 July 2025. Paragraph 2 exempts models released under a free and open-source licence with public parameters from duties (a) and (b), unless the model carries systemic risk. Our guide to general-purpose AI covers the classification; here the point is the content. Annex XI, section 1, asks for a general description (tasks, acceptable use policies, release date and distribution, architecture and number of parameters, modalities, licence) and for the development elements: the integration means, the design and training methodology with the rationale for key choices, the data (type, provenance, curation, number of data points, measures to detect bias), the compute in floating point operations and training time, and the energy consumption. Section 2 adds, for models with systemic risk, the evaluation strategies and results, the adversarial testing performed, and the system architecture. The GPAI Code of Practice, final text of 10 July 2025, endorsed by the Commission and the AI Board on 1 August 2025, turns those lists into a form. Measure 1.1: document every field of the Model Documentation Form when the model is placed on the market, update it, and keep previous versions for ten years. Measure 1.2: publish a contact point, give downstream providers the fields that concern them within 14 days of a request, and give the AI Office what it asks for. Measure 1.3: control the quality of the documentation and protect it against alteration. The Commission’s page hosts the Code and the form. The Model Documentation Form is, in effect, a model card with a third column. For each field, three boxes say whether the information goes to the AI Office, to national competent authorities, or proactively to downstream providers. The fields are more specific than any public template: the legal name of the provider; a hash or endpoint that proves the model’s authenticity; the release date and the date it was placed on the Union market; the models it was fine-tuned from; the parameter count with two significant figures for the AI Office and as a range for everyone else; maximum input and output sizes; the licence and the acceptable use policy; the intended uses in about 200 words and the AI system types the model may and may not enter; the training process in about 400 words with the rationale for the key decisions; the data by provenance category (web crawling, private third-party datasets, user data, public datasets, synthetic data) with the number of data points; the measures taken to detect unsuitable sources, including child sexual abuse material, and identifiable bias; the training time in hardware days; the compute in floating point operations; the energy in megawatt-hours. Information handed to authorities falls under the trade-secret protection of Article 78.

High-risk AI systems: Annex IV, Article 13 and the deployer’s card

For providers of high-risk systems the documentation duty sits in Article 11 and Annex IV: a technical file drawn up before the system is placed on the market and kept up to date, with nine points. A general description (intended purpose, provider, versions, interactions with hardware and software, instructions for use). The development process: methods, third-party pre-trained models, design specifications, architecture, data requirements with provenance, labelling and cleaning, the assessment of human oversight measures, the pre-determined changes, the validation and testing procedures with their accuracy, resilience and discriminatory-impact metrics, and the cybersecurity measures. The monitoring and control information, including performance for specific persons or groups. Why the chosen metrics are appropriate. The Article 9 risk management system. The lifecycle changes. The harmonised standards applied. The EU declaration of conformity. The Article 72 post-market monitoring plan. Article 18 requires the file to be kept for ten years. Article 13(3) then defines the card the deployer actually receives, the instructions for use: the identity of the provider; the intended purpose; the level of accuracy with its metrics, the resilience of the system and its cybersecurity; the known risks and the reasonably foreseeable misuse; the performance for specific groups; the input data specifications; the pre-determined changes; the human oversight measures; the resources needed, the expected lifetime and the maintenance; and the logging mechanisms. Every one of those items has a place in a model card and almost none of them appears in the average card. The dates moved in 2026. The Digital Omnibus on AI, Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force since 27 July 2026, deferred the stand-alone high-risk obligations of Annex III to 2 December 2027 and the embedded systems of Annex I to 2 August 2028. It also extended to small mid-caps the right, previously reserved to SMEs, to use the Commission’s simplified technical documentation form: the same nine areas, less text, no waived point. Our guides to high-risk classification and to conformity assessment explain who falls under the regime and who checks the file. Deployers are not outside it: Article 26 requires them to use the system according to the instructions for use, and Article 27 asks some of them to run a fundamental rights impact assessment that starts from exactly that card.

Gap audit: the Hugging Face template against the legal field lists

The table below is our reading, not an official mapping. It places each section of the Hugging Face template next to the closest field of Annex IV, Annex XI or the Model Documentation Form and names the gap. <table header-row=”true”> <tr> <td>Template section</td> <td>What the legal lists ask for</td> <td>Gap</td> </tr> <tr> <td>Model Details</td> <td>Legal name, Union market release date, authenticity hash, model dependencies</td> <td>Missing: the template has developers, date, licence and base model, not the legal entity, the EU date or the hash</td> </tr> <tr> <td>Uses</td> <td>Acceptable use policy, intended uses, AI system types the model may and may not enter</td> <td>Partly present: the policy link and the system types are absent</td> </tr> <tr> <td>Bias, Risks and Limitations</td> <td>Measures to detect unsuitable data sources and identifiable bias</td> <td>Missing: the template records outcomes, the form asks for methods</td> </tr> <tr> <td>Training Details</td> <td>Decision rationale, data by provenance category, number of data points, curation, CSAM detection</td> <td>Missing: training data is a one-line field with a dataset link</td> </tr> <tr> <td>Evaluation</td> <td>Evaluation strategies, adversarial testing, discriminatory-impact metrics with thresholds</td> <td>Partly present: results exist, thresholds and disaggregation rarely do</td> </tr> <tr> <td>Environmental Impact</td> <td>Energy in megawatt-hours with two significant figures and the method</td> <td>Partly present: the CO2 field is close, the unit differs</td> </tr> <tr> <td>Technical Specifications</td> <td>Parameter count, maximum input and output size, compute in floating point operations, hardware days</td> <td>Partly present: architecture and hardware exist, the figures rarely do</td> </tr> <tr> <td>No section</td> <td>Annex IV points 2(e) to 2(h) and 3 to 9: human oversight, pre-determined changes, cybersecurity, risk management, lifecycle changes, standards, declaration of conformity, post-market plan</td> <td>Missing</td> </tr> <tr> <td>No section</td> <td>Article 13(3)(e) and (f): expected lifetime, maintenance, logging</td> <td>Missing</td> </tr> </table> Counted this way, a well-filled Hugging Face card covers about half of Annex XI and almost none of the governance points of Annex IV. That is not a defect of the template. It was designed for transparency between practitioners, and the parts the law does not ask for, the citation, the glossary, the authors, the code snippet that gets a model running, are what make a card readable. Keep them. The practical move is to add the legal fields underneath the model card, in the regulator’s names, rather than replace a document people actually open with a form nobody does.

Beyond the EU: where else a model card is expected

  • United States, medical devices. The FDA’s draft guidance of 7 January 2025 on AI-enabled device software functions includes an example model card for users and healthcare providers (Appendix E) and an example 510(k) summary with a completed card (Appendix F). The agency writes that it does not require a model card or a specific format, and that research shows a card can increase user trust. The document was still a draft at the time of writing.
  • California. SB 53, the Transparency in Frontier Artificial Intelligence Act, signed on 29 September 2025 and in force since 1 January 2026, requires developers of frontier models (trained with more than 10\^26 operations) to publish a transparency report before or at deployment: release date, supported languages and modalities, intended uses, general restrictions, contact. Large frontier developers (more than 500 million dollars in annual revenue) add summaries of their catastrophic-risk assessments. Penalties reach 1 million dollars per violation, enforced by the Attorney General (Clifford Chance summary). The skeleton of that report is a model card. The state’s other transparency statute is covered in our guide to the California AI Transparency Act.
  • New York. The RAISE Act, signed on 19 December 2025, amended on 27 March 2026 and in force from 1 January 2027, requires large developers to publish their safety protocols and to report safety incidents within 72 hours, with penalties up to 1 million dollars for a first violation and 3 million for the next (Wiley summary).
  • NIST. The AI RMF Playbook lists model cards and system cards among its suggested documentation practices. It is voluntary, and it is what US customers will ask you to map to; our guide to the NIST AI RMF shows where.
  • European supervisors. The AI auditing checklist commissioned by the European Data Protection Board (Gemma Galdon Clavell, January 2023) opens its audit with a section titled “Model Card”, cross-referenced to GDPR articles on purpose, data categories, impact assessment and automated decisions. The Dutch Ministry of Infrastructure’s AI Impact Assessment, version 2.0 of December 2024, asks whether there is a model card or equivalent documentation. Both treat the card as the first thing an auditor reads.
  • Standards and supply chain. ISO/IEC 42001 lists control areas for AI system technical documentation, system documentation and information for users, and data provenance; a certification auditor asks where that documentation lives and how it is kept current (see our guide to ISO 42001). On the engineering side, CycloneDX 1.5, released in June 2023, added a machine-learning bill of materials whose model component mirrors a model card, so the card can travel inside the same inventory as the software it ships with.

How to write a model card that survives an audit: six steps

  1. One model card per model version, immutable once released. Every retraining, fine-tune or evaluation rerun gets a new card and a change log line that names what changed in data, training, evaluation or intended use. Annex IV point 6 asks for lifecycle changes; Measure 1.1 of the Code keeps previous versions for ten years. Record produced: the version history.
  2. Write the legal fields first, in the regulator’s vocabulary. Legal name of the provider, release date and Union market date, dependencies, licence, acceptable use policy, intended uses and out-of-scope uses, the AI system types the model may and may not enter. These are the fields a downstream provider will ask for under Annex XII within 14 days. Record produced: the general information block.
  3. Evaluate disaggregated and report what Article 13 wants. Accuracy with its metrics and decision thresholds, resilience, cybersecurity, performance for the groups named in the Factors section, the datasets and dates behind each number. Record produced: an evaluation report that matches the card, which is the subject of our AI benchmarking guide.
  4. Document the data by provenance category. Web crawling, private third-party datasets, user data, public datasets, synthetic data; the number of data points with significant figures; the curation steps; the measures that detected unsuitable sources and identifiable bias. Record produced: the data lineage that a data governance framework already asks for.
  5. Three roles sign, one owner keeps it current. The developer, the sociotechnic and the project organiser each sign their sections; a named owner reviews the card at every release and at least once a year; the review date goes on the card. Record produced: the approval trail an AI audit looks for first.
  6. Link the model card to the system record. The inventory entry, the risk register, the instructions for use, the incident log and the vendor contract all point to the card version in force. For purchased models, ask the vendor for the Annex XII fields, write the 14-day duty into the contract and add a quarterly update clause, the way our vendor due diligence checklist suggests. Record produced: the cross-references that turn a document into a file.

Six steps, six records. Together they form the documentation a regulator can already request under Article 53, and the backbone of the Annex IV file that high-risk providers owe from December 2027.

Model cards in the AI system record

A model card is about a model. A regulator asks about a system: the one that screens job applications, prices a policy or triages a patient. The two meet in the registry. Each AI system record in AI Sigil lists its model components, each component carries its model card fields and version, and a change in the card reopens the questions that depend on it: the risk assessment, the instructions for use, the controls that cite a performance figure. The same record holds the vendor’s card next to your own evaluation of it, so the gap between what was claimed and what was measured stays visible. Keeping the card in the inventory rather than in a repository nobody outside engineering opens is what makes the answer to “show me the documentation for this model” a report instead of a search. Our guide to the AI inventory describes what regulators expect to find in that record.

FAQ

What is a model card in simple terms? A model card is a short, structured document that travels with a trained AI model and answers a fixed set of questions: who built it, what it is for, what it must not be used for, what data it learned from, how it was tested, how well it performed for different groups, what its known risks and limitations are, and who to contact. The format was proposed by Google researchers in 2019 and is now the README of every model repository on Hugging Face. It is written for people who did not build the model, including buyers, auditors and regulators. Is a model card mandatory? No law requires the model card format as such. What the law requires is the information. Since 2 August 2025 the EU AI Act obliges providers of general-purpose AI models to keep technical documentation with the Annex XI elements and to give downstream providers the Annex XII elements. From 2 December 2027 providers of high-risk systems owe an Annex IV technical file and Article 13 instructions for use. In California, frontier developers have published transparency reports since 1 January 2026. A model card is the most practical container for all of these, provided the legal fields are added. What is the difference between a model card and a system card? A model card describes the weights: training, evaluation, limitations of the model itself. A system card describes a deployed product built around a model: prompts, guardrails, tools, retrieval, and the safety evaluations run on the whole. OpenAI introduced the system card with GPT-4 in March 2023 and Anthropic uses the same name. If you deploy a model inside an application with its own prompts and tools, you need both documents, and the system card is the one that maps to the instructions for use of Article 13. Does a Hugging Face model card satisfy the EU AI Act? Not on its own. A well-filled Hugging Face card covers roughly half of the Annex XI fields: it lacks the legal name, the Union market release date, the authenticity hash, the acceptable use policy, the data provenance categories, the number of data points, the detection measures for unsuitable sources and bias, the compute in floating point operations and the energy in megawatt-hours. For high-risk systems it covers none of the governance points of Annex IV, such as human oversight, risk management, pre-determined changes and the post-market monitoring plan. It is the right starting document, not the finished file. Who should write the model card? Three roles, following the Hugging Face guidebook. The developer writes the training details, the technical specifications and the evaluation results. The sociotechnic, a lawyer, ethicist or rights advocate, writes the bias, risks and limitations and the out-of-scope uses. The project organiser writes the model details, the intended uses and the contact point and owns the card afterwards. A card signed by one role reads like that role: all technique, or all caution. A named owner with a review date is what an auditor looks for. How often should a model card be updated? At every model version, and at least once a year even when the weights have not changed, because the evaluation data, the intended uses and the legal context do. Each update is a new version of the card with a change log; previous versions are kept, since the GPAI Code of Practice keeps documentation for ten years and Annex IV asks for the history of lifecycle changes. A model card without a version number describes no model in particular.

Conclusion

The model card is the most widely adopted way of describing an AI model to someone who did not build it. It was designed for transparency between practitioners, not for compliance, and the numbers show the limits of goodwill: an average transparency score of 41 out of 100 among the largest developers, and limitation and evaluation sections that most authors leave empty. The EU AI Act has now written a longer version of the card into law for model providers, with field names, word counts and a ten-year retention period, and will do the same for providers of high-risk systems in December 2027. US states ask frontier developers for transparency reports with the same skeleton. The work for a governance team is not to replace the model card but to complete it, version it, sign it and connect it to the system it serves. AI Sigil keeps each model’s card next to its AI system record, its risks and its controls, so the documentation a regulator asks for already exists when the request arrives.

Model Cards: From Hugging Face Template to AI Act Evidence

What a model card contains, what the Hugging Face template misses against EU AI Act Annex IV and XI, and six steps to turn a card into audit evidence.

OECD AI Principles: From Soft Law to an Evidence File

The OECD AI Principles are not law, yet they define what an AI system is under the EU AI Act. What they say, what changed in 2026, and how to evidence them.

GDPR-Compliant AI: The Evidence File Regulators Expect

GDPR-compliant AI in 2026: legal basis after EDPB Opinion 28/2024, web scraping guidelines, Article 4a bias data, DPIA vs FRIA, and the 8 records to keep.

Agentic AI Security: The OWASP Top 10 as Audit Evidence

Agentic AI security read as an auditor would: the ten OWASP agentic risks mapped to EU AI Act duties, the records that prove control, and the deadlines.

NIST AI 600-1: The Generative AI Profile as an Evidence Map

NIST AI 600-1 explained: the 12 generative AI risks, 211 suggested actions, its 2026 status, the Texas TRAIGA defense and the EU AI Act mapping.

Model Drift: When a Compliant Model Stops Complying

Model drift quietly turns a validated AI model into a non-compliant one. Learn the types, detection metrics and what EU AI Act Articles 15 and 72 require.