Reading the Model Card and the ONC Source Attributes
A vendor emails you two PDFs the night before the governance committee meets. The first is a four-page "model card" with a training-data section, a performance table, and a paragraph headed Limitations. The second is a screen capture of the tool's "source attributes," a tidy, standardized panel your EHR now surfaces because of a federal rule. Most committees skim both, nod, and move on. But these two documents are not decoration. They are the closest thing you will ever get to a straight answer to the only question that matters before deployment: does this apply to my patients? A clinical leader who can actually read them turns a rubber-stamp into a decision. This lesson teaches you to read them.
Two Artifacts, One Question
Transparency in clinical AI now arrives in two distinct forms, and it helps to keep them straight. The model card is the vendor's own documentation: a structured summary of what the model is, what data it was trained and tested on, how it performed, what it is intended for, and where it should not be used. It is a voluntary industry convention that predates regulation, borrowed from the broader machine-learning world, and its quality varies enormously because no law dictates its contents. A great model card is candid and specific. A weak one is a marketing brochure wearing a lab coat.
The ONC source attributes are different in kind, because they are required. Under the ONC HTI-1 rule, certified health IT had to meet new Decision Support Intervention criteria by December 31, 2024, with ongoing maintenance obligations from January 1, 2025. The rule renamed clinical decision support (CDS) to decision support intervention (DSI), and it created a specific category, predictive DSI (PDSI), defined in regulation as technology that supports decision-making based on algorithms or models that derive relationships from training data and then produce an output resulting in a prediction, classification, recommendation, evaluation, or analysis. For every predictive DSI a developer supplies in its certified product, the system must expose a standardized set of source attributes: a defined list of facts about the intervention's development, validation, intended use, fairness, and applicability. Think of it as a nutrition label. It does not tell you the food is good for you; it tells you, in a consistent format you can learn to read, exactly what is in it so you can decide.
Both documents exist to answer one question, from two directions. The model card is the vendor telling you about the tool. The source attributes are a regulated, standardized minimum the certified system must surface so you are not dependent on the vendor's generosity. Together, read properly, they let a clinical leader judge applicability: whether a tool built and tested somewhere else, on someone else's patients, can be trusted on yours. The regulatory framing is worth internalizing, because it is the language your informatics and compliance colleagues will use: the ONC goal is to give clinical users enough information to judge whether a predictive DSI is fair, appropriate, valid, effective, and safe, sometimes abbreviated FAVES, for the patients in front of them.
It is worth being clear about what these artifacts do not do, because misunderstanding that is where committees go wrong. They do not certify safety. They do not promise the tool works. They do not replace your own local validation or your pilot. A source-attribute panel can be fully populated for a tool that is entirely wrong for your patients, and a model card can be beautifully written for a model that will fail on the very population you serve. The artifacts are inputs to a judgment, not the judgment itself. Their value is that they make the relevant facts visible and, in the case of the source attributes, they make those facts appear in a predictable place and format so that the burden shifts from prying information loose to actually thinking about what it means. A clinical leader who understands this reads them not to be reassured but to find the specific places where the tool and their patients diverge.
The Thirty-One Source Attributes, Enumerated
Here is the part most committees never see, because most lessons on this subject never spell it out. The HTI-1 rule expands the required disclosure to thirty-one source attributes for predictive DSIs (in addition to thirteen for evidence-based DSIs), codified at 45 CFR 170.315(b)(11). They are grouped into nine categories. You do not need to memorize the regulatory subparagraph numbers; you need to know what a clinician is now entitled to demand to see, and where in the panel to find it. The table below is your reading map. When a vendor says a field is proprietary or leaves it blank, you now know exactly which of the thirty-one facts is missing and can name it in writing.
| Category (per 170.315(b)(11)) | What the attributes require the developer to disclose | What you, the clinician, do with it |
|---|---|---|
| 1. Details and output of the intervention | Developer name and contact; funding source for the development; description of the value the output produces; and whether the output is a prediction, classification, recommendation, evaluation, or analysis. | Identify who built it and who paid, and pin down exactly what kind of output you are being asked to trust. |
| 2. Purpose of the intervention | Intended use; intended patient population(s); intended user(s); and the intended decision-making role (informs, augments, or replaces clinical management). | Run your first applicability test: does the intended population and setting match your patients, and is the tool meant to inform or to replace judgment? |
| 3. Cautioned, out-of-scope use | Tasks, situations, or populations where the user is cautioned against applying the intervention; and known risks, inappropriate settings, inappropriate uses, or known limitations. | Read the limitations to predict where the tool will fail in your population before it does. |
| 4. Development details and input features | Exclusion and inclusion criteria for the training data; the input variables used; a description of the demographic representativeness of the training data; and the relevance of that training data to your intended deployed setting. | Answer "was this built on patients like mine," using quantified representativeness, not a "large, diverse dataset" slogan. |
| 5. Fairness in development | The approach taken to ensure the output is fair; and the approaches used to manage, reduce, or eliminate bias. | See whether bias was addressed by design or merely asserted, and demand specifics if the field is boilerplate. |
| 6. External validation process | The data source, clinical setting, or environment where validity and fairness were assessed outside the training source; the party that conducted the external testing; the demographic representativeness of that external data; and a description of the external validation process. | Distinguish a model tested only on its own data from one tested somewhere real and independent. |
| 7. Quantitative measures of performance | Validity and fairness in test data from the same source as training; validity and fairness in external data; and references or citations to studies of the intervention's effect on outcomes such as morbidity, mortality, or length of stay. | Demand the numbers, internal and external, and check whether anyone has shown the tool changes outcomes, not just ranks patients. |
| 8. Ongoing maintenance of implementation and use | The process and frequency by which validity is monitored over time; validity in local data; the process and frequency by which fairness is monitored; and fairness in local data. | Confirm the tool is watched for drift after go-live, including on your own local data, not just at purchase. |
| 9. Update and continued validation schedule | The process and frequency by which the intervention is updated; and the frequency by which performance is corrected when validity or fairness risks are found. | Learn how the tool changes under you, and whether a silent update could invalidate the version you approved. |
Notice what this list quietly guarantees. A vendor can no longer hide the funding source, the demographic makeup of the training data, whether external validation happened and who did it, or whether anyone monitors the model for drift after deployment. Those were once favors you had to beg for. They are now categories the certified system must surface. The rule cannot force the answers to be good, but it forces them to be present or conspicuously absent, and a conspicuous absence is a finding you can act on.
You are now entitled to thirty-one specific facts about any predictive DSI in your EHR. If a vendor will not give you one of them, you can name the exact source attribute that is missing. That is what the rule bought clinicians.
Model Card, Source Attributes, and FDA Clearance Are Three Different Things
Clinical leaders repeatedly collapse three separate things into one comforting blur, and the confusion is dangerous because each answers a different question and none of them answers "safe for my patients." Keep them apart.
FDA clearance or authorization is a regulatory determination that a tool meeting the definition of a medical device may be marketed for a defined intended use. By early 2026 the FDA had authorized more than 1,350 AI and machine-learning-enabled devices, the large majority in radiology. Clearance tells you the device met a regulatory bar for its stated intended use. It does not tell you it works in your workflow, on your population, at your prevalence, and many predictive tools embedded in an EHR are not FDA-regulated devices at all.
ONC HTI-1 transparency, the source attributes, is not an approval and not a safety finding. It is a disclosure requirement. It compels certified health IT to surface a standardized set of facts so a buyer can evaluate the tool. A fully populated source-attribute panel means the facts are visible, not that the tool is good.
Local validation is the only one of the three that actually tests the tool on your patients, in your setting, at your prevalence, with your workflow. It is your responsibility, not the vendor's and not the FDA's. Nothing in FDA clearance or ONC transparency substitutes for it. The source attributes tell you where to aim your local validation; they do not perform it. A tool can be FDA-cleared, have a beautifully complete source-attribute panel, and still fail badly on your patients, and only local validation will catch that before a patient does.
Reading the Model Card Like a Skeptic
A model card is only as honest as the vendor chose to be, so read it the way you would read a study you suspect was written to sell something. Four sections carry most of the signal, and each maps onto a source attribute you can now demand if the card is vague.
The training data section
Start here, because everything downstream depends on it. You want the population described in enough detail to compare to yours: how many patients, from where, over what years, with what demographic and clinical composition. This is source-attribute category 4, development details and input features, which requires a description of demographic representativeness and of the training data's relevance to your setting. A red flag is vagueness dressed as reassurance, phrases like "a large, diverse dataset" with no numbers. Diversity is a claim you can only evaluate if it is quantified. If the training data was drawn from a few academic centers a decade ago, the tool encodes those patients and that era of practice, and your safety-net clinic in 2026 is not that.
The performance table
Look for the same clinical metrics the previous lesson demanded: sensitivity, specificity, PPV, NPV, and, ideally, those numbers broken out by subgroup, not just an aggregate AUROC. This is source-attribute category 7, quantitative measures of performance, which explicitly separates validity in internal test data from validity in external data. A model card that reports only AUROC has told you the model can rank patients but not how it will behave when it fires an alert on a real one. And check whether the performance came from the same population as the training data or from an external validation set, because a model that has only ever been tested on data resembling its training set has not really been tested at all. Treat every headline number, "95 percent accuracy," "AUROC 0.92," as a figure to verify against these categories, never a claim to repeat.
The intended-use statement
This is the most legally and clinically load-bearing sentence in the document, and it is easy to skip. It is source-attribute category 2, purpose of the intervention, which requires the intended use, the intended patient population, the intended user, and the intended decision-making role. Deploying outside that intended use is off-label use of a clinical tool, and it moves you into territory the validation never covered. When the intended use says "to support, not replace, clinician judgment for adult inpatients," those words are a fence, and stepping over it is a decision you are making without evidence. The intended decision-making role is especially telling: a tool designed to inform is a very different liability profile from one a vendor markets as able to replace a step of clinical management.
The limitations section
Read this section first for what it contains and second for whether it exists at all. It is source-attribute category 3, cautioned out-of-scope use, which requires the developer to name the tasks, situations, or populations where the tool should not be applied and the known risks and limitations. A substantive limitations section that names the populations where performance is weaker, the conditions the model was not trained on, and the failure modes the developer has seen is a sign of a serious vendor. A limitations section that is empty, generic, or absent tells you the vendor either does not know their tool's failure modes or chose not to write them down, and both are reasons to slow down. A useful trick: read the limitations section as if you were the vendor's own safety officer who had been told to be honest. If what you see falls far short of what that person would have to admit, someone has been editing for the sales cycle, and the real limitations are still out there waiting to find your patients.
Reading the two artifacts against each other
The most revealing move a skilled evaluator makes is to read the model card and the source attributes side by side and look for the places they disagree or fail to line up. The source attributes are a regulated floor, so they should not contradict the vendor's own model card, and when they do, that gap is a finding you escalate. If the source-attribute fairness field (category 5) is blank but the model card reports subgroup metrics, use the card's data and require the vendor to complete the regulated field for the record. If the model card claims broad applicability but the intended-use attribute is narrow, the narrow one governs, because it is the boundary the certified system is attesting to. Cross-reading turns two separate documents into a single, more honest picture than either produces alone, and it is exactly the kind of scrutiny that a serious vendor expects and a glossy one hopes you will skip.
A model card does not tell you a tool is safe. It tells you enough to decide whether it is safe for your patients, if you actually read it. The vagueness is the finding.
A Worked Example: Reading a Card and Spotting the Gap
A community hospital serving a largely older, lower-income, racially diverse population is evaluating a predictive DSI that flags patients at high risk of clinical deterioration. The demo is persuasive. The CMIO pulls up the source attributes in the EHR and opens the vendor's model card side by side, and reads them against her own patients rather than against her hope.
The intended-use statement (category 2) says the tool is validated for adult inpatients in acute-care settings. Good, that matches. The development-and-validation section (category 4) says the model was trained on 480,000 admissions from three health systems between 2016 and 2021. She reads the representativeness detail: the cohort was 78 percent white, skewed younger than her patients, and drawn from systems with a very different payer mix. The performance category (7) reports an AUROC of 0.85 and, to the vendor's credit, subgroup numbers, which show sensitivity dropping meaningfully for the oldest patients and for one racial subgroup that happens to be a large share of her hospital. Crucially, the external-validation category (6) is thin: the only validation cited is on data from the same three systems, so what the card calls validation is really internal testing. The cautioned-use section (category 3) notes that performance in safety-net settings has not been separately established. The maintenance section (category 8) says the model is re-validated annually by the vendor but does not commit to monitoring validity in local data or to notifying customers of changes.
Now watch what a skilled reading produces. She does not reject the tool, and she does not rubber-stamp it. Reading the enumerated attributes in order, she finds the gap the demo hid: the tool's weakest documented performance lands exactly on her core population, the oldest patients and the racial subgroup her hospital serves most, and its "validation" is internal only. So the artifacts have not told her "yes" or "no"; they have told her precisely where to aim her own local validation before she trusts it, and precisely which subgroups her pilot must prove out. That is what reading these documents is for. They do not make the decision. They tell you where the risk is concentrated so your decision is aimed at it, and so that when governance asks "why did you design the pilot this way," the answer is written in the source attributes, not invented afterward.
When the Attributes Reveal a Non-Representative Population, or Reveal Nothing
Two failure patterns recur, and both have a disciplined response. The first is the honest bad news: the attributes are fully populated and they tell you the training population does not look like yours. This is not a reason to panic and not a reason to proceed blindly. A model trained on a non-representative population predictably underperforms for the patients already underserved, so the correct move is to treat the mismatch as a specification for your local validation. Require the pilot to measure performance in the specific subgroups the attributes flag as weak or absent from training, set a pre-specified threshold for acceptable performance in those subgroups, and refuse to scale beyond the pilot until that threshold is met. The attributes told you where the tool is likely to fail; your job is to prove or disprove that on your own data before a patient inherits the gap.
The second pattern is the empty populated field. Sometimes you will open the source attributes and find the required fields technically present but substantively empty: an intended-use line that says everything and therefore nothing, a validation section that says "internal testing," a fairness field left at a bare minimum. The regulation compels the presence of the categories; it cannot compel candor within them. When the artifacts are thin, that thinness is itself the most important datum you have, and the correct response is not to shrug and proceed. It is to treat the gap as an open question that procurement must resolve before the tool advances: send the specific unanswered questions back to the vendor in writing, naming the exact source attribute that is incomplete, require the missing population and subgroup detail, and record both the request and the response in the governance file for that tool.
This matters for a reason beyond safety. Under the accreditation direction the Joint Commission and CHAI have set, and under the ONC transparency regime, your institution is expected to have evaluated applicability and risk before deployment. A completed reading of the model card and source attributes, with the gaps you found and the answers you demanded, is the artifact that proves you did. When a surveyor or a plaintiff asks how you knew a predictive tool was appropriate for your patients, the strongest possible answer is a documented, skeptical reading of exactly these two artifacts, showing what you checked, what was missing, and what you required before you trusted it. The nutrition label only protects the patient if someone actually reads it, and the record only protects the institution if that reading is written down.
There is also a discipline worth building into your governance rhythm: the reading is not a one-time event. A model card and its source attributes describe the tool as it exists today, and both can change when the vendor updates the model, which is precisely why the update-schedule attribute (category 9) and the contractual right to model-change notification matter so much. Treat a substantive model change as a trigger to re-read the artifacts, because the intended use may have shifted, the validation cohort may have been refreshed, or a new limitation may have appeared. A tool you read carefully in 2026 is not a tool you have read carefully forever. The habit that keeps patients safe is periodic, skeptical re-reading tied to change, not a single approval filed and forgotten.
Key Takeaways
- Two transparency artifacts answer one question, from two directions: the model card is the vendor's own documentation of training data, performance, and limits; the ONC source attributes are a required, standardized nutrition-label panel the certified EHR must surface for every predictive DSI.
- HTI-1 requires thirty-one source attributes for predictive DSIs (codified at 45 CFR 170.315(b)(11)), grouped into nine categories: intervention details and output; purpose and intended use; cautioned out-of-scope use and known limitations; development details and input features; fairness in development; external validation; quantitative performance; ongoing maintenance; and update schedule.
- Because of the rule you are now entitled to specific facts once treated as favors: the funding source, the demographic representativeness of the training data, whether external validation occurred and who did it, and whether the model is monitored for drift; a missing attribute can be named exactly in writing.
- FDA clearance (a device authorization for an intended use), ONC transparency (a disclosure of facts), and local validation (a test on your own patients) are three different things, and only local validation actually tells you the tool works in your setting.
- Read the model card like a skeptic: quantify the training population, demand clinical metrics and subgroup performance rather than a bare AUROC, treat the intended-use statement as a fence, read the limitations section for content and existence, and verify every accuracy number rather than repeating it.
- Read the model card and source attributes against each other: where they disagree, the narrow, regulated attribute governs, and the gap is a finding to escalate.
- When the attributes reveal a non-representative training population, make that mismatch the specification for your pilot, with pre-specified subgroup thresholds you must meet before scaling.
- A documented, skeptical reading of both artifacts, with gaps found and answers demanded, is the record that proves you evaluated applicability before deployment, which is what accreditors, surveyors, and courts will ask you to show.
Skill.re