What Is The NIST AI Risk Management Framework?
- 22 hours ago
- 29 min read

A hospital system spends eight months building a diagnostic AI tool, only to shelve it days before launch because nobody can answer a simple question: who is accountable if the model gives a wrong recommendation and a patient is harmed? That scenario plays out across banks, insurers, and manufacturers every year, not because the AI is broken, but because the organization never built a structured way to think about AI risk before writing the first line of code. The NIST AI Risk Management Framework exists to close that exact gap, and understanding it well is now a basic competency for anyone building, buying, or approving AI systems.
TL;DR
The NIST AI Risk Management Framework (AI RMF) is a voluntary, non-regulatory framework published by the U.S. National Institute of Standards and Technology to help organizations manage risks across the AI lifecycle (NIST, 2023).
As of August 2026, AI RMF 1.0 (NIST AI 100-1, released January 26, 2023) remains the only finalized version; NIST has stated the document is being revised, but no finalized 2.0 exists yet (NIST, 2026).
The framework is organized around four functions: GOVERN, MAP, MEASURE, and MANAGE, which work together rather than in a strict sequence (NIST, 2023).
NIST AI 600-1, the Generative AI Profile, applies the same four functions to generative AI and large language model risks and was finalized on July 26, 2024 (NIST, 2024).
The AI RMF is not a law, not certifiable, and not a checklist; ISO/IEC 42001 is the certifiable management-system counterpart, while the EU AI Act is a binding law with its own separate timeline (ISO, 2023; European Union, 2024).
Effective adoption starts with governance and context (GOVERN and MAP) before an organization jumps into testing metrics (MEASURE) or picking controls (MANAGE).
What Is The NIST AI Risk Management Framework?
The NIST AI Risk Management Framework (AI RMF) is a voluntary framework from the U.S. National Institute of Standards and Technology that helps organizations manage risks in designing, developing, deploying, and using AI systems. Built around four functions, GOVERN, MAP, MEASURE, and MANAGE, it promotes trustworthy AI without imposing binding legal requirements or certification.
Table of Contents
What Is the NIST AI Risk Management Framework?
The NIST AI Risk Management Framework, usually shortened to AI RMF, is a voluntary framework published by the U.S. National Institute of Standards and Technology to help organizations manage risks tied to designing, developing, deploying, evaluating, and using artificial intelligence systems (NIST, 2023). NIST is a non-regulatory agency inside the U.S. Department of Commerce, and the AI RMF was released on January 26, 2023, as NIST AI 100-1 after a public comment process that included two draft versions and multiple workshops (NIST, 2023).
As of August 3, 2026, AI RMF 1.0 is still the only finalized version of the core framework. NIST's own AI RMF page states plainly that "the AI RMF 1.0 is being revised," but no finalized 1.1 or 2.0 document has been published (NIST, 2026). Anything describing a formal "AI RMF 2.0" as already released is describing a rumor, not a NIST publication. NIST has instead extended the original framework through companion documents called Profiles, most notably the Generative AI Profile released July 26, 2024, and a concept note for a Trustworthy AI in Critical Infrastructure Profile released April 7, 2026, which remains an early-stage concept document rather than a finished profile (NIST, 2026).
Managing AI risk is different from eliminating it. The framework does not promise a risk-free AI system, and NIST is explicit that some AI risks can be minimized but not fully removed given current technology (NIST, 2023). The goal is deliberate, documented risk decisions, not a guarantee of zero harm.
A common misconception treats the AI RMF as guidance for machine learning models alone. In practice, NIST frames AI risk as sociotechnical: it spans the model, the data pipelines, the human operators, the organizational processes around deployment, and the people affected by the system's outputs (NIST, 2023). A model that scores well on accuracy benchmarks can still cause real harm if deployed without adequate human oversight, clear escalation paths, or attention to who is affected downstream.
Attribute | Detail |
Publisher | National Institute of Standards and Technology (NIST), U.S. Department of Commerce |
Current status (Aug. 2026) | AI RMF 1.0 finalized (Jan. 26, 2023); a revision is underway with no finalized replacement yet |
Primary document | NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0) |
Legal status | Voluntary; not a law and not certifiable |
Core structure | Four functions: GOVERN, MAP, MEASURE, MANAGE |
Intended users | Any organization designing, developing, deploying, evaluating, or using AI, across sectors and sizes |
Key companion resources | AI RMF Playbook, AI RMF Roadmap, Crosswalks, Generative AI Profile (NIST AI 600-1), NIST AI Resource Center (AIRC) |
Why NIST Created the AI RMF
Conventional software risk management assumes a system behaves the same way every time it runs a given input through fixed logic. AI systems, especially those built on machine learning, break that assumption. Their behavior depends on training data, can shift after deployment as real-world conditions change, and can produce outputs that were never explicitly programmed (NIST, 2023). NIST refers to this as the sociotechnical nature of AI: technical performance and human, organizational, and societal context are inseparable when judging whether a system is safe to use.
Several properties make AI risk management harder than traditional IT risk management: context dependence, where the same model can be low risk in one deployment and high risk in another; emergent behavior that appears only after real users interact with a system at scale; data and model limitations invisible until an incident occurs; automation bias, where people over-trust system outputs; and scale, where a single flawed model can affect millions of decisions before anyone notices (NIST, 2023).
NIST developed the AI RMF between 2021 and 2023, including a Request for Information, two public drafts, several workshops, and formal comment periods before the final release on January 26, 2023 (NIST, 2023). Participation from private industry, federal agencies, academic researchers, and civil-society organizations is why the framework reads as consensus guidance rather than one company's position.
Who the Framework Is For
NIST uses the term AI actor to describe anyone who plays a role in the AI lifecycle, from initial planning through decommissioning. This includes both people who build AI systems and people who are affected by them, which matters because responsibilities differ sharply depending on where an organization sits in the AI value chain (NIST, 2023).
Executive leadership and boards: set risk tolerance, fund governance programs, and are ultimately accountable for AI-related harm.
Risk, compliance, and legal teams: translate regulatory obligations and organizational policy into concrete AI controls.
Security teams: extend existing cybersecurity practice to AI-specific threats such as data poisoning and prompt injection.
Product managers: define intended use, acceptable use, and go/no-go criteria for AI features.
Data scientists and machine learning engineers: build, tune, and document models against agreed requirements.
Model validators and independent evaluators: test claims made by developers before and after deployment.
Internal audit: verifies that governance processes are actually followed, not just written down.
Procurement teams and vendors: manage risk introduced by third-party models, data, and AI-enabled tools.
System operators and end users: apply the system in daily work and are often the first to notice drift or misuse.
Affected individuals and communities: do not build the system but bear the consequences of its decisions.
A large enterprise developing its own foundation model faces different obligations than a small business buying a vendor's AI model and plugging it into a workflow. NIST's framework accommodates both, but the specific controls, evidence, and ownership will look different for a developer versus a deployer versus a downstream user (NIST, 2023).
How NIST Defines AI Risk and Trustworthy AI
NIST defines risk as a measure of the extent to which an entity is negatively influenced by a potential circumstance or event, combining the likelihood of occurrence with the magnitude of impact (NIST, 2023). Risk is not purely negative in the framework's language; AI can also create positive outcomes, and risk management includes weighing benefits against potential harms rather than only cataloguing dangers.
Several distinctions matter for applying this definition correctly. Known risks are ones an organization has already identified and can plan for; unknown risks may only surface after deployment, through monitoring or incident reports. Systemic risks affect an entire sector or society, while contextual risks are specific to one deployment. Risk tolerance is the amount and type of risk an organization is willing to accept in pursuit of its objectives, and it should be set deliberately by leadership rather than left to individual engineers. Residual risk is what remains after mitigation, and every AI system carries some.
Trustworthiness in the AI RMF is not a single score. NIST is explicit that checking one characteristic, such as fairness, does not prove a system is trustworthy overall; a model can be accurate but insecure, or explainable but privacy-invasive (NIST, 2023). Different stakeholders can also reasonably disagree about what "trustworthy" means for a given system: a hospital administrator, a patient, and a regulator may each weigh safety, transparency, and privacy differently, and part of governance is making those trade-offs explicit rather than assuming consensus.
Model performance and system trustworthiness are related but distinct. A model can hit strong accuracy numbers on a held-out test set while the surrounding system, including human oversight, data pipelines, and deployment controls, remains poorly governed. NIST's trustworthy AI characteristics apply to the full sociotechnical system, not only to the underlying model (NIST, 2023).
The Seven Characteristics of Trustworthy AI
NIST AI 100-1 lists seven interrelated characteristics that, taken together, describe trustworthy AI systems. None of them is sufficient alone, and NIST notes that maximizing one can sometimes come at the expense of another, which is why trade-offs need to be documented rather than assumed away (NIST, 2023).
Valid and reliable
Validity means a system performs its intended function accurately; reliability means it does so consistently over time. A fraud model that performs well in testing but degrades as spending patterns shift is unreliable even if once valid. Evidence: benchmark results, confidence intervals, post-deployment accuracy tracking. Limitation: strong lab metrics do not guarantee real-world reliability once conditions change.
Safe
A safe system does not endanger life, health, property, or the environment under reasonable use, including foreseeable misuse. This matters most in physical systems like autonomous vehicles or medical devices, but also applies to financial or psychological harm. Evidence: safety testing, incident logs, fail-safe mechanisms. Limitation: testing cannot cover every real-world scenario in advance.
Secure and resilient
Security covers resistance to attacks such as data poisoning and adversarial inputs; resilience covers recovering and continuing to function after an incident. Evidence: penetration testing, red-team reports, incident-response drills. Limitation: adversarial techniques evolve quickly, so security needs continuous updating, not one-time certification.
Accountable and transparent
Accountability means clear ownership for AI decisions; transparency means stakeholders can access appropriate information about how a system works. Evidence: documented ownership, model cards, disclosure to affected users. Limitation: full transparency can conflict with legitimate confidentiality, so organizations must calibrate what is disclosed to whom.
Explainable and interpretable
Explainability describes representing the mechanisms behind a system's behavior; interpretability describes a human's ability to understand a specific output in context. A credit model can be technically explainable yet still hard for a loan officer to interpret for one customer. Evidence: feature-importance reports, usability testing. Limitation: some high-performing architectures are inherently harder to explain, creating a real accuracy-versus-explainability trade-off.
Privacy-enhanced
Privacy-enhanced systems protect human autonomy and dignity by limiting unnecessary data collection throughout the AI lifecycle. Evidence: data minimization, anonymization or differential-privacy techniques, privacy impact assessments. Limitation: privacy techniques can reduce data available for training, sometimes at a cost to accuracy or fairness.
Fair, with harmful bias managed
Fairness addresses equity and harmful bias, including bias baked into historical data, bias from model design, and bias in how outputs get used. Evidence: disaggregated performance metrics across groups, documented bias-testing procedures. Limitation: multiple mathematical fairness definitions cannot all be satisfied at once, so an organization must choose and justify which criteria matter most for a given use case (NIST, 2023).
Characteristic | Core question it answers |
Valid and reliable | Does it perform its function accurately and consistently over time? |
Safe | Does it avoid endangering life, health, property, or the environment? |
Secure and resilient | Can it resist attacks and recover from incidents? |
Accountable and transparent | Is ownership clear, and can stakeholders access appropriate information? |
Explainable and interpretable | Can humans understand how and why it produced a given output? |
Privacy-enhanced | Does it protect personal data and human autonomy? |
Fair, bias managed | Are harmful biases identified and actively managed? |
How the NIST AI RMF Is Organized
The AI RMF has two main parts. Part one covers foundational information: framing AI risks and introducing the seven trustworthy AI characteristics. Part two is the AI RMF Core, organized into four functions broken into categories and subcategories describing specific outcomes (NIST, 2023). A category groups related outcomes under a function; each subcategory states one actionable outcome. The framework states desired outcomes, while the companion AI RMF Playbook suggests, without mandating, specific actions to reach them (NIST, 2023).
GOVERN is cross-cutting: it informs and is informed by the other three functions rather than sitting before or after them in a strict sequence. NIST is explicit that organizations should revisit earlier functions as new information emerges from later ones (NIST, 2023). A team might complete an initial MAP exercise, start MEASURE testing, discover a missed context, and cycle back before finishing.
Function | Focus | Typical output |
GOVERN | Culture, policy, and accountability across the whole AI risk program | Policies, roles, risk taxonomy, oversight structures |
MAP | Context, purpose, stakeholders, and foreseeable impacts of a specific system | System context documentation, stakeholder and risk mapping |
MEASURE | Testing, metrics, and monitoring of AI risk and performance | Test results, benchmarks, bias and security evaluations |
MANAGE | Prioritizing and responding to identified risks | Risk treatment decisions, deployment controls, incident response |
GOVERN: Build AI Risk Governance
GOVERN covers the policies, processes, and culture that shape how an organization manages AI risk overall. NIST's categories under GOVERN address organizational risk-management processes, roles and responsibilities, workforce diversity and competence, organizational commitment to safety and equity, and processes for engaging with external stakeholders and third parties (NIST, 2023). It is the function most often skipped by teams eager to start building, and its absence tends to surface as confusion later: nobody owns a risk decision, or an incident sits unescalated because no process defines who should be told.
Practical GOVERN work includes setting a documented AI policy that states what uses are acceptable, building a governance structure with clear decision rights, and establishing a risk taxonomy so different teams describe risk the same way. Workforce training matters too: people approving AI deployments need enough technical literacy to ask the right questions, not necessarily to build models themselves.
Common governance artifacts an organization can point to as evidence include:
A written AI policy and acceptable-use guidelines
An AI system inventory listing every model and its intended use
A responsibility matrix naming the owner for each system
Defined approval gates before an AI system moves from testing to production
A documented exception process for deviations from policy
An incident-escalation procedure specific to AI failures
A vendor and third-party AI questionnaire
Training records showing relevant staff have completed AI risk training
A charter for the group or committee overseeing AI risk decisions
Weak governance undermines everything downstream. If nobody owns a risk decision, MAP exercises produce documentation nobody reads, MEASURE results sit in a folder without triggering action, and MANAGE decisions get made ad hoc by whoever is in the room when a problem appears (NIST, 2023).
MAP: Understand Context and Impact
MAP is about establishing the context in which an AI system will operate before deciding how much risk it poses. NIST's MAP categories cover context establishment, categorization of the AI system, understanding of benefits and costs, and identification of risks and impacts across the lifecycle (NIST, 2023). Skipping MAP is the most common shortcut teams take, and it is also the reason MEASURE activities frequently test the wrong things: you cannot design a meaningful test for a risk you never identified.
A MAP exercise should be able to answer concrete questions: What is this system's intended purpose, and who are its intended users? What deployment context will it operate in, and what related dependencies or third-party components does it rely on? What is reasonably foreseeable misuse, even if it is not the intended use? Who are the people affected by this system's decisions, and have their perspectives been considered? What human oversight exists, and where are the boundaries of the system's authority?
An AI system card documenting intended use and known limitations
A use-case description reviewed by both technical and business stakeholders
An impact assessment covering benefits and potential harms
A stakeholder map identifying affected individuals and groups, not just users
A data-flow diagram showing where training and inference data originates
A misuse analysis covering reasonably foreseeable off-label use
Context cannot be assumed from the model type alone. Two chatbots built on the same underlying language model carry very different risk profiles if one answers internal HR questions for employees and the other gives medical guidance to the public.
MEASURE: Analyze and Assess AI Risk
MEASURE covers the testing, evaluation, and monitoring needed to understand whether an AI system behaves as intended and where it falls short. NIST's MEASURE categories address appropriate methods and metrics, evaluation for trustworthy characteristics, mechanisms for tracking risks over time, and processes for gathering feedback from stakeholders. This work is often called TEVV, short for testing, evaluation, verification, and validation, and it spans the entire lifecycle rather than a single pre-launch checkpoint (NIST, 2023).
Four distinct things get measured, and conflating them is a frequent mistake. Model performance covers metrics like accuracy or F1 score on held-out data. System performance covers how the model behaves once wired into a real application with real user inputs, which often differs from lab conditions. Real-world impact covers what actually happens to people and the business after deployment, sometimes only visible weeks or months later. Governance-process effectiveness covers whether the GOVERN policies are actually followed, which requires auditing the process itself, not just the model.
Choosing metrics well means avoiding reliance on a single score. A model can hit strong overall accuracy while performing far worse for one demographic subgroup, a gap only visible through disaggregated analysis. Thresholds and baselines should be set deliberately, with documented confidence levels, and independent evaluation, ideally by someone other than the model's own developers, adds credibility that self-reported numbers cannot.
Quantitative evidence: benchmark scores, fairness metrics by subgroup, latency, uptime, false-positive and false-negative rates, drift-detection statistics
Qualitative evidence: red-team findings, user feedback themes, expert review notes, documented edge cases
Two related failure modes NIST's own red-teaming and concept-drift literature both flag are worth tracking over time: models that were valid at launch can quietly degrade as real-world data shifts, and models tested only once before deployment give a false sense of durability.
MANAGE: Prioritize and Treat AI Risk
MANAGE covers deciding what to do about the risks identified through MAP and MEASURE. NIST's MANAGE categories address risk prioritization, resource allocation for risk treatment, and processes for responding to and recovering from incidents (NIST, 2023). Standard risk-treatment options apply: mitigate the risk directly, accept it because it falls within tolerance, avoid it by not deploying, or transfer it, for example through insurance or contractual terms with a vendor.
Prioritization should weigh severity, likelihood, which populations are affected, how reversible the harm is, how easily an issue would be detected in practice, and the organization's stated risk tolerance from GOVERN. A low-severity but high-frequency issue may deserve more urgent attention than a severe but vanishingly rare one, depending on that calculus.
Example risk | Possible response | Owner | Evidence | Monitoring indicator |
Model shows lower accuracy for one demographic group | Mitigate: retrain with rebalanced data, add human review for that segment | Model owner + risk lead | Disaggregated test results, retraining log | Subgroup accuracy tracked monthly |
Generative assistant occasionally fabricates citations | Mitigate: add retrieval grounding, require citation verification step | Product owner | Evaluation report, UI change log | Confabulation rate in sampled outputs |
Vendor model has undisclosed training-data sources | Transfer/avoid: renegotiate contract terms or seek alternative vendor | Procurement + legal | Vendor questionnaire, contract addendum | Vendor risk review at renewal |
Low-severity UI confusion affecting few users | Accept: document and monitor, no immediate action | Product owner | Risk register entry | User complaint volume |
Practical MANAGE controls include human review checkpoints before high-stakes decisions, a documented rollback or shutdown capability if a deployed system misbehaves, defined incident-response steps specific to AI failures, and a channel for communicating known limitations to affected users. Continuous improvement closes the loop: incidents and monitoring data should feed back into GOVERN policy updates and MAP context documentation.
How the Four Functions Work Together
GOVERN, MAP, MEASURE, and MANAGE are mutually reinforcing rather than four sequential project phases you complete once and move past. NIST frames them as continuous activities that inform one another throughout an AI system's life (NIST, 2023). A monitoring finding under MEASURE can trigger a MANAGE response, which can prompt a GOVERN policy update, which then reshapes how the next system gets MAPped.
Idea intake: a business unit proposes an AI use case; GOVERN policy determines whether it needs formal risk review.
Context mapping: the team documents intended use, stakeholders, and foreseeable misuse (MAP).
Development and testing: engineers build the system while MEASURE activities run in parallel, not only at the end.
Pre-deployment review: MANAGE decision-makers weigh MEASURE results against GOVERN-set risk tolerance to approve, modify, or reject deployment.
Deployment with monitoring: MEASURE continues in production; drift or incidents trigger MANAGE response.
Periodic review: findings feed back into GOVERN policy and the system's MAP documentation is updated.
Retirement: when a system is decommissioned, GOVERN records the decision and lessons feed into future MAP exercises.
Treating the functions as a rigid waterfall, MAP once, MEASURE once, MANAGE once, is a common misapplication. NIST's own language stresses that organizations should revisit earlier functions as new information surfaces (NIST, 2023).
AI RMF Profiles, the Playbook, and the NIST AI Resource Center
A Profile applies the AI RMF's four functions to a specific use case, sector, or technology, translating general outcomes into more specific guidance. NIST supports both current-state profiles, describing how an organization manages AI risk today, and target-state profiles, describing a desired future state, which together let an organization run a gap analysis and build a roadmap for closing the difference (NIST, 2023).
The companion AI RMF Playbook, published alongside the framework, offers suggested actions, transparency questions, and reference resources tied to each subcategory in the Core. NIST does not present the Playbook as a mandatory checklist; organizations are expected to select actions that fit their context, size, and risk tolerance rather than implement every suggestion (NIST, 2023).
The NIST AI Resource Center, known as AIRC, hosts the Playbook, use cases submitted by other organizations, crosswalks mapping the AI RMF to other frameworks, and technical resources related to AI risk. A crosswalk shows conceptual overlap between the AI RMF and another standard; it does not mean the two documents are identical or that satisfying one automatically satisfies the other (NIST, 2023). Similarly, a third-party resource appearing on an AIRC use-case page reflects that the organization submitted it, not that NIST has vetted or endorsed its content.
The NIST Generative AI Profile
NIST AI 600-1, the Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, was finalized on July 26, 2024. It applies the same GOVERN, MAP, MEASURE, and MANAGE structure to risks that are specific to, or significantly exacerbated by, generative AI and large language models (NIST, 2024).
The Generative AI Profile identifies twelve risk categories unique to or amplified by generative AI, including confabulation, or the generation of false content presented as fact; harmful bias and homogenization; dangerous, violent, or hateful content; data privacy; information integrity; information security; intellectual property concerns; obscene, degrading, or abusive content; value-chain and component integration risks; human-AI configuration risks such as automation bias and overreliance; environmental impacts; and CBRN information risks (NIST, 2024).
Applying MAP to a generative AI use case means examining not just the base model but the surrounding pipeline: retrieval sources, plugins or tools the model can invoke, and how outputs reach end users. MEASURE for generative systems often includes confabulation rate testing, prompt-injection resistance testing, and content-safety evaluation, in addition to conventional accuracy metrics. MANAGE responses commonly include output filtering, mandatory citation or provenance disclosure, and human review gates for high-stakes generated content. Organizations evaluating an internal AI security posture for generative tools typically layer this profile on top of standard cybersecurity controls rather than replacing them.
How to Implement the NIST AI RMF Step by Step
Implementing an organization-wide AI risk program is different from assessing a single AI system. The steps below describe building the program; a single-system assessment is a smaller exercise that a mature program eventually runs for every new use case.
Establish executive sponsorship (CEO, CIO, or risk officer) with real authority to block a risky launch. Mistake to avoid: assigning ownership to staff who cannot say no. Small teams: one accountable executive beats a committee with no owner.
Define scope: which systems and business units the program covers first. Mistake: trying to cover everything in year one instead of the highest-risk uses.
Inventory AI systems and use cases, including informally adopted shadow AI tools. Mistake: counting only internally built models and missing embedded vendor tools.
Classify systems by risk severity and likelihood. Mistake: applying one generic risk tier to every system regardless of use case.
Assign a named, accountable owner per system through a governance committee. Mistake: leaving ownership ambiguous between business and technical teams.
Map stakeholders, context, benefits, and harms for priority systems (MAP). Mistake: mapping only intended users and skipping people affected indirectly.
Select which trustworthy-AI characteristics matter most for each system and why. Mistake: treating all seven as equally critical everywhere.
Define tests, metrics, and evidence to retain (MEASURE), ideally with independent evaluators. Mistake: relying only on a developer's self-reported results.
Assess and prioritize risks by severity, likelihood, and reversibility. Mistake: fixing whatever is easiest rather than what causes the most harm.
Implement controls and approval gates before production deployment (MANAGE). Mistake: treating pre-launch sign-off as the only checkpoint.
Monitor deployed systems continuously for drift, incidents, and user feedback. Mistake: tracking uptime but ignoring fairness or accuracy drift.
Review and update the program, feeding incidents back into policy and future MAP work. Mistake: never revisiting policy after rollout.
A phased rollout for a first-time program might target early governance basics within 30 days, a documented inventory and risk classification by 60 days, initial MAP and MEASURE work on the highest-priority system by 90 days, and a full first cycle through MANAGE with a documented review by 180 days. These are illustrative benchmarks; actual timelines depend heavily on organization size, existing risk infrastructure, and how many AI systems are already in production.
Practical Example: Applying the AI RMF to a Generative AI Assistant
The following is an illustrative hypothetical example, not a real company or documented case, meant to show how the four functions connect in practice.
A mid-sized insurance company considers an internal generative AI assistant to help claims adjusters draft correspondence and summarize policy documents. Intended use: drafting support, not final decision-making. Stakeholders: adjusters, claimants who receive the correspondence, compliance staff, and the vendor providing the model. Benefit: faster, more consistent responses.
Foreseeable misuse includes adjusters pasting sensitive claimant health data into prompts, and the assistant fabricating a policy detail that does not exist. The company must confirm whether the vendor retains prompts for further training, a MAP-stage question with direct privacy implications, and treat the vendor relationship itself as a governance concern.
Key risks: confabulated details reaching a claimant, health data leaking through prompts, and adjusters over-relying on drafted text without review. Governance decision: no AI-drafted correspondence goes to a claimant without human sign-off. MEASURE work samples outputs for factual accuracy against source policy documents. MANAGE controls mask sensitive fields before they reach the prompt, require human review, and include a kill switch if quality drops below an agreed threshold.
Monitoring tracks the error rate caught in review and claimant complaints. Residual risk remains: reviewers can still miss a fabricated detail. The final decision authorizes a limited pilot with one team rather than a full rollout, an unresolved trade-off between adoption speed and confidence in the evidence gathered so far.
Evidence, Metrics, and Documentation to Maintain
Disciplined AI risk management leaves a paper trail. NIST's framework does not mandate a specific document set, but the following categories reflect what most mature implementations retain: an AI system inventory; system and model cards; data documentation describing sources and known limitations; impact assessments; a risk register; test plans and results; evaluation reports, including fairness and privacy reviews; threat models and red-team reports; vendor risk assessments; deployment approval records; change logs; monitoring dashboards; incident records; user feedback logs; and retirement decisions when a system is decommissioned.
Metric or indicator | What it measures | Data source | Owner | Review frequency | Limitation |
Subgroup accuracy gap | Fairness across demographic segments | Test results, production logs | Data science + risk lead | Quarterly | Requires demographic labels, which may be unavailable or sensitive to collect |
Confabulation rate | Share of generated outputs with fabricated facts | Sampled output review | Product owner | Monthly for high-use systems | Manual review does not scale to every output |
Incident count and severity | Frequency and impact of AI-related failures | Incident-management system | Risk lead | Monthly | Under-reporting if staff do not know how to flag an AI incident |
Human override rate | How often reviewers reject or edit AI output | Workflow logs | Product owner | Monthly | High override rate could reflect either poor model quality or overly cautious reviewers |
Vendor risk score | Third-party model and data risk posture | Vendor questionnaires, contracts | Procurement + legal | At contract renewal | Self-reported vendor answers are not independently verified by default |
Governance policy adherence | Whether required approval gates were followed | Internal audit sampling | Internal audit | Semiannual | Sampling may miss isolated but serious lapses |
Documentation alone does not prove risk is well managed. A thick binder of unread risk assessments provides no protection if nobody acts on the findings; evidence matters only when it feeds real decisions in GOVERN and MANAGE (NIST, 2023).
NIST AI RMF vs. Other AI Governance Standards and Laws
Organizations frequently need to reconcile the AI RMF with other standards and laws. They serve different functions and are not interchangeable, and using one does not automatically satisfy the requirements of another. This is general educational information, not legal advice; organizations should consult qualified counsel for specific compliance obligations.
Instrument | Type | Purpose | Mandatory? | Certification? |
NIST AI RMF | Voluntary risk-management framework | Structure AI risk management using GOVERN, MAP, MEASURE, MANAGE | No | Not certifiable |
ISO/IEC 42001:2023 | International management-system standard | Specify requirements for an auditable AI management system (AIMS) | No, but adopted voluntarily by many organizations | Yes, third-party certifiable (ISO, 2023) |
ISO/IEC 23894:2023 | International risk-management guidance | Detailed AI-specific risk-management guidance, extending ISO 31000 | No | Not certifiable |
NIST Cybersecurity Framework 2.0 | Voluntary risk-management framework | Manage cybersecurity risk broadly; complements AI-specific security controls | No | Not certifiable |
EU AI Act (Regulation 2024/1689) | Binding law | Regulate AI systems in the EU market by risk tier, with obligations and penalties | Yes, for systems in scope | Conformity assessment required for high-risk systems, not a voluntary certification |
ISO/IEC 42001 is a certifiable management-system standard, meaning an accredited body can audit and certify an organization against it, similar to how ISO 27001 certifies information-security management. ISO/IEC 23894 provides detailed AI risk-management guidance that many organizations use as a companion methodology when implementing either the AI RMF or ISO 42001's risk-assessment clauses (ISO, 2023). The two ISO standards and the NIST AI RMF are widely described as complementary: work performed to satisfy AI RMF outcomes often carries directly into an ISO 42001 audit, but completing one does not automatically satisfy the other's specific clause requirements.
The EU AI Act, Regulation (EU) 2024/1689, entered into force August 1, 2024, with penalties up to €35 million or seven percent of global turnover for the most serious violations (European Union, 2024). Its original timeline set high-risk Annex III obligations to apply from August 2, 2026. EU negotiators reached a provisional agreement in May 2026 to defer those obligations to December 2, 2027, but as of this writing that agreement had not been formally adopted, so organizations should confirm current status rather than assume the delay is final (European Union, 2026). Either way, the Act is binding law with real penalties, unlike the AI RMF's voluntary status.
Benefits, Limitations, and Common Mistakes
The AI RMF's biggest practical benefit is a shared vocabulary. Teams from legal, engineering, and business functions can discuss AI risk using the same four functions and seven characteristics instead of talking past each other. Its voluntary, flexible design also lets organizations of very different sizes and sectors adopt it without a one-size-fits-all compliance burden, and its lifecycle framing pushes teams to think past launch day toward ongoing monitoring.
The same flexibility creates real limitations. Broad, outcome-based language requires interpretation, which means two organizations can both claim AI RMF alignment while implementing very different controls. Measurement gaps persist, especially for societal-level impacts that are hard to quantify. The framework depends heavily on organizational maturity: a company with weak general risk management will struggle to apply AI-specific guidance well. And because there is no certification, an organization cannot point to an independent seal of approval the way it could with ISO 42001.
Treating the framework, or the Playbook, as a rigid checklist rather than outcome-based guidance to adapt.
Focusing risk assessment only on model accuracy while ignoring governance and downstream impact.
Skipping engagement with affected communities and relying solely on internal technical review.
Starting MEASURE testing before finishing MAP context work, which leads to testing the wrong things.
Treating governance as paperwork rather than as decisions that actually shape system design.
Leaving risk decisions without a named, accountable owner.
Ignoring vendor and third-party models when they carry as much risk as internally built systems.
Applying generic risk thresholds across very different use cases instead of context-specific ones.
Stopping monitoring once a system reaches production instead of continuing to test after deployment.
Assuming a crosswalk to another standard proves equivalence rather than conceptual overlap.
Applying identical controls to every AI system regardless of its actual risk classification.
FAQ
What does NIST AI RMF stand for?
NIST AI RMF stands for the National Institute of Standards and Technology Artificial Intelligence Risk Management Framework. NIST is a U.S. Department of Commerce agency, and the RMF is its voluntary framework, published as NIST AI 100-1 on January 26, 2023, for managing risks across the AI lifecycle (NIST, 2023). It is not a law, regulation, or certification scheme.
Is the NIST AI RMF mandatory?
No. The framework is explicitly voluntary and non-regulatory (NIST, 2023). Some U.S. federal agencies reference or require AI RMF alignment in specific contracts or guidance, and it is often cited by regulators as a benchmark of good practice, but there is no general legal requirement for private organizations to adopt it, and no penalty exists for choosing not to.
Who should use the framework?
Any organization that designs, develops, deploys, evaluates, or uses AI systems can use it, regardless of sector or size (NIST, 2023). This includes technology companies building models, businesses deploying vendor AI tools, and public-sector agencies. Responsibilities differ by role: a model developer and a company simply using a purchased AI tool will apply different parts of the framework.
What are the four NIST AI RMF functions?
The four functions are GOVERN, which covers policy and accountability; MAP, which covers understanding context and impact; MEASURE, which covers testing and evaluation; and MANAGE, which covers prioritizing and treating identified risks (NIST, 2023). GOVERN is cross-cutting and interacts with all three other functions rather than acting as a separate first step.
What are the trustworthy-AI characteristics?
NIST lists seven: valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed (NIST, 2023). No single characteristic proves overall trustworthiness on its own, and improving one can sometimes trade off against another.
Is NIST AI RMF only for organizations in the United States?
No. Although NIST is a U.S. government agency, the framework is technology-neutral and has been referenced internationally, including through official translations into Arabic and Japanese (NIST, 2023). Multinational organizations often use it alongside regional requirements such as the EU AI Act.
Does NIST certify AI RMF compliance?
No. NIST does not audit or certify organizations against the AI RMF, and there is no official "NIST-certified" status for a company or AI system. Organizations self-assess and can voluntarily state alignment, but this differs sharply from ISO/IEC 42001, which supports third-party accredited certification (ISO, 2023).
Is the NIST AI RMF a checklist?
No. The framework describes outcomes through categories and subcategories, while the companion Playbook suggests possible actions to reach those outcomes. NIST does not present the Playbook as a mandatory checklist, and organizations are expected to select and adapt actions based on their own context and risk tolerance (NIST, 2023).
What is an AI RMF Profile?
A Profile applies the framework's four functions to a specific use case, sector, or technology, such as generative AI. NIST supports both current-state profiles describing existing practice and target-state profiles describing a goal, which together allow gap analysis and roadmap planning (NIST, 2023).
What is the NIST AI RMF Playbook?
The Playbook is a companion resource published alongside the framework that offers suggested actions, transparency questions, and reference materials tied to each subcategory in the AI RMF Core. It is guidance to adapt, not a mandatory compliance checklist (NIST, 2023).
How does NIST AI RMF apply to generative AI?
NIST AI 600-1, the Generative AI Profile finalized July 26, 2024, applies GOVERN, MAP, MEASURE, and MANAGE to risks specific to or amplified by generative AI, identifying twelve risk categories such as confabulation, data privacy, and information security (NIST, 2024). Organizations use the same four-function structure but focus testing and controls on generative-AI-specific failure modes.
How does NIST AI RMF differ from ISO/IEC 42001?
The AI RMF is voluntary risk-management guidance with no certification. ISO/IEC 42001, published in December 2023, is an international, certifiable management-system standard that an accredited body can audit (ISO, 2023). The two are considered complementary: AI RMF work often supports an ISO 42001 audit, but completing one does not automatically satisfy the other's specific requirements.
Does using NIST AI RMF ensure compliance with the EU AI Act?
No. The EU AI Act is a binding law with specific legal obligations and penalties, while the AI RMF is voluntary guidance (European Union, 2024). Aligning with the AI RMF can support a broader AI governance program but does not by itself satisfy the AI Act's conformity-assessment or documentation requirements. This is general information, not legal advice, and organizations should consult qualified legal counsel for compliance questions.
How often should an organization review its AI risks?
NIST frames risk management as continuous rather than a one-time event, with MEASURE and MANAGE activities running throughout deployment, not only before launch (NIST, 2023). Many organizations set quarterly reviews for active systems and trigger additional review after any significant model update, incident, or change in deployment context.
Key Takeaways
As of August 2026, AI RMF 1.0 (NIST AI 100-1, January 2023) remains the only finalized core framework; NIST has confirmed a revision is underway but has not finalized a replacement.
The framework is voluntary and non-certifiable; ISO/IEC 42001 is the certifiable management-system alternative for organizations that want third-party audit.
GOVERN is cross-cutting: weak governance undermines MAP, MEASURE, and MANAGE even when those three functions are executed well individually.
No single trustworthy-AI characteristic proves overall system trustworthiness, and improving one can trade off against another.
The Generative AI Profile (NIST AI 600-1, July 2024) identifies twelve risk categories specific to or amplified by generative AI, including confabulation and data privacy.
The Playbook offers suggested, adaptable actions, not a mandatory checklist, and organizations should select actions based on context and risk tolerance.
Documentation only reduces risk when it drives real GOVERN and MANAGE decisions, not when it sits unread in a folder.
The EU AI Act is a binding law with financial penalties; the AI RMF is voluntary guidance, and the two should not be treated as interchangeable.
Actionable Next Steps
No formal AI governance yet: name one accountable executive, build a basic AI system inventory, and draft a one-page acceptable-use policy before adopting any new AI tool.
Already using AI without a formal program: run a MAP exercise on your highest-risk system first, then retrofit GOVERN policy based on what that exercise reveals.
Procuring third-party AI: add an AI-specific vendor questionnaire to procurement covering training-data sources, data retention, and known limitations before signing new contracts.
Deploying generative AI specifically: review the twelve risk categories in NIST AI 600-1, prioritize confabulation and data-privacy testing, and require human review before high-stakes outputs reach end users.
Mature organization integrating multiple standards: map AI RMF subcategories against ISO/IEC 42001 clauses and, if operating in the EU, confirm current EU AI Act deadlines directly with legal counsel rather than relying on older news coverage.
Glossary
AI actor: Anyone who plays a role in the AI lifecycle, including developers, deployers, evaluators, operators, and people affected by an AI system.
AI lifecycle: The full span of an AI system's existence, from planning and design through development, deployment, monitoring, and retirement.
AI risk: A measure of the extent to which an entity may be negatively affected by an AI-related event, combining likelihood and impact.
Artificial intelligence system: A machine-based system that can, for a given set of human-defined objectives, generate outputs such as predictions, recommendations, or decisions.
Affected individual: A person impacted by an AI system's outputs, whether or not they directly use the system.
Bias: Systematic deviation that unfairly disadvantages individuals or groups; NIST distinguishes systemic, computational, and human forms of bias.
Confabulation: A generative AI system producing false or fabricated content presented as though it were factual, sometimes called hallucination.
Current Profile: A profile describing how an organization manages AI risk today, used as a baseline for gap analysis.
Explainability: The ability to represent the underlying mechanisms of how an AI system operates.
Fairness: The property of an AI system being equitable, with harmful bias identified and actively managed.
Generative AI: AI systems capable of generating new content, such as text, images, audio, or code, based on learned patterns.
GOVERN: The AI RMF function covering culture, policy, and accountability for AI risk across an organization.
Human oversight: The presence of meaningful human review or control over an AI system's decisions or actions.
Impact assessment: A structured evaluation of an AI system's potential benefits and harms to individuals, groups, and society.
Interpretability: The ability of a person to understand the meaning of an AI system's output in a specific context.
MANAGE: The AI RMF function covering prioritization and treatment of identified AI risks.
MAP: The AI RMF function covering context establishment and identification of an AI system's benefits and impacts.
MEASURE: The AI RMF function covering testing, evaluation, and ongoing monitoring of AI risk and performance.
Model drift: A decline in an AI model's performance over time as real-world data diverges from its training data, closely related to concept drift.
Profile: A document that applies the AI RMF's four functions to a specific use case, sector, or technology.
Residual risk: The risk that remains after mitigation controls have been applied.
Risk tolerance: The amount and type of risk an organization is willing to accept in pursuit of its objectives.
Sociotechnical system: A system whose behavior and risk depend on the interaction of technology with people, organizations, and society, not the technology alone.
Target Profile: A profile describing an organization's desired future state for managing a specific AI risk area.
TEVV: Testing, evaluation, verification, and validation, the umbrella term for MEASURE-function activities across an AI system's lifecycle.
Trustworthiness: A composite property of an AI system reflecting validity, safety, security, accountability, explainability, privacy, and fairness together, not any single trait alone.
Validation: Confirming that an AI system meets the needs of its intended use and users.
Verification: Confirming that an AI system has been built correctly according to its specifications.
Sources & References
National Institute of Standards and Technology. "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," NIST AI 100-1. Published January 26, 2023. https://doi.org/10.6028/NIST.AI.100-1 (accessed August 3, 2026).
National Institute of Standards and Technology. "AI Risk Management Framework." NIST ITL webpage, last modified June 10, 2026. https://www.nist.gov/itl/ai-risk-management-framework (accessed August 3, 2026).
National Institute of Standards and Technology. "NIST AI RMF Playbook." NIST AI Resource Center. https://airc.nist.gov/AI_RMF_Knowledge_Base/Playbook (accessed August 3, 2026).
National Institute of Standards and Technology. "AI Risk Management Framework Roadmap." https://www.nist.gov/itl/ai-risk-management-framework/roadmap-nist-artificial-intelligence-risk-management-framework-ai (accessed August 3, 2026).
National Institute of Standards and Technology. "Crosswalks to the NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0)." https://www.nist.gov/itl/ai-risk-management-framework/crosswalks-nist-artificial-intelligence-risk-management-framework (accessed August 3, 2026).
National Institute of Standards and Technology. "AI Risk Management Framework FAQs." https://www.nist.gov/itl/ai-risk-management-framework/ai-risk-management-framework-faqs (accessed August 3, 2026).
National Institute of Standards and Technology. "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," NIST AI 600-1. Published July 26, 2024. https://doi.org/10.6028/NIST.AI.600-1 (accessed August 3, 2026).
National Institute of Standards and Technology. "Concept Note: AI RMF Profile on Trustworthy AI in Critical Infrastructure." Published April 7, 2026. https://www.nist.gov/programs-projects/concept-note-ai-rmf-profile-trustworthy-ai-critical-infrastructure (accessed August 3, 2026).
National Institute of Standards and Technology. "Trustworthy and Responsible AI Resource Center (AIRC)." Launched March 30, 2023. https://airc.nist.gov/Home (accessed August 3, 2026).
International Organization for Standardization. "ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system." Published December 2023. https://www.iso.org/standard/42001 (accessed August 3, 2026).
International Organization for Standardization. "ISO/IEC 23894:2023, Information technology — Artificial intelligence — Guidance on risk management." Published 2023. https://www.iso.org/standard/77304.html (accessed August 3, 2026).
National Institute of Standards and Technology. "The NIST Cybersecurity Framework (CSF) 2.0." Published February 26, 2024. https://doi.org/10.6028/NIST.CSWP.29 (accessed August 3, 2026).
European Union. "Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act)." Entered into force August 1, 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj (accessed August 3, 2026).
European Commission. "Timeline for the Implementation of the EU AI Act." AI Act Service Desk. https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act (accessed August 3, 2026).
Vassilev, A. et al. "Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations," NIST AI 100-2e2025. National Institute of Standards and Technology, published March 2025. https://doi.org/10.6028/NIST.AI.100-2e2025 (accessed August 3, 2026).