On April 17, 2026 the Fed, OCC and FDIC replaced SR 11-7 with SR 26-2, OCC Bulletin 2026-13 and FDIC FIL-15-2026, and rescinded the 2021 BSA/AML model risk statement in the same move. Footnote 3 of the 12-page attachment says generative AI and agentic AI models are not within the scope of the guidance, while non-generative, non-agentic AI models still are. The guidance never defines those terms and never mentions embeddings, rerankers, classifiers or components of a larger system, so classifying each part of an LLM stack is your documented judgment. Our reading: guardrail classifiers, rerankers and risk-scoring layers that gate decisions are in scope, embedding models are the ambiguous row, and the generator and agent loop are out of scope but still governed by your own risk practices, because footnote 1 keeps supervisory action open for unsafe or unsound practices. The guidance is expected to be most relevant above $30 billion in total assets, and the agencies promised an AI-focused request for information that had not been published as of late September 2026. Start this week by re-anchoring your policy citations and walking one production LLM system component by component, recording a classification rationale for each part in the inventory.
If your model risk policy opens with a citation to SR 11-7, it cites a letter that no longer exists. On April 17, 2026 the Federal Reserve, the OCC and the FDIC replaced it, and the replacement (SR 26-2 model risk management guidance, issued in parallel as OCC Bulletin 2026-13 and FDIC FIL-15-2026) settles one question for AI platform teams in a single footnote: generative AI and agentic AI models are not within its scope.
That footnote does not settle what model risk, compliance and platform leads actually have to decide, which is what goes into the inventory. A production LLM system at a bank can be several models at once. A self-hosted retrieval and agent deployment can include an open-weight generator, an embedding model, a reranker, one or more guardrail classifiers, a scoring layer on the output and an agent loop calling tools. Some of those are generative. Several are not, and the guidance says non-generative AI models are still in scope.
This post applies the 2026 model definition to each of those components, says where we think each one lands, and lists the evidence to keep for the parts that fall outside. It is not a line-by-line diff against SR 11-7, and it does not re-explain model risk management to people who run it. The per-component calls are our reading, labelled as such, because the agencies did not make them.
What SR 26-2 replaced on April 17, 2026
All three agencies issued and rescinded on the same day. None of the three documents names a separate effective date.
One rescission deserves a second look. The 2021 BSA/AML statement is the document that addressed how model risk management guidance applies to bank systems supporting BSA/AML compliance. It is gone, so any AML validation standard that cites it needs a new anchor, even if the controls it describes stay exactly where they are.
The Fed letter says the revision reflects supervisory experience and industry feedback accumulated over the past fifteen years. The attachment is 12 pages in seven sections: introduction, purpose and scope, model risk overview, development and use, validation and monitoring, governance and controls, and vendor products. For an AI team, almost everything consequential happens on page 3.
| Agency | Issued | Rescinded |
|---|---|---|
| Federal Reserve | SR 26-2, Revised Guidance on Model Risk Management | SR 11-7 (April 4, 2011); SR 21-8, the interagency BSA/AML model risk statement (April 9, 2021) |
| OCC | Bulletin 2026-13, Model Risk Management: Revised Guidance | Bulletin 2011-12; Bulletin 2021-19; Bulletin 1997-24 on credit scoring models, including its appendix; the Model Risk Management booklet of the Comptroller's Handbook |
| FDIC | FIL-15-2026, Agencies Revise the Interagency Model Risk Management Guidance | FIL-22-2017 (the June 2017 adoption of the 2011 guidance); FIL-27-2021 |
The scope footnote, word for word
Section II defines a model this way:
For the purposes of this guidance, the term "model" refers to a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates. The term "model" in this guidance excludes simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes and software where there are no statistical, economic, or financial theories underpinning their design or use.
Footnote 3, attached to that definition, reads in full:
Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. Nonetheless, a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models.
The OCC bulletin repeats the first two sentences in both its highlights and its background section, so this is not a buried aside.
Now read what the text does not do. It does not define "generative AI model", "agentic AI model" or "non-generative, non-agentic AI models". It never mentions embeddings, rerankers, retrieval, prompts, or components of a composite system. The agencies drew a line between two categories and left the institution to decide which side of the line each piece of a real architecture falls on.
A senior official's gloss came ten days later. In remarks delivered April 27, 2026 and published May 1, the Fed's Vice Chair for Supervision said the revised guidance "now applies narrowly to traditional models and basic AI applications", and that supervisors had, over time, expanded the previous guidance beyond its original purpose. That tells you the direction of intent. A speech is not guidance, though, and "basic AI applications" is not a term an examiner can test against. Your classification has to rest on the definition itself.
Decomposing an LLM system against the model definition
The definition gives you three tests. Does the component produce quantitative estimates? Does it apply statistical theory, as opposed to being a deterministic rule-based process? And under footnote 3, is it generative or agentic? Run every component through all three.
The last column below is Particula's reading of the definition, not agency text. No agency has classified any of these components.
The classifiers are the clear case
A fine-tuned discriminative classifier that decides whether a request goes to the generator, or whether an output gets blocked, is a non-generative AI model producing a quantitative estimate that drives a decision. Footnote 3's last sentence covers it. The same applies to a scoring layer that converts an LLM response into a confidence band and routes low-confidence items to review. Guardrail stacks built from these parts, like the ones we compared in NeMo Guardrails, Llama Guard and Guardrails AI, can contain at least one in-scope model even when the team thinks of the whole stack as "the LLM". Llama Guard is the borderline row. It is a fine-tuned Llama language model that returns its safety verdict as generated text, so its architecture is generative while its job is classification. We would not argue either side with confidence. Write down which way you classified it and why, and if you call it out of scope, apply the out-of-scope evidence set below to it anyway. AML triage is the same pattern at higher stakes. In an AML alert triage pipeline, any statistical or non-generative AI model that scores, ranks, suppresses or closes an alert is in scope, while the LLM that drafts the evidence narrative is the generative part. The rescission of the 2021 statement does not change a scoring model's status under the 2026 definition.
The reranker produces a score
A cross-encoder reranker outputs a relevance score per passage, and the order it produces determines what the generator sees. That is a quantitative estimate from a statistical model. Whether it matters depends on materiality, not on whether it is a model: a reranker in front of an internal policy search tool and a reranker selecting evidence for a credit memo are the same model with very different exposure. If you are still deciding whether the reranker earns its place at all, the trade-offs are in when RAG reranking actually improves retrieval.
The embedding model is the genuinely ambiguous row
An embedding model is a trained neural network. It is not generative in function. By elimination that puts it on the non-generative side of footnote 3. But its output is a vector, not an estimate anyone uses directly in a decision, and the definition speaks of processing input data "into quantitative estimates". Reasonable validators will disagree. Do not resolve that ambiguity by leaving it off the list. Section III gives you the tool: "Banking organizations may deem certain models immaterial based on model exposure and purpose. In those cases, model risk management may consist of identifying those models and monitoring model performance and conditions under which the use of those models may become material to the banking organization in the future." Inventory it, record the classification rationale, set it as low materiality if that is your judgment, and monitor retrieval quality. A silent omission is the only answer you cannot defend.
Assess the in-scope parts together
Section III also says "Sound practice involves assessing model risk both individually and in aggregate." A guardrail classifier, a reranker and a scoring layer can each be low materiality alone and material together when they jointly decide what a customer-facing system says. Link their inventory entries to the system they serve.
| Component | What it outputs | Quantitative estimate? | Generative or agentic? | Our reading |
|---|---|---|---|---|
| Discriminative classifier as guardrail or router (intent, toxicity, PII detection) | A class label and a probability score | Yes, and the score gates what happens next | Neither | In inventory |
| LLM used as a classifier (for example Llama Guard) | A verdict emitted as generated text | Functionally yes | Generative architecture, classification use | Document the classification decision |
| Reranker | A relevance score per candidate passage | Yes | Neither | In inventory, materiality set by what the ranking feeds |
| Embedding model | A vector | Not directly; it is an intermediate representation | Neither in function | Document the classification decision |
| Retrieval thresholds and business rules | Include or exclude decisions | No | Neither | Not a model: deterministic rule, governed as software |
| Confidence or risk-scoring layer on LLM output | A score or a band that routes to a human or an action | Yes | Neither | In inventory |
| Generator (the LLM) | Free text | No, in the definition's sense | Generative | Out of MRM scope, governed under your own risk practices |
| Agent loop (planner plus tool calls) | Actions and tool invocations | No | Agentic | Out of MRM scope, governed under your own risk practices |
| Prompt templates | Instructions to the generator | No | Configuration of a generative model | Out of scope as a model; version-controlled as part of the generator's change log |
Out of scope is not uncontrolled
Footnote 3's middle sentence is the one to put on the slide: "Nonetheless, a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document."
Then read Section I and its footnote together. Section I says the guidance "does not set forth enforceable standards or prescriptive requirements; accordingly, non-compliance with this guidance will not result in supervisory criticism against a banking organization." Footnote 1 adds: "However, supervisory action may result for any violations of law or unsafe or unsound practices stemming from insufficient management of model risk." The FDIC letter's wording is narrower still: "non-compliance with this guidance alone will not result in supervisory criticism." Out of the guidance's scope is not out of the safety and soundness standard.
So the generator and the agent loop need an evidence set, and a defensible one reuses the vocabulary the guidance keeps for in-scope models: conceptual soundness, outcomes analysis, ongoing monitoring, documentation and effective challenge. Your validators already know how to review against those words. Seven items cover it.
None of this requires a model inventory entry for the generator. It requires that when someone asks how the bank knows the system is fit for purpose, the answer is a folder, not a conversation.
Vendor models, third-party risk and why self-hosting changes validation
Section VII keeps the vendor principle intact. Vendor products, "including data, parameter values, or complete models", can present unique validation challenges, and because some components may be proprietary, "banking organizations may not receive from the vendor the underlying code, data, or methodology that they would have if a model were developed internally. Nevertheless, the principles of model risk management remain applicable." It goes on: sound practice involves "conducting ongoing monitoring and outcome analysis to assess whether vendor models are accurate, remain fit for purpose, and continue to be reliable." For customized vendor models, sound practice also involves documenting, justifying and evaluating the adjustments as part of model validation.
That splits one LLM system across different governing texts.
For the hosted generator, the operative third-party text is still the 2023 interagency guidance on third-party relationships. On September 11, 2026 the OCC, with the Fed, the FDIC and the NCUA, proposed guidance to revise and replace it (OCC Bulletin 2026-46). It was published in the Federal Register on September 15, with comments due November 16, 2026. The proposal bulletin does not mention AI or models. Until a final version lands, plan against the 2023 text.
This is where running the stack on your own infrastructure changes the answer. For the in-scope components (classifiers, rerankers, scoring layers) on the bank's own hardware, the bank holds the weights, pins them by hash, runs outcomes analysis on its own labelled data and decides when a version changes. That is the evidence conceptual soundness and ongoing monitoring ask for, and it removes most of the "may not receive" gap Section VII describes. For the generator, self-hosting keeps the eval results, change log and review samples inside the perimeter and removes the silent model update, where a hosted endpoint changes behavior between two of your eval runs without a version event on your side.
Be honest about the limit. Open weights give you the model and your own test evidence. They are not the same as open training data, so unless a release includes its pretraining data, a foundation model's conceptual soundness rests on documented evaluation and benchmarking rather than on development data in the Section VII sense. The guidance's own allowance for benchmarking, written for in-scope models, is a reasonable precedent to borrow.
| Component type | Governing text | What it means in practice |
|---|---|---|
| In-scope vendor component (hosted classifier, hosted reranker, vendor risk score) | SR 26-2 Section VII | Understand conceptual soundness, design, development data and performance as far as the vendor allows; run your own outcome analysis |
| Out-of-scope hosted generator | Third-party risk management plus your own governance practices | Vendor due diligence, contract terms on model changes, your own eval evidence |
| Out-of-scope self-hosted generator | Your own governance practices | The evidence checklist above, with the weights and the logs in your perimeter |
Who SR 26-2 applies to: the $30 billion line
The OCC's note for community banks says the guidance "is applicable to all community banks, subject to the limitations discussed in the guidance." Its background also cites OCC Bulletin 2025-26, which said the OCC's model risk guidance "does not, and should not be interpreted to, require community banks to perform annual model validation." The FDIC letter applies to all FDIC-supervised institutions and says the guidance generally does not apply to banks at $30 billion or less "to the extent such institutions do not have significant exposure to model risk."
Read together, the smaller-bank position is a general exclusion with an exposure test attached, not a blanket exemption, and the OCC has separately said annual validation is not required of community banks. A community bank that deploys a self-hosted LLM with two classifiers and a reranker has added models, and whether they amount to significant model risk exposure is a judgment the bank should document rather than assume. The guidance also states no inventory requirement; Section VI describes maintaining a comprehensive set of information on models as "common industry practice". It is still the cheapest way to show an examiner that the classification work was done.
| Total assets | What the guidance says | What to assume |
|---|---|---|
| Over $30 billion | Expected to be most relevant | The scope analysis above applies in full |
| $30 billion or less | Generally excluded; models at such banks "typically are subject to internal risk management and governance practices appropriate for the size and risk profile of these banking organizations" | Your own practices govern, sized to your risk |
| $30 billion or less with significant model risk exposure | "May be relevant" where exposure comes "because of the prevalence and complexity of their models or because of activities outside the scope of traditional community banking" | A bank running a multi-model AI stack should treat itself as in this row until it has argued otherwise |
Before the promised RFI: what to do this quarter
The OCC bulletin says "the agencies plan to issue in the near future a request for information that addresses model risk management generally and considers, in particular, banks' use of AI, including generative AI and agentic AI and AI-based models." That request had not been published as of late September 2026. Whether it proposes a separate framework for generative and agentic AI or folds them back into model risk management is unknown, and anyone describing its contents is guessing.
That uncertainty argues for doing the classification work now rather than waiting. When the RFI arrives, banks that already have a component-level map of their LLM systems can answer it from evidence. Banks that do not will be answering from memory.
The work is small enough to start this week:
This is the scoping work we do when we build self-hosted LLM systems for regulated teams: the model risk and validation staff you already have make the calls, and the system is built so each component arrives with the weights pinned, the eval evidence attached and the change log running from day one. More on running AI inside regulated businesses is in our AI for business guides.
FAQ
Quick answers to the questions this post tends to raise.



