Guidance, plainly
Fifteen years on, the Federal Reserve's model-risk guidance still decides how banks buy, build, and defend models. Its quiet center of gravity is evidence: documentation and records strong enough to support effective challenge. Here's what that means, and where most evidence falls short.
Three pillars, one recurring demand: records a stranger can trust.
SR 11-7 — the Federal Reserve's 2011 supervisory guidance on model risk management, issued jointly with the OCC — organizes model risk into three pillars: sound development, implementation, and use; effective validation; and governance, policies, and controls. Every pillar leans on the same substrate: documentation. The guidance expects records detailed enough that a knowledgeable third party, not involved in building the model, can understand how it operates, what it assumes, and where it breaks.
Validation, in the guidance's terms, has three parts: evaluation of conceptual soundness, ongoing monitoring — including benchmarking and process verification while the model runs — and outcomes analysis, comparing what the model said to what actually happened, backtesting being the canonical form. The first part is about how the model was built. The second and third are about what the model actually did in production — and that is where evidence quality is usually weakest.
The guidance's most-quoted phrase is an evidentiary standard in disguise.
"Effective challenge" — critical analysis by objective, informed parties with the incentives, competence, and influence to change the model — is the phrase everyone cites. Less noticed: challenge is only as strong as what the challenger can examine. A validator handed a summary spreadsheet of production decisions is not challenging the model; they are challenging the spreadsheet, and the spreadsheet was made by the people being challenged.
When the first line produces both the decisions and the record of the decisions, ongoing monitoring rests on the assumption that the record is complete and unedited. Internal audit is asked to verify processes it cannot independently re-run.
A record that proves its own integrity: sealed at decision time, anchored by a timestamp authority outside the bank and outside the vendor, reproducible by anyone with the receipt. The validator no longer has to trust the first line's bookkeeping — they can check it.
This is the property regulators reward across domains: evidence whose credibility does not depend on the credibility of the party producing it.
Where a sealed decision record slots into each expectation.
Every Corobate receipt names its inputs, their provenance grades, the rule configuration, and the engine version — so the question "what was the model actually doing that quarter" has a sealed answer, not an archaeology project across deploy logs.
Receipts accumulate into a production record that includes the abstentions — the decisions the system declined to make when evidence was missing or stale. A monitoring program that can see honest refusals is measuring the model's real operating envelope, not just its confident moments.
Backtesting against sealed receipts means the "what did we predict" side of the comparison cannot drift after the fact. Our backtests page walks the mechanics: last year's decisions become a scoreboard nobody can quietly rewrite.
How sealed receipts turn outcomes analysis into a scoreboard that improves next year's decisions.
Provenance classes, gated confidence, honest abstention, the seal, the anchor — with a live demo.
The European record-keeping duty asks for the same underlying property: records an outsider can believe.
Attestation receipt, effective challenge's evidentiary cousin terms, and the rest — defined in a paragraph each.