We Graded Our AI Platform Against the Central Banks' Own Rulebook. Here's the Report Card.

We Graded Our AI Platform Against the Central Banks' Own Rulebook. Here's the Report Card.
What a BIS governance framework for central banks tells us about where AI tooling is strong, and where the whole industry still has work to do.
If you work in finance, you've watched AI go from a science-fair curiosity to something that now sits inside real decisions: forecasting, fraud detection, document review, client service. And you've probably felt the unease that comes with it. How do we actually know this stuff is under control?
Regulators are asking the same question. The EU AI Act is phasing in, the NIST AI Risk Management Framework is becoming a reference standard, and boards have moved past "does it work?" to something harder: "can you prove it's governed?"
In January 2025, a very credible group weighed in. The Bank for International Settlements, which is basically the central bank for central banks, published Governance of AI adoption in central banks. It was written by a task force drawn from the central banks of Brazil, Canada, Chile, Colombia, Mexico, Peru, and the United States. It's one of the clearest, most practical governance yardsticks the financial sector has right now. So we did something a little nerve-wracking. We used it to grade our own AI platform, honestly, and we're sharing the report card.
What the BIS actually asks for
Strip away the jargon and the report comes down to one idea. Don't invent a brand-new rulebook for AI. Extend the disciplined risk management you already run. It lays out ten practical actions, everything from keeping an inventory of your AI tools, to assessing risk before deployment, to monitoring systems in production and reporting incidents. It leans on the familiar three lines of defence model. And it spells out the qualities trustworthy AI needs to have: secure, private, explainable, reliable, ethical, accountable, and subject to human oversight.
Here's the part that turned out to matter most. The BIS splits governance into two halves that track a model's life. There's the work you do before you deploy, like inventorying it, documenting it, assessing its risks, and approving it. And there's the work you do after you deploy, like monitoring it, catching incidents, and reviewing it over time. Hold onto that split, because it's the whole story.
The ten recommendations, scored
Here's the full report card at a glance. We graded each of the BIS's ten recommended actions as Exceeds, Meets, Partial, or Gap. One thing worth saying plainly first: a few of these are things an institution does, not something software can do for you. A platform can only supply the scaffolding. We've flagged those with an asterisk.
The ten recommendations, scored
Here's the full report card. We graded each of the BIS's ten recommended actions as Exceeds, Meets, Partial, or Gap. One caveat worth stating plainly: a few of these are things an institution does, not something software can do for you — a platform can only supply the scaffolding. Those are marked with an asterisk.
1. Establish an interdisciplinary AI committee — Partial.* The platform supplies the reviewer and approver roles a committee works through; actually standing up the committee is the institution's job.
2. Define principles for responsible AI use — Gap.* No in-product way to encode a responsible-AI policy and hold deployments to it.
3. Establish an AI framework and update guidance — Meets. Required documentation plus a staged approval flow act as a working assurance framework.
4. Maintain an AI tools inventory — Exceeds. Every model, language model, and tool is registered and discoverable in one place.
5. Map AI tools and stakeholders — Meets. Use cases and access rights are documented; a dedicated stakeholder map is lighter.
6. Assess risks and controls before go-live — Meets. Nothing reaches production without a completed model card and independent sign-off.
7. Perform regular monitoring (drift, bias, fairness) — Gap. Cost, usage, and errors are tracked; live model-quality monitoring is not yet built in.
8. Report anomalies and incidents — Partial. Run logs and an emergency stop exist; a structured incident workflow does not.
9. Develop and improve workforce skills — Gap.* Training and upskilling sit with the institution, not the platform.
10. Perform ongoing reviews and adaptations — Partial. New model versions are re-reviewed; time-based re-review of live models isn't enforced.
* Primarily an organizational responsibility — the platform can support it, but can't own it.
Read the table top to bottom and the pattern is hard to miss. The strong marks cluster in the pre-deployment rows, numbers 3 through 6. The gaps cluster in the post-deployment rows, numbers 7 through 10. That split is the thing to understand.
Where the tooling is already strong
Graded against that yardstick, the pre-deployment half is in good shape, and not just for us. This is broadly true of serious enterprise AI platforms. Put in terms a risk or compliance officer actually cares about:
A living inventory. Every model, language model, and tool is registered and discoverable. You can answer "what AI is running in our organization?", a question that sounds trivial right up until an auditor asks it and nobody can.
A documentation gate before anything goes live. Nothing reaches production without a completed "model card," which is a structured fact sheet covering the business case, the data, known biases, privacy exposure, and how to read the model's outputs. Think of it as the prospectus for a model.
An independent approval step. A compliance reviewer has to sign off before a model is published, with the power to reject and a documented reason for doing so. That's a real second line of defence, built into the software instead of bolted on in a spreadsheet.
Granular access control and a kill switch. Permissions are scoped tightly, and an unsafe tool can be pulled from live workflows on the spot. Every change gets logged.
A full audit trail. Each run captures its inputs, outputs, timing, and cost, plus a visual trace of exactly how an answer was produced. That's the traceability regulators increasingly expect.
For a finance audience, the headline is encouraging. The machinery for onboarding AI in a disciplined way, meaning inventory, documentation, approval, access control, and auditability, is real. In places it's ahead of the BIS baseline.
Where the gaps are, and why they matter
The honest part of the report card is the post-deployment half. This is where our platform, and frankly most of the market, has the furthest to go.
Live model-quality monitoring. Tracking cost, errors, and usage in production is straightforward. Continuously watching whether a model's quality is quietly degrading, whether its outputs are drifting, or whether bias is creeping in, is a lot harder, and it isn't a solved, built-in capability yet. This is the single biggest gap.
Incident management. When something goes wrong, the raw logs and an emergency stop are there. What isn't there yet is a structured "report it, investigate it, find the root cause, route it to an oversight committee" workflow.
Risk tiering. The BIS, and the EU AI Act, both want AI use cases sorted by risk level, so a high-stakes model faces tougher controls than a low-stakes one. Today the bar is the same for both.
Why should a finance professional care about these specifics? Because this is exactly where the regulation is heading. The EU AI Act's obligations for "high-risk" systems, and the NIST framework's push to measure and manage risk continuously, are all about the post-deployment life of a model, not the launch. A model that was perfectly safe on the day it shipped can drift into being wrong, unfair, or non-compliant six months later, and nobody would necessarily notice. Monitoring is the control that catches that. Its absence is the industry's collective blind spot.
The takeaway for finance leaders
Two things stand out. First, "AI governance" isn't one thing. It's a lifecycle, and it's worth pushing your vendors and your own teams to prove they cover both halves, not just the reassuring pre-launch checklist. Second, the good news: the hard foundational work is largely done. The platforms already capturing every model, every decision, and every input are sitting on exactly the raw material that live monitoring needs. Closing the gap is an evolution, not a rebuild.
If you're sizing up an AI platform, or your own AI program, here's a simple pocket test. Can it tell me what AI we run? Can it prove each one was reviewed before launch? Can it show me why a given decision was made? And can it alert me when a model starts to go wrong? The first three are fast becoming table stakes. The fourth is where the leaders will pull away from the pack.
This piece is based on an internal scorecard mapping our platform's capabilities against the BIS Consultative Group on Risk Management's Governance of AI adoption in central banks (January 2025). It's an honest read, strengths and gaps alike, because in governance the gaps are the part worth talking about.