The finding
On September 30, Bank of England Governor Andrew Bailey set out where AI governance should begin: “A sensible starting point is rigorous model testing, conducted before and after deployment.” Regulation, in his view, comes later. Understanding, testing and points of intervention “must come first.”
He is right, and the “after deployment” part matters most. A model that passed once is not guaranteed to pass again.
The same day, the Bank published the record of its Financial Policy Committee’s September meeting. It discusses AI 41 times. It never uses the words independent, evaluation, validation, audit, assurance or benchmark. The risks it describes, including frontier models taking unexpected actions in test environments, are routed through cyber and operational resilience.
Testing is now a named starting point. Who does the testing, and what it has to prove, is in neither document.
The Governor said test. Nobody has said who.
The implication
For an operations team, “test before and after deployment” becomes three practical questions. Which cases? How many runs? Run by whom, on what configuration?
A vendor’s accuracy figure answers none of them. A single run on the vendor’s own cases tells you how the model did once, on cases the vendor chose. Our own work this year found models changing their answers between identical runs, and the same model scoring very differently depending on which host served it.
The labs have started defining the testers themselves. Their Frontier Model Forum describes third-party assessment and names two assessors, METR and Apollo Research. Both are serious organizations. Neither tests settlement, margin or collateral workflows.
The test that matters is the one on your workflow, repeated, with the conditions recorded.
The residual finding
Testing has a failure mode of its own: the test can tell the model it is a test.
This week we scanned the current version of all eight AAL datasets, 2,000 cases, for anything in the model-facing text that names the benchmark. We found two problems in our own work.
In AAL-D-002, one case of 250 had two sentences in its context naming the dataset, including that it was “the final case.” We fixed it in version 1.0.1, and the correct answer did not change. GPT-4o answered that case correctly on all three runs. Removing it moves GPT-4o’s amount accuracy from 76.08% to 75.88%, and no other figure by more than 0.2 points.
In AAL-D-008, every case identifier shown to the model includes the benchmark name. Both frontier models scored 100% there, so there is no headroom to tell whether it mattered. We disclosed it rather than guess.
The prompts for five other datasets carry no such cue. For one, AAL-D-005, we could not check: its run scripts are not in our current files, and its page now says so.
A test that announces itself is still measuring something. It may not be what you meant.
The control consideration
If you run AI evaluations internally, scan every field the model sees for anything that names the test: dataset names, case IDs, file paths, words like “benchmark” or “evaluation.” Show the model neutral identifiers and keep the real case ID on the scoring side. Then disclose what you find, including the parts that don’t change the answer.
Our scan script is public, in the AAL benchmark repository alongside the datasets and scorers.
If your test data can tell the model it’s being tested, fix that before you trust the score.
From AI Alpha Labs this week
The AAL benchmark repository is public: generators, deterministic scorers, drivers, datasets and results.
AAL-D-002 v1.0.1 is published, with the correction disclosed on its page.
The AAL-D-008 page now discloses the identifier issue.
Our Methodology page now describes current practice, including what we don’t do.
