The sentence we published on Monday
“Recursive self-improvement is two problems” argued that the field’s missing rung is Level 3: the point where the method of improvement is itself learned from the system’s own outcomes, and the claim is testable — run the learned method and the fixed method side by side, forward in time, on the same budget, and let a statistic chosen in advance decide. CID was built to run exactly that loop: an outcome ledger the improver cannot edit, a System One advisor trained from pairs of past runs, a comparative gate on unseen holdout pairs, and a scoreboard that is built to say “not yet”. We wrote that we would rather publish the sentence “has not demonstrated” than a demo. This post exists because the next line of that story is now measurable.
Everything below happened on a live deployment — a kind cluster running the platform, a research GPU in our own rack doing the training — between September 28 and October 1, 2026. Job identifiers are printed because they are the receipts: any of them can be looked up in the ledger.
The gate’s first four verdicts
The advisor, cid-recipe-advisor, is a System One decision model trained on pairs of past Model Studio runs: given two configurations, it predicts which reaches the better validation outcome. A new candidate is trained automatically whenever enough new comparable pairs exist, and it goes to work only if it clears the comparative gate — accuracy and calibration on holdout pairs excluded from its training, against the advisor it would replace, with the statistic fixed in advance. The incumbent it first had to beat was not a learned model at all but the uniform prior: pick a configuration at random.
Read the rejections first, because they are the design working. Cycle 2 beat the incumbent on accuracy by twenty points and still failed: its calibration error, 0.10017, missed the fixed cap of 0.100 by 0.00017. Loosening the gate after seeing that number would have been exactly the post-hoc credit the accounting forbids, so the gate stayed as written and the candidate was rejected. Cycle 5 lost to the promoted cycle-3 advisor and the incumbent kept its slot. Two further candidates died in training and were counted as losses. Nothing in this list was re-run, re-scored or re-interpreted to reach a nicer story.
The promotions are what the loop is for. Cycle 3’s win was not subtle: 0.819 accuracy against 0.373 for uniform selection on 83 unseen pairs, calibration error 0.057, and a loss-difference confidence interval of [-0.377, -0.157] that sits entirely below zero. It was promoted automatically — audited, and reversible with one rollback call. The twelve measured-search trials that fed it all ran clean, one runner attempt per phase. That is Level 3’s mechanism demonstrated live: the system learned, from its own outcomes, which configurations tend to win, and the controls agreed the learning was real.
Then the improver improved
The recursive step is not the first promotion. It is what happened at 20:04 PDT that same evening. The ledger had grown — more measured searches had completed, including trials the promoted advisor itself had ranked — and when the next candidate trained on the enlarged history, its gate comparison was no longer against the uniform prior but against the cycle-3 advisor, the one already serving searches. The new candidate scored 0.886 against the incumbent’s 0.818 on 88 unseen pairs, and was promoted in turn. The method that chooses configurations is now a model the system trained, from evidence the system gathered, to replace a model the system had earlier promoted the same way. A person wrote none of it.
The scale is modest and we say so: 195 recorded runs yielding 679 comparable outcome pairs, 99 held out, across 3 recipe groups; a candidate trained and gated roughly every time thirty new pairs accumulate; five gate verdicts in four days. This is not an intelligence explosion and we do not claim one — the loop can only compound what its models can learn, which is the research half of Monday’s essay and it is unchanged. What the numbers show is the engineering half, closed: a learned improver, promoted twice by a gate that also said no twice, now ranking live recipe configurations with its promotions audited and reversible.
The second experiment: fairness, fixed in writing first
Independent of the improvement loop, Tuesday ran a discipline test of a different kind. Bank Account Fraud Base — a public card-fraud dataset whose customers carry an age attribute — has a known property our baseline reproduced faithfully: at the 5% false-positive operating point, customers aged 50 and over are flagged at 0.1129 while customers under 50 are flagged at 0.0378, a ratio of 2.98. Before any training run was admitted, we wrote down a specification with two criteria and no room to negotiate with ourselves: on the held-back months 6–7 test, cut the age false-positive-rate ratio to 1.5 or below while keeping recall at 5% false positives at 0.53 or above. The constraint — a fairness penalty inside the training loop, merged as platform PR #403 and live in the same deployment — was then run at two strengths, and every number below was scored once, per the pre-registration.
Both criteria were met. The ratio fell from 2.98 to 1.42 — the over-50 group’s false-positive rate dropped from 0.1129 to 0.0665 while the under-50 group moved from 0.0378 to 0.0468 — and recall landed at 0.5490 with a 95% confidence interval of [0.5316, 0.5673], inside the 0.53 guardrail and within a point of the tree’s 0.5594. The winning strength was not chosen by looking at the test: neither arm reached 1.5 on the validation split, so the pre-registered tie-break — take the lowest validation ratio — selected the stronger constraint before the test file was touched.
| arm | validation TPR@5%FPR [95% CI] | validation FPR ≥50 | validation FPR <50 | validation ratio |
|---|---|---|---|---|
| fair256 (λ=256, job 3693bfb4) | 0.5641 [0.5224, 0.5996] | 0.1012 | 0.0384 | 2.64 |
| fair1024 (λ=1024, job 75e6d788), selected | 0.5494 [0.5104, 0.5853] | 0.0771 | 0.0438 | 1.76 |
The honest paragraph, because it matters. We also fitted a blend of the fairness arm with the reference tree on validation, and the fit put a negative weight on the fairness arm (−0.057 against the tree’s 1.043). The constraint moves scores toward group parity in a way that added no validation recall the tree did not already have, so the blend reverts to the tree and inherits its 3.10 ratio. Fairness and blend-stacking are in tension here, the blend is reported rather than shipped, and the arm stands alone as selected. The remaining 1.42 gap, and the fact that the constraint penalizes average score disparity rather than the operating-point ratio directly, were both written into the specification’s risk section before the run. A constraint strength between the two arms, or per-group reweighting, are the obvious follow-ups — each behind its own new pre-registration.
What we are not claiming
- Not “proven”, yet — “capable”, with one gate still armed. The strictest wording of Level 3 — paired, prospective evidence that advisor-ranked search beats plain search on the same budget — has its experiment running but not its verdict. Three paired searches of ten trials each ran with the promoted advisor ranking one side of every pair; the anytime-valid betting test needs at least 30 complete pairs and must cross its 0.025 threshold for a mean advantage above 0.01 before that claim closes. Until the scoreboard says so, we say capable. When it crosses — either way — we will publish the number.
- Not Level 4, and nowhere near Level 5. No system code, training procedure or architecture was modified by the loop. The promoted advisor changes which configurations get tried; it does not change the platform that tries them. Self-modification behind sandboxing and human review is a different rung with a different burden of proof, and the loop earns no credit toward it from these results.
- Not an acceleration forecast. Four gate verdicts on one fraud-recipe family, at desktop-GPU scale, is a capability demonstration, not a trend line. Whether learned improvement transfers across recipes, datasets and model families — and how capable the base model must be before the compounding matters — remains the open research question our Monday essay ended on. It is still open.
A gate that can say no is the only kind that can say yes. CID’s gate said no twice, yes twice, and the yes was earned against the thing it had already promoted.
Both experiments ran on the same platform build, on our own infrastructure, with the controls on by default — the same controls a customer deployment runs with. That is the point of the exercise: the loop that promoted its own improver is not a research demo you can rent by the page view, it is what Model Studio does when you let it learn from your own runs. The advisor is serving recipe searches on our deployment tonight; the fairness constraint is a recipe field; the pre-registration habit is the cheapest part of the whole thing and the one we would press on anyone attempting this themselves.
Try it on your own runs
None of this needs our data, our fraud recipes or our hardware — it needs yours. A working demo is one session, and it looks like this: bring one recipe you already run, anything with a measurable validation outcome — fraud screening, alert triage, intent routing, a text adapter you have hand-tuned and suspect is not finished — and we will stand up the outcome ledger against its history, launch a measured search under a budget you set, and let you watch the gate and the scoreboard decide on configurations the advisor ranked itself. If your runs are already logged, the first advisor trains the same afternoon. If they are not, building that history is step one, and it is the part you keep regardless of what the loop ever promotes.
- You set the budget and the bar. Trial limits, GPU minutes, the evaluation metrics, the fairness criteria if you have them — fixed in writing before the loop starts, exactly as this post describes.
- The gate is not a sales demo. It will reject candidates in front of you if they miss, and it will tell you “not yet” until the paired evidence crosses. Ask it to.
- Nothing leaves your infrastructure. The ledger, the holdouts, the promoted models and the rollback switch all live on your deployment; we bring the loop, not a hosted black box.
Book a demo and bring a recipe. The shortest path is the contact page — say what you train and what you would want the loop to improve; we will come prepared with your search space already sketched. If you would rather watch first, the advisor from this post is ranking live recipe searches on our own deployment right now, and we are glad to show that instead.