Flask → Starlette
The language stays, the framework it is organized around is replaced. 2,455 lines, 9 hours per model, and 2,830 checks across 11 modules.
- Runs
- 26
- Past rung 1
- 24
- Past rung 2
- 0
- Past rung 3
- 0
- Best composite
- —
one per configuration
of 26 — the migration happened
of those 24 — every check perfect
no run got this far
nothing scored here
The instrument
Acceptance gates
5
decide whether the migration happened
Behavioral modules
11
ported suites, run against the new tree
Frozen checks
2,830
all must pass to open rung three
Verifiers
6
hostile programs, written against the spec
Where the 26 runs stopped
- 1
- Never migrated
- 24
- Migrated, behavior lost
- 1
- Perfect suite, gate rejected
- 0
- Verifier found a difference
- 0
- Accepted
1 run was perfect on all 2,830 checks and rejected anyway — the migration had not happened. No run reached rung three, so no verifier was paid to attack this task.
By model
| Model | Runs | Gate | Ceil. | Where they stopped | Mean | Best |
|---|---|---|---|---|---|---|
| claude-opus-5 | 5 | 5 | 0 | | 0.0 | — |
| claude-sonnet-5 | 5 | 5 | 0 | | 0.0 | — |
| dsv4-flash | 1 | 1 | 0 | | 0.0 | — |
| glm-5.2 | 1 | 1 | 0 | | 0.0 | — |
| gpt-5.6-luna | 6 | 5 | 0 | | 0.0 | — |
| gpt-5.6-sol | 6 | 6 | 0 | | 0.0 | — |
| kimi-k3 | 1 | 1 | 0 | | 0.0 | — |
| qwen3.8-max | 1 | 0 | 1 | | 0.0 | — |
Every run
best composite firstOne row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.
| Model | Effort | Outcome | Gate | Checks | Verif. | Score | Shape |
|---|---|---|---|---|---|---|---|
| qwen3.8-max | max | blind | 2/5 | 100.0% | — | 0 | |
| claude-opus-5 | high | partial | 5/5 | 99.9% | — | 0 | |
| claude-opus-5 | max | partial | 5/5 | 99.9% | — | 0 | |
| gpt-5.6-luna | max | partial | 5/5 | 99.8% | — | 0 | |
| kimi-k3 | max | partial | 5/5 | 99.7% | — | 0 | |
| claude-opus-5 | low | partial | 5/5 | 99.4% | — | 0 | |
| claude-opus-5 | medium | partial | 5/5 | 99.4% | — | 0 | |
| claude-opus-5 | xhigh | partial | 5/5 | 99.2% | — | 0 | |
| gpt-5.6-sol | max | partial | 5/5 | 99.2% | — | 0 | |
| dsv4-flash | max | partial | 5/5 | 99.1% | — | 0 | |
| glm-5.2 | max | partial | 5/5 | 99.1% | — | 0 | |
| gpt-5.6-sol | xhigh | partial | 5/5 | 98.8% | — | 0 | |
| gpt-5.6-sol | medium | partial | 5/5 | 98.5% | — | 0 | |
| gpt-5.6-sol | high | partial | 5/5 | 98.4% | — | 0 | |
| gpt-5.6-luna | xhigh | partial | 5/5 | 97.5% | — | 0 | |
| claude-sonnet-5 | xhigh | partial | 5/5 | 96.7% | — | 0 | |
| gpt-5.6-sol | low | partial | 5/5 | 94.9% | — | 0 | |
| gpt-5.6-sol | none | partial | 5/5 | 93.9% | — | 0 | |
| gpt-5.6-luna | high | partial | 5/5 | 93.3% | — | 0 | |
| claude-sonnet-5 | medium | partial | 5/5 | 90.9% | — | 0 | |
| claude-sonnet-5 | low | partial | 5/5 | 90.3% | — | 0 | |
| claude-sonnet-5 | high | partial | 5/5 | 87.2% | — | 0 | |
| claude-sonnet-5 | max | partial | 5/5 | 82.5% | — | 0 | |
| gpt-5.6-luna | medium | partial | 5/5 | 76.3% | — | 0 | |
| gpt-5.6-luna | low | failed | 4/5 | 68.2% | — | 0 | |
| gpt-5.6-luna | none | partial | 5/5 | 53.7% | — | 0 |