C → Java
The implementation language moves, the artifact does not. 22,732 lines, 20 hours per model, and 4,162 checks across 18 modules.
- Runs
- 26
- Past rung 1
- 18
- Past rung 2
- 4
- Past rung 3
- 0
- Best composite
- 90
one per configuration
of 26 — the migration happened
of those 18 — every check perfect
of those 4 — accepted
highest score on this task
The instrument
Acceptance gates
10
decide whether the migration happened
Behavioral modules
18
ported suites, run against the new tree
Frozen checks
4,162
all must pass to open rung three
Verifiers
6
hostile programs, written against the spec
Where the 26 runs stopped
- 8
- Never migrated
- 14
- Migrated, behavior lost
- 0
- Perfect suite, gate rejected
- 4
- Verifier found a difference
- 0
- Accepted
4 runs were perfect on all 4,162 checks, and all passed the gate. Of the 4 that reached rung three, 0 survived all six verifiers.
By model
| Model | Runs | Gate | Ceil. | Where they stopped | Mean | Best |
|---|---|---|---|---|---|---|
| kimi-k3 | 1 | 1 | 1 | | 90.0 | 90 |
| claude-opus-5 | 5 | 5 | 3 | | 48.0 | 90 |
| claude-sonnet-5 | 5 | 5 | 0 | | 0.0 | — |
| dsv4-flash | 1 | 1 | 0 | | 0.0 | — |
| glm-5.2 | 1 | 0 | 0 | | 0.0 | — |
| gpt-5.6-luna | 6 | 4 | 0 | | 0.0 | — |
| gpt-5.6-sol | 6 | 1 | 0 | | 0.0 | — |
| qwen3.8-max | 1 | 1 | 0 | | 0.0 | — |
Every run
best composite firstOne row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.
| Model | Effort | Outcome | Gate | Checks | Verif. | Score | Shape |
|---|---|---|---|---|---|---|---|
| claude-opus-5 | xhigh | broken | 10/10 | 100.0% | 5/6 | 90 | |
| kimi-k3 | max | broken | 10/10 | 100.0% | 5/6 | 90 | |
| claude-opus-5 | medium | broken | 10/10 | 100.0% | 4/6 | 80 | |
| claude-opus-5 | high | broken | 10/10 | 100.0% | 3/6 | 70 | |
| claude-sonnet-5 | max | partial | 10/10 | 100.0% | — | 0 | |
| claude-sonnet-5 | xhigh | partial | 10/10 | 99.9% | — | 0 | |
| claude-sonnet-5 | low | partial | 10/10 | 99.7% | — | 0 | |
| gpt-5.6-luna | max | partial | 10/10 | 99.5% | — | 0 | |
| qwen3.8-max | max | partial | 10/10 | 99.3% | — | 0 | |
| claude-opus-5 | low | partial | 10/10 | 99.1% | — | 0 | |
| claude-opus-5 | max | partial | 10/10 | 98.2% | — | 0 | |
| claude-sonnet-5 | high | partial | 10/10 | 97.6% | — | 0 | |
| dsv4-flash | max | partial | 10/10 | 96.4% | — | 0 | |
| glm-5.2 | max | failed | 9/10 | 93.9% | — | 0 | |
| gpt-5.6-sol | low | failed | 9/10 | 67.8% | — | 0 | |
| gpt-5.6-sol | high | partial | 10/10 | 66.7% | — | 0 | |
| gpt-5.6-sol | max | failed | 8/10 | 42.7% | — | 0 | |
| gpt-5.6-luna | high | partial | 10/10 | 42.4% | — | 0 | |
| gpt-5.6-sol | xhigh | failed | 9/10 | 9.2% | — | 0 | |
| gpt-5.6-luna | medium | partial | 10/10 | 3.8% | — | 0 | |
| gpt-5.6-luna | none | failed | 8/10 | 3.4% | — | 0 | |
| gpt-5.6-sol | medium | failed | 8/10 | 2.7% | — | 0 | |
| gpt-5.6-luna | low | failed | 8/10 | 2.1% | — | 0 | |
| claude-sonnet-5 | medium | partial | 10/10 | 1.9% | — | 0 | |
| gpt-5.6-sol | none | failed | 8/10 | 1.7% | — | 0 | |
| gpt-5.6-luna | xhigh | partial | 10/10 | 0.5% | — | 0 |