x86-64 → tri-arch
The host the code assumes changes. 89,429 lines, 6 hours per model, and 1,487 checks across 12 modules.
- Runs
- 26
- Past rung 1
- 18
- Past rung 2
- 16
- Past rung 3
- 4
- Best composite
- 100
one per configuration
of 26 — the migration happened
of those 18 — every check perfect
of those 16 — accepted
highest score on this task
The instrument
Acceptance gates
6
decide whether the migration happened
Behavioral modules
12
ported suites, run against the new tree
Frozen checks
1,487
all must pass to open rung three
Verifiers
6
hostile programs, written against the spec
Where the 26 runs stopped
- 4
- Never migrated
- 2
- Migrated, behavior lost
- 4
- Perfect suite, gate rejected
- 12
- Verifier found a difference
- 4
- Accepted
20 runs were perfect on all 1,487 checks; 4 were rejected anyway — the migration had not happened. Of the 16 that reached rung three, 4 survived all six verifiers.
By model
| Model | Runs | Gate | Ceil. | Where they stopped | Mean | Best |
|---|---|---|---|---|---|---|
| kimi-k3 | 1 | 1 | 1 | | 100.0 | 100 |
| claude-opus-5 | 5 | 4 | 5 | | 76.0 | 100 |
| gpt-5.6-sol | 6 | 4 | 4 | | 38.3 | 100 |
| claude-sonnet-5 | 5 | 5 | 5 | | 72.0 | 80 |
| gpt-5.6-luna | 6 | 2 | 3 | | 13.3 | 80 |
| glm-5.2 | 1 | 1 | 1 | | 70.0 | 70 |
| dsv4-flash | 1 | 1 | 1 | | 60.0 | 60 |
| qwen3.8-max | 1 | 0 | 0 | | 0.0 | — |
Every run
best composite firstOne row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.
| Model | Effort | Outcome | Gate | Checks | Verif. | Score | Shape |
|---|---|---|---|---|---|---|---|
| claude-opus-5 | low | accepted | 6/6 | 100.0% | 6/6 | 100 | |
| claude-opus-5 | high | accepted | 6/6 | 100.0% | 6/6 | 100 | |
| gpt-5.6-sol | max | accepted | 6/6 | 100.0% | 6/6 | 100 | |
| kimi-k3 | max | accepted | 6/6 | 100.0% | 6/6 | 100 | |
| claude-opus-5 | xhigh | broken | 6/6 | 100.0% | 5/6 | 90 | |
| claude-opus-5 | max | broken | 6/6 | 100.0% | 5/6 | 90 | |
| claude-sonnet-5 | low | broken | 6/6 | 100.0% | 4/6 | 80 | |
| claude-sonnet-5 | high | broken | 6/6 | 100.0% | 4/6 | 80 | |
| claude-sonnet-5 | xhigh | broken | 6/6 | 100.0% | 4/6 | 80 | |
| gpt-5.6-luna | max | broken | 6/6 | 100.0% | 4/6 | 80 | |
| gpt-5.6-sol | high | broken | 6/6 | 100.0% | 4/6 | 80 | |
| claude-sonnet-5 | max | broken | 6/6 | 100.0% | 3/6 | 70 | |
| glm-5.2 | max | broken | 6/6 | 100.0% | 3/6 | 70 | |
| dsv4-flash | max | broken | 6/6 | 100.0% | 2/6 | 60 | |
| claude-sonnet-5 | medium | broken | 6/6 | 100.0% | 1/6 | 50 | |
| gpt-5.6-sol | xhigh | broken | 6/6 | 100.0% | 1/6 | 50 | |
| claude-opus-5 | medium | blind | 5/6 | 100.0% | — | 0 | |
| gpt-5.6-luna | high | blind | 5/6 | 100.0% | — | 0 | |
| gpt-5.6-luna | xhigh | blind | 5/6 | 100.0% | — | 0 | |
| gpt-5.6-sol | medium | blind | 5/6 | 100.0% | — | 0 | |
| gpt-5.6-sol | none | partial | 6/6 | 99.7% | — | 0 | |
| qwen3.8-max | max | failed | 5/6 | 99.7% | — | 0 | |
| gpt-5.6-luna | low | failed | 4/6 | 99.0% | — | 0 | |
| gpt-5.6-sol | low | failed | 5/6 | 99.0% | — | 0 | |
| gpt-5.6-luna | medium | partial | 6/6 | 98.9% | — | 0 | |
| gpt-5.6-luna | none | failed | 4/6 | 98.6% | — | 0 |