Navers lab
← All tasks
fw06 Framework rewrite Task 13 / 20

gorilla/mux → net/http

go-simple-upload-server gorilla/mux net/http ServeMux

The language stays, the framework it is organized around is replaced. 803 lines, 6 hours per model, and 485 checks across 9 modules.

Runs
26

one per configuration

Past rung 1
13

of 26 — the migration happened

Past rung 2
13

of those 13 — every check perfect

Past rung 3
9

of those 13 — accepted

Best composite
100

highest score on this task

The instrument

Acceptance gates

6

decide whether the migration happened

Behavioral modules

9

ported suites, run against the new tree

Frozen checks

485

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

4
Never migrated
0
Migrated, behavior lost
9
Perfect suite, gate rejected
4
Verifier found a difference
9
Accepted

22 runs were perfect on all 485 checks; 9 were rejected anyway — the migration had not happened. Of the 13 that reached rung three, 9 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
claude-opus-5 5 5 5
100.0 100
kimi-k3 1 1 1
100.0 100
qwen3.8-max 1 1 1
100.0 100
gpt-5.6-sol 6 3 5
33.3 100
claude-sonnet-5 5 1 5
20.0 100
gpt-5.6-luna 6 2 3
20.0 70
dsv4-flash 1 0 1
0.0
glm-5.2 1 0 1
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 low accepted 6/6 100.0% 6/6 100
claude-opus-5 medium accepted 6/6 100.0% 6/6 100
claude-opus-5 high accepted 6/6 100.0% 6/6 100
claude-opus-5 xhigh accepted 6/6 100.0% 6/6 100
claude-opus-5 max accepted 6/6 100.0% 6/6 100
claude-sonnet-5 medium accepted 6/6 100.0% 6/6 100
gpt-5.6-sol max accepted 6/6 100.0% 6/6 100
kimi-k3 max accepted 6/6 100.0% 6/6 100
qwen3.8-max max accepted 6/6 100.0% 6/6 100
gpt-5.6-luna xhigh broken 6/6 100.0% 3/6 70
gpt-5.6-sol low broken 6/6 100.0% 2/6 60
gpt-5.6-luna max broken 6/6 100.0% 1/6 50
gpt-5.6-sol high broken 6/6 100.0% 0/6 40
claude-sonnet-5 low blind 5/6 100.0% 0
claude-sonnet-5 high blind 5/6 100.0% 0
claude-sonnet-5 xhigh blind 5/6 100.0% 0
claude-sonnet-5 max blind 5/6 100.0% 0
dsv4-flash max blind 5/6 100.0% 0
glm-5.2 max blind 5/6 100.0% 0
gpt-5.6-luna high blind 5/6 100.0% 0
gpt-5.6-sol medium blind 5/6 100.0% 0
gpt-5.6-sol xhigh blind 5/6 100.0% 0
gpt-5.6-luna medium failed 5/6 98.4% 0
gpt-5.6-luna low failed 5/6 97.5% 0
gpt-5.6-sol none failed 5/6 92.8% 0
gpt-5.6-luna none failed 4/6 92.4% 0