Navers lab
← All tasks
fw04 Framework rewrite Task 11 / 20

Gin → chi

ChartMuseum Gin chi v5 (net/http)

The language stays, the framework it is organized around is replaced. 5,401 lines, 12 hours per model, and 11,429 checks across 14 modules.

Runs
26

one per configuration

Past rung 1
25

of 26 — the migration happened

Past rung 2
6

of those 25 — every check perfect

Past rung 3
3

of those 6 — accepted

Best composite
100

highest score on this task

The instrument

Acceptance gates

5

decide whether the migration happened

Behavioral modules

14

ported suites, run against the new tree

Frozen checks

11,429

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

0
Never migrated
19
Migrated, behavior lost
1
Perfect suite, gate rejected
3
Verifier found a difference
3
Accepted

7 runs were perfect on all 11,429 checks; 1 was rejected anyway — the migration had not happened. Of the 6 that reached rung three, 3 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
claude-opus-5 5 5 4
76.0 100
gpt-5.6-sol 6 6 1
16.7 100
dsv4-flash 1 1 1
80.0 80
claude-sonnet-5 5 5 0
0.0
glm-5.2 1 1 0
0.0
gpt-5.6-luna 6 6 0
0.0
kimi-k3 1 1 0
0.0
qwen3.8-max 1 0 1
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 high accepted 5/5 100.0% 6/6 100
claude-opus-5 xhigh accepted 5/5 100.0% 6/6 100
gpt-5.6-sol max accepted 5/5 100.0% 6/6 100
claude-opus-5 medium broken 5/5 100.0% 5/6 90
claude-opus-5 max broken 5/5 100.0% 5/6 90
dsv4-flash max broken 5/5 100.0% 4/6 80
qwen3.8-max max blind 2/5 100.0% 0
glm-5.2 max partial 5/5 99.9% 0
gpt-5.6-luna max partial 5/5 99.7% 0
gpt-5.6-sol xhigh partial 5/5 99.6% 0
claude-opus-5 low partial 5/5 99.6% 0
gpt-5.6-sol medium partial 5/5 99.6% 0
claude-sonnet-5 high partial 5/5 99.6% 0
claude-sonnet-5 max partial 5/5 99.6% 0
gpt-5.6-sol high partial 5/5 99.6% 0
kimi-k3 max partial 5/5 99.6% 0
gpt-5.6-luna xhigh partial 5/5 99.5% 0
gpt-5.6-luna high partial 5/5 99.5% 0
claude-sonnet-5 medium partial 5/5 99.4% 0
claude-sonnet-5 low partial 5/5 99.4% 0
gpt-5.6-luna medium partial 5/5 99.3% 0
gpt-5.6-luna none partial 5/5 99.2% 0
gpt-5.6-sol none partial 5/5 96.8% 0
gpt-5.6-luna low partial 5/5 96.0% 0
gpt-5.6-sol low partial 5/5 93.4% 0
claude-sonnet-5 xhigh partial 5/5 85.1% 0