Navers lab
← All tasks
lang03 Language rewrite Task 3 / 20

Python → Go

sqlparse 0.5.3 Python Go (stdlib)

The implementation language moves, the artifact does not. 4,393 lines, 15 hours per model, and 8,530 checks across 15 modules.

Runs
26

one per configuration

Past rung 1
14

of 26 — the migration happened

Past rung 2
0

of those 14 — every check perfect

Past rung 3
0

no run got this far

Best composite

nothing scored here

The instrument

Acceptance gates

8

decide whether the migration happened

Behavioral modules

15

ported suites, run against the new tree

Frozen checks

8,530

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

12
Never migrated
14
Migrated, behavior lost
0
Perfect suite, gate rejected
0
Verifier found a difference
0
Accepted

No run was perfect on all 8,530 checks. No run reached rung three, so no verifier was paid to attack this task.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
claude-opus-5 5 4 0
0.0
claude-sonnet-5 5 1 0
0.0
dsv4-flash 1 0 0
0.0
glm-5.2 1 0 0
0.0
gpt-5.6-luna 6 4 0
0.0
gpt-5.6-sol 6 4 0
0.0
kimi-k3 1 1 0
0.0
qwen3.8-max 1 0 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 xhigh partial 8/8 97.4% 0
claude-opus-5 low partial 8/8 96.9% 0
claude-opus-5 medium partial 8/8 96.1% 0
claude-opus-5 high partial 8/8 95.3% 0
claude-sonnet-5 high failed 6/8 94.3% 0
claude-opus-5 max failed 7/8 88.7% 0
gpt-5.6-sol max partial 8/8 85.8% 0
gpt-5.6-luna max partial 8/8 84.9% 0
claude-sonnet-5 max partial 8/8 84.9% 0
claude-sonnet-5 low failed 6/8 84.3% 0
qwen3.8-max max failed 7/8 83.0% 0
claude-sonnet-5 medium failed 7/8 82.6% 0
gpt-5.6-luna xhigh partial 8/8 79.5% 0
kimi-k3 max partial 8/8 78.0% 0
dsv4-flash max failed 7/8 76.9% 0
glm-5.2 max failed 7/8 76.0% 0
gpt-5.6-sol medium partial 8/8 74.3% 0
gpt-5.6-luna high partial 8/8 72.5% 0
gpt-5.6-sol low partial 8/8 66.5% 0
gpt-5.6-sol xhigh partial 8/8 55.5% 0
gpt-5.6-luna medium partial 8/8 53.1% 0
gpt-5.6-luna none failed 7/8 39.0% 0
gpt-5.6-sol high failed 7/8 31.1% 0
gpt-5.6-luna low failed 7/8 28.5% 0
gpt-5.6-sol none failed 7/8 25.2% 0
claude-sonnet-5 xhigh failed 2/8 3.4% 0