Navers lab
← All tasks
lang02 Language rewrite Task 2 / 20

C → Java

zlib 1.3.1 C89 Java 17

The implementation language moves, the artifact does not. 22,732 lines, 20 hours per model, and 4,162 checks across 18 modules.

Runs
26

one per configuration

Past rung 1
18

of 26 — the migration happened

Past rung 2
4

of those 18 — every check perfect

Past rung 3
0

of those 4 — accepted

Best composite
90

highest score on this task

The instrument

Acceptance gates

10

decide whether the migration happened

Behavioral modules

18

ported suites, run against the new tree

Frozen checks

4,162

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

8
Never migrated
14
Migrated, behavior lost
0
Perfect suite, gate rejected
4
Verifier found a difference
0
Accepted

4 runs were perfect on all 4,162 checks, and all passed the gate. Of the 4 that reached rung three, 0 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
kimi-k3 1 1 1
90.0 90
claude-opus-5 5 5 3
48.0 90
claude-sonnet-5 5 5 0
0.0
dsv4-flash 1 1 0
0.0
glm-5.2 1 0 0
0.0
gpt-5.6-luna 6 4 0
0.0
gpt-5.6-sol 6 1 0
0.0
qwen3.8-max 1 1 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 xhigh broken 10/10 100.0% 5/6 90
kimi-k3 max broken 10/10 100.0% 5/6 90
claude-opus-5 medium broken 10/10 100.0% 4/6 80
claude-opus-5 high broken 10/10 100.0% 3/6 70
claude-sonnet-5 max partial 10/10 100.0% 0
claude-sonnet-5 xhigh partial 10/10 99.9% 0
claude-sonnet-5 low partial 10/10 99.7% 0
gpt-5.6-luna max partial 10/10 99.5% 0
qwen3.8-max max partial 10/10 99.3% 0
claude-opus-5 low partial 10/10 99.1% 0
claude-opus-5 max partial 10/10 98.2% 0
claude-sonnet-5 high partial 10/10 97.6% 0
dsv4-flash max partial 10/10 96.4% 0
glm-5.2 max failed 9/10 93.9% 0
gpt-5.6-sol low failed 9/10 67.8% 0
gpt-5.6-sol high partial 10/10 66.7% 0
gpt-5.6-sol max failed 8/10 42.7% 0
gpt-5.6-luna high partial 10/10 42.4% 0
gpt-5.6-sol xhigh failed 9/10 9.2% 0
gpt-5.6-luna medium partial 10/10 3.8% 0
gpt-5.6-luna none failed 8/10 3.4% 0
gpt-5.6-sol medium failed 8/10 2.7% 0
gpt-5.6-luna low failed 8/10 2.1% 0
claude-sonnet-5 medium partial 10/10 1.9% 0
gpt-5.6-sol none failed 8/10 1.7% 0
gpt-5.6-luna xhigh partial 10/10 0.5% 0