Navers lab
← All tasks
build02 Build-toolchain rewrite Task 19 / 20

Maven → Gradle

Gson 2.10.1 Maven Gradle (offline)

What produces the artifact changes, and the package is the observable. 19,015 lines, 6 hours per model, and 2,669 checks across 10 modules.

Runs
26

one per configuration

Past rung 1
20

of 26 — the migration happened

Past rung 2
13

of those 20 — every check perfect

Past rung 3
0

of those 13 — accepted

Best composite
70

highest score on this task

The instrument

Acceptance gates

6

decide whether the migration happened

Behavioral modules

10

ported suites, run against the new tree

Frozen checks

2,669

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

6
Never migrated
7
Migrated, behavior lost
0
Perfect suite, gate rejected
13
Verifier found a difference
0
Accepted

13 runs were perfect on all 2,669 checks, and all passed the gate. Of the 13 that reached rung three, 0 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
gpt-5.6-sol 6 4 4
35.0 70
glm-5.2 1 1 1
60.0 60
claude-opus-5 5 5 4
42.0 60
claude-sonnet-5 5 5 2
20.0 60
kimi-k3 1 1 1
40.0 40
gpt-5.6-luna 6 4 1
6.7 40
dsv4-flash 1 0 0
0.0
qwen3.8-max 1 0 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
gpt-5.6-sol high broken 6/6 100.0% 3/6 70
claude-opus-5 xhigh broken 6/6 100.0% 2/6 60
claude-sonnet-5 medium broken 6/6 100.0% 2/6 60
glm-5.2 max broken 6/6 100.0% 2/6 60
gpt-5.6-sol xhigh broken 6/6 100.0% 2/6 60
claude-opus-5 low broken 6/6 100.0% 1/6 50
claude-opus-5 medium broken 6/6 100.0% 1/6 50
claude-opus-5 high broken 6/6 100.0% 1/6 50
claude-sonnet-5 high broken 6/6 100.0% 0/6 40
gpt-5.6-luna xhigh broken 6/6 100.0% 0/6 40
gpt-5.6-sol medium broken 6/6 100.0% 0/6 40
gpt-5.6-sol max broken 6/6 100.0% 0/6 40
kimi-k3 max broken 6/6 100.0% 0/6 40
claude-opus-5 max partial 6/6 100.0% 0
gpt-5.6-luna max partial 6/6 100.0% 0
claude-sonnet-5 low partial 6/6 99.9% 0
claude-sonnet-5 max partial 6/6 99.9% 0
gpt-5.6-luna medium partial 6/6 99.9% 0
gpt-5.6-sol none failed 5/6 99.9% 0
gpt-5.6-luna high failed 5/6 99.9% 0
claude-sonnet-5 xhigh partial 6/6 99.9% 0
gpt-5.6-luna low failed 5/6 99.4% 0
gpt-5.6-luna none partial 6/6 98.2% 0
dsv4-flash max failed 5/6 95.4% 0
gpt-5.6-sol low failed 5/6 93.7% 0
qwen3.8-max max failed 5/6 0.4% 0