Navers lab
← All tasks
lang01 Language rewrite Task 1 / 20

C → Rust

cmark 0.31.1 C11 Rust 1.90

The implementation language moves, the artifact does not. 22,619 lines, 30 hours per model, and 4,184 checks across 15 modules.

Runs
26

one per configuration

Past rung 1
6

of 26 — the migration happened

Past rung 2
0

of those 6 — every check perfect

Past rung 3
0

no run got this far

Best composite

nothing scored here

The instrument

Acceptance gates

8

decide whether the migration happened

Behavioral modules

15

ported suites, run against the new tree

Frozen checks

4,184

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

15
Never migrated
6
Migrated, behavior lost
5
Perfect suite, gate rejected
0
Verifier found a difference
0
Accepted

5 runs were perfect on all 4,184 checks and rejected anyway — the migration had not happened. No run reached rung three, so no verifier was paid to attack this task.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
claude-opus-5 5 0 4
0.0
claude-sonnet-5 5 0 1
0.0
dsv4-flash 1 0 0
0.0
glm-5.2 1 0 0
0.0
gpt-5.6-luna 6 3 0
0.0
gpt-5.6-sol 6 3 0
0.0
kimi-k3 1 0 0
0.0
qwen3.8-max 1 0 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 medium blind 7/8 100.0% 0
claude-opus-5 high blind 7/8 100.0% 0
claude-opus-5 xhigh blind 7/8 100.0% 0
claude-opus-5 max blind 7/8 100.0% 0
claude-sonnet-5 high blind 7/8 100.0% 0
claude-opus-5 low failed 7/8 100.0% 0
claude-sonnet-5 max failed 7/8 100.0% 0
kimi-k3 max failed 7/8 100.0% 0
claude-sonnet-5 xhigh failed 7/8 100.0% 0
dsv4-flash max failed 7/8 99.9% 0
gpt-5.6-sol max partial 8/8 99.9% 0
gpt-5.6-sol medium failed 1/8 99.6% 0
gpt-5.6-luna xhigh partial 8/8 94.6% 0
gpt-5.6-sol low partial 8/8 59.1% 0
gpt-5.6-luna max partial 8/8 1.3% 0
gpt-5.6-luna none failed 7/8 1.2% 0
gpt-5.6-luna high failed 7/8 1.1% 0
gpt-5.6-luna low failed 6/8 1.1% 0
gpt-5.6-luna medium partial 8/8 1.1% 0
gpt-5.6-sol high failed 6/8 1.1% 0
gpt-5.6-sol none failed 4/8 0.5% 0
claude-sonnet-5 low failed 7/8 0.3% 0
gpt-5.6-sol xhigh partial 8/8 0.3% 0
qwen3.8-max max failed 7/8 0.3% 0
glm-5.2 max failed 7/8 0.3% 0
claude-sonnet-5 medium failed 7/8 0.2% 0