Navers lab
← All tasks
lang06 Language rewrite Task 6 / 20

C++ → C#

jsonnet 0.20.0 C++11 C# / .NET 8

The implementation language moves, the artifact does not. 39,842 lines, 20 hours per model, and 2,608 checks across 19 modules.

Runs
26

one per configuration

Past rung 1
10

of 26 — the migration happened

Past rung 2
1

of those 10 — every check perfect

Past rung 3
0

of the 1 — accepted

Best composite
60

highest score on this task

The instrument

Acceptance gates

9

decide whether the migration happened

Behavioral modules

19

ported suites, run against the new tree

Frozen checks

2,608

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

14
Never migrated
9
Migrated, behavior lost
2
Perfect suite, gate rejected
1
Verifier found a difference
0
Accepted

3 runs were perfect on all 2,608 checks; 2 were rejected anyway — the migration had not happened. Of the 1 that reached rung three, 0 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
claude-opus-5 5 3 3
12.0 60
claude-sonnet-5 5 2 0
0.0
dsv4-flash 1 1 0
0.0
glm-5.2 1 0 0
0.0
gpt-5.6-luna 6 2 0
0.0
gpt-5.6-sol 6 1 0
0.0
kimi-k3 1 0 0
0.0
qwen3.8-max 1 1 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 xhigh broken 9/9 100.0% 2/6 60
claude-opus-5 high blind 8/9 100.0% 0
claude-opus-5 max blind 6/9 100.0% 0
kimi-k3 max failed 7/9 100.0% 0
claude-opus-5 medium partial 9/9 99.9% 0
claude-opus-5 low partial 9/9 99.8% 0
qwen3.8-max max partial 9/9 99.7% 0
claude-sonnet-5 max failed 8/9 97.1% 0
claude-sonnet-5 xhigh failed 7/9 97.0% 0
dsv4-flash max partial 9/9 96.6% 0
claude-sonnet-5 medium failed 7/9 96.1% 0
glm-5.2 max failed 8/9 95.9% 0
claude-sonnet-5 high partial 9/9 95.7% 0
claude-sonnet-5 low partial 9/9 95.2% 0
gpt-5.6-luna max partial 9/9 85.7% 0
gpt-5.6-luna xhigh failed 8/9 61.7% 0
gpt-5.6-luna high partial 9/9 36.5% 0
gpt-5.6-luna medium failed 8/9 27.1% 0
gpt-5.6-luna low failed 7/9 7.3% 0
gpt-5.6-sol high failed 7/9 6.7% 0
gpt-5.6-sol medium failed 7/9 6.2% 0
gpt-5.6-sol low failed 6/9 6.1% 0
gpt-5.6-sol none failed 7/9 6.0% 0
gpt-5.6-luna none failed 7/9 4.1% 0
gpt-5.6-sol xhigh failed 8/9 0.0% 0
gpt-5.6-sol max partial 9/9 0.0% 0