Navers lab
← All tasks
lang07 Language rewrite Task 7 / 20

JavaScript → TypeScript

JSONata 2.0.6 JavaScript TypeScript 5.9

The implementation language moves, the artifact does not. 9,247 lines, 12 hours per model, and 13,977 checks across 15 modules.

Runs
26

one per configuration

Past rung 1
14

of 26 — the migration happened

Past rung 2
3

of those 14 — every check perfect

Past rung 3
3

of those 3 — accepted

Best composite
100

highest score on this task

The instrument

Acceptance gates

10

decide whether the migration happened

Behavioral modules

15

ported suites, run against the new tree

Frozen checks

13,977

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

10
Never migrated
11
Migrated, behavior lost
2
Perfect suite, gate rejected
0
Verifier found a difference
3
Accepted

5 runs were perfect on all 13,977 checks; 2 were rejected anyway — the migration had not happened. Of the 3 that reached rung three, 3 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
gpt-5.6-sol 6 2 3
33.3 100
claude-opus-5 5 5 1
20.0 100
claude-sonnet-5 5 4 0
0.0
dsv4-flash 1 1 0
0.0
glm-5.2 1 0 0
0.0
gpt-5.6-luna 6 0 1
0.0
kimi-k3 1 1 0
0.0
qwen3.8-max 1 1 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 xhigh accepted 10/10 100.0% 6/6 100
gpt-5.6-sol high accepted 10/10 100.0% 6/6 100
gpt-5.6-sol max accepted 10/10 100.0% 6/6 100
gpt-5.6-luna max blind 8/10 100.0% 0
gpt-5.6-sol medium blind 8/10 100.0% 0
gpt-5.6-sol xhigh failed 8/10 100.0% 0
dsv4-flash max partial 10/10 100.0% 0
gpt-5.6-luna high failed 8/10 100.0% 0
gpt-5.6-luna xhigh failed 8/10 100.0% 0
claude-sonnet-5 high partial 10/10 99.9% 0
glm-5.2 max failed 8/10 99.9% 0
gpt-5.6-luna medium failed 8/10 99.9% 0
claude-opus-5 max partial 10/10 99.9% 0
kimi-k3 max partial 10/10 99.9% 0
claude-sonnet-5 low partial 10/10 99.9% 0
qwen3.8-max max partial 10/10 99.9% 0
claude-sonnet-5 xhigh partial 10/10 99.9% 0
claude-opus-5 medium partial 10/10 99.9% 0
claude-sonnet-5 medium failed 8/10 99.9% 0
gpt-5.6-luna none failed 8/10 99.9% 0
gpt-5.6-luna low failed 7/10 99.9% 0
claude-opus-5 high partial 10/10 99.8% 0
claude-opus-5 low partial 10/10 99.8% 0
gpt-5.6-sol none failed 7/10 99.8% 0
claude-sonnet-5 max partial 10/10 90.8% 0
gpt-5.6-sol low failed 8/10 0.0% 0