Navers lab
← All tasks
pf02 Platform port Task 16 / 20

CommonJS → V8 realm

Stylus 0.63.0 CommonJS + Node ESM in a V8 realm

The host the code assumes changes. 16,012 lines, 10 hours per model, and 2,570 checks across 12 modules.

Runs
26

one per configuration

Past rung 1
12

of 26 — the migration happened

Past rung 2
0

of those 12 — every check perfect

Past rung 3
0

no run got this far

Best composite

nothing scored here

The instrument

Acceptance gates

7

decide whether the migration happened

Behavioral modules

12

ported suites, run against the new tree

Frozen checks

2,570

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

14
Never migrated
12
Migrated, behavior lost
0
Perfect suite, gate rejected
0
Verifier found a difference
0
Accepted

No run was perfect on all 2,570 checks. No run reached rung three, so no verifier was paid to attack this task.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
claude-opus-5 5 5 0
0.0
claude-sonnet-5 5 4 0
0.0
dsv4-flash 1 1 0
0.0
glm-5.2 1 0 0
0.0
gpt-5.6-luna 6 2 0
0.0
gpt-5.6-sol 6 0 0
0.0
kimi-k3 1 0 0
0.0
qwen3.8-max 1 0 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
gpt-5.6-luna none failed 3/7 99.9% 0
qwen3.8-max max failed 3/7 99.9% 0
claude-opus-5 xhigh partial 7/7 57.5% 0
claude-opus-5 medium partial 7/7 57.5% 0
claude-opus-5 high partial 7/7 57.5% 0
claude-opus-5 max partial 7/7 57.5% 0
glm-5.2 max failed 6/7 57.4% 0
gpt-5.6-sol max failed 6/7 57.4% 0
dsv4-flash max partial 7/7 57.3% 0
claude-opus-5 low partial 7/7 57.2% 0
claude-sonnet-5 xhigh partial 7/7 57.2% 0
gpt-5.6-luna max partial 7/7 56.4% 0
gpt-5.6-sol xhigh failed 5/7 56.3% 0
gpt-5.6-luna xhigh partial 7/7 55.4% 0
claude-sonnet-5 high partial 7/7 54.0% 0
claude-sonnet-5 low failed 6/7 53.4% 0
gpt-5.6-luna low failed 4/7 50.7% 0
gpt-5.6-luna high failed 6/7 50.4% 0
gpt-5.6-luna medium failed 6/7 50.2% 0
gpt-5.6-sol medium failed 5/7 50.0% 0
gpt-5.6-sol none failed 6/7 50.0% 0
gpt-5.6-sol high failed 6/7 49.8% 0
gpt-5.6-sol low failed 4/7 41.4% 0
claude-sonnet-5 max partial 7/7 40.2% 0
kimi-k3 max failed 5/7 36.0% 0
claude-sonnet-5 medium partial 7/7 33.6% 0