Navers lab
← All tasks
pf03 Platform port Task 17 / 20

x86-64 → tri-arch

QuickJS x86-64 host layout tri-arch, endian-clean

The host the code assumes changes. 89,429 lines, 6 hours per model, and 1,487 checks across 12 modules.

Runs
26

one per configuration

Past rung 1
18

of 26 — the migration happened

Past rung 2
16

of those 18 — every check perfect

Past rung 3
4

of those 16 — accepted

Best composite
100

highest score on this task

The instrument

Acceptance gates

6

decide whether the migration happened

Behavioral modules

12

ported suites, run against the new tree

Frozen checks

1,487

all must pass to open rung three

Verifiers

6

hostile programs, written against the spec

Where the 26 runs stopped

4
Never migrated
2
Migrated, behavior lost
4
Perfect suite, gate rejected
12
Verifier found a difference
4
Accepted

20 runs were perfect on all 1,487 checks; 4 were rejected anyway — the migration had not happened. Of the 16 that reached rung three, 4 survived all six verifiers.

By model

Model Runs Gate Ceil. Where they stopped Mean Best
kimi-k3 1 1 1
100.0 100
claude-opus-5 5 4 5
76.0 100
gpt-5.6-sol 6 4 4
38.3 100
claude-sonnet-5 5 5 5
72.0 80
gpt-5.6-luna 6 2 3
13.3 80
glm-5.2 1 1 1
70.0 70
dsv4-flash 1 1 1
60.0 60
qwen3.8-max 1 0 0
0.0

Every run

best composite first

One row per configuration. Shape is the run's tool sequence in 24 slices — pale is shell, dark is an edit — and each log is the full session as recorded.

ModelEffortOutcomeGateChecksVerif.ScoreShape
claude-opus-5 low accepted 6/6 100.0% 6/6 100
claude-opus-5 high accepted 6/6 100.0% 6/6 100
gpt-5.6-sol max accepted 6/6 100.0% 6/6 100
kimi-k3 max accepted 6/6 100.0% 6/6 100
claude-opus-5 xhigh broken 6/6 100.0% 5/6 90
claude-opus-5 max broken 6/6 100.0% 5/6 90
claude-sonnet-5 low broken 6/6 100.0% 4/6 80
claude-sonnet-5 high broken 6/6 100.0% 4/6 80
claude-sonnet-5 xhigh broken 6/6 100.0% 4/6 80
gpt-5.6-luna max broken 6/6 100.0% 4/6 80
gpt-5.6-sol high broken 6/6 100.0% 4/6 80
claude-sonnet-5 max broken 6/6 100.0% 3/6 70
glm-5.2 max broken 6/6 100.0% 3/6 70
dsv4-flash max broken 6/6 100.0% 2/6 60
claude-sonnet-5 medium broken 6/6 100.0% 1/6 50
gpt-5.6-sol xhigh broken 6/6 100.0% 1/6 50
claude-opus-5 medium blind 5/6 100.0% 0
gpt-5.6-luna high blind 5/6 100.0% 0
gpt-5.6-luna xhigh blind 5/6 100.0% 0
gpt-5.6-sol medium blind 5/6 100.0% 0
gpt-5.6-sol none partial 6/6 99.7% 0
qwen3.8-max max failed 5/6 99.7% 0
gpt-5.6-luna low failed 4/6 99.0% 0
gpt-5.6-sol low failed 5/6 99.0% 0
gpt-5.6-luna medium partial 6/6 98.9% 0
gpt-5.6-luna none failed 4/6 98.6% 0