Navers lab
← All tasks
JobShop Task 15 / 47

abz

Classical job-shop scheduling on the ABZ benchmark family (Adams, Balas, Zawack 1988): sequence operations on machines to minimize makespan or tardiness—strongly NP-hard combinatorial optimization. Scoring compares against published bounds/optima for relative gap, a core manufacturing scheduling challenge.

Model leaderboard

Updated v1 results · 2026-09-15 · raw score, higher is better.

# Participant Raw score Medal
1 Claude Opus 4.6 96.1035 Gold
2 GPT-5.4 91.231431 Silver
3 GLM-5 88.4924 Bronze
4 DeepSeek V3.2 88.3614 —
5 Grok 4.20 87.6717 —
6 Gemini 3.1 Pro Preview 86.751 —
7 Seed 2.0 Pro 86.672 —
8 Qwen3 Coder Next 85.603 —

Framework results · original paper

Historical results from the original evaluators · normalized score (0–100).

# Participant Score
1 Claude Opus 4.6 + ABMCTS 100.0
2 Claude Opus 4.6 + ShinkaiEvolve 90.0
3 Claude Opus 4.6 + OpenEvolve 83.2
4 GPT-OSS + ShinkaiEvolve 28.0
5 GPT-OSS + ABMCTS 14.9
6 GPT-OSS + OpenEvolve 0.0

Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.