JobShop
Task 17 / 47
ta
JSSP on Taillard's TA family—large, widely used instances for makespan minimization that stress heuristics and parallel search at scale. Scoring against Taillard best-known solutions benchmarks industrial-grade job-shop solvers.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | Claude Opus 4.6 | 90.8322 | Gold |
| 2 | GLM-5 | 86.8095 | Silver |
| 3 | GPT-5.4 | 86.160701 | Bronze |
| 4 | Gemini 3.1 Pro Preview | 85.7065 | — |
| 5 | Qwen3 Coder Next | 85.5489 | — |
| 6 | Grok 4.20 | 84.9136 | — |
| 7 | DeepSeek V3.2 | 84.9043 | — |
| 8 | Seed 2.0 Pro | 83.9694 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | Claude Opus 4.6 + ShinkaiEvolve | 100.0 |
| 2 | Claude Opus 4.6 + OpenEvolve | 97.1 |
| 3 | Claude Opus 4.6 + ABMCTS | 68.4 |
| 4 | GPT-OSS + ShinkaiEvolve | 66.3 |
| 5 | GPT-OSS + ABMCTS | 24.7 |
| 6 | GPT-OSS + OpenEvolve | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.