Overall Leaderboard
Across 47 engineering tasks, the current Medal Score credits a model for reaching each task's gold/silver/bronze thresholds. The eight-model results below use the updated scores from 2026-09-15. Source scores and thresholds.
Frontier Models
Eight frontier models evaluated on engineering problems with OpenEvolve, using the released submissions and updated evaluator results. This update combines rescored submissions with replacement candidates from shorter reruns; it is not a new uniform 100-iteration experiment.
Medal score
A model earns 1.00 / 0.67 / 0.33 for reaching the available gold / silver / bronze thresholds, then its credit is averaged over all tasks. Invalid results earn zero credit; unavailable medal tiers award no credit. The denominator remains 47 for v1 and 10 for the original, fixed v1-lite subset.
v1 · 47 tasks
| # | Model | Medal | 🥇 | 🥈 | 🥉 |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | 0.533 | 14 | 15 | 3 |
| 2 | GPT-5.4 | 0.454 | 18 | 4 | 2 |
| 3 | GLM-5 | 0.347 | 7 | 8 | 12 |
| 4 | Gemini 3.1 Pro Preview | 0.284 | 7 | 7 | 5 |
| 5 | DeepSeek V3.2 | 0.269 | 6 | 6 | 8 |
| 6 | Grok 4.20 | 0.220 | 6 | 5 | 3 |
| 7 | Seed 2.0 Pro | 0.206 | 6 | 4 | 3 |
| 8 | Qwen3 Coder Next | 0.170 | 5 | 3 | 3 |
v1-lite · 10 tasks
| # | Model | Medal | 🥇 | 🥈 | 🥉 |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6 | 0.501 | 3 | 3 | 0 |
| 2 | Gemini 3.1 Pro Preview | 0.300 | 1 | 2 | 2 |
| 2 | GLM-5 | 0.300 | 1 | 2 | 2 |
| 4 | DeepSeek V3.2 | 0.299 | 1 | 1 | 4 |
| 5 | GPT-5.4 | 0.267 | 2 | 1 | 0 |
| 6 | Grok 4.20 | 0.167 | 1 | 1 | 0 |
| 7 | Seed 2.0 Pro | 0.100 | 1 | 0 | 0 |
| 8 | Qwen3 Coder Next | 0.066 | 0 | 0 | 2 |
Average rank · original paper
Historical mean within-task rank over 47 tasks (lower is better), across the paper's nine models. This table predates the updated results above. The released raw scores do not include gpt-oss-120b, so a revised nine-model table is not available.
| # | Model | Avg. rank |
|---|---|---|
| 1 | GPT-5.4 | 3.54 |
| 2 | Claude Opus 4.6 | 3.63 |
| 3 | GLM-5 | 4.34 |
| 4 | DeepSeek V3.2 | 4.76 |
| 5 | gpt-oss-120b | 4.81 |
| 6 | Gemini 3.1 Pro Preview | 5.53 |
| 7 | Grok 4.20 | 5.82 |
| 8 | Seed 2.0 Pro | 5.86 |
| 9 | Qwen3 Coder Next | 6.71 |
Performance profile · original paper
The original paper's Dolan–Moré performance profile, using the original scores rather than the updated results above: each curve shows how often a model stays near the best score on the bench within a given slack. Higher on the left means stronger and more consistent across the 47 tasks.
Browse the benchmark task by task → open Tasks