Navers lab
← Frontier-Engineering

Overall Leaderboard

47 tasks 5 categories

Across 47 engineering tasks, the current Medal Score credits a model for reaching each task's gold/silver/bronze thresholds. The eight-model results below use the updated scores from 2026-09-15. Source scores and thresholds.


Frontier Models

Eight frontier models evaluated on engineering problems with OpenEvolve, using the released submissions and updated evaluator results. This update combines rescored submissions with replacement candidates from shorter reruns; it is not a new uniform 100-iteration experiment.

Medal score

A model earns 1.00 / 0.67 / 0.33 for reaching the available gold / silver / bronze thresholds, then its credit is averaged over all tasks. Invalid results earn zero credit; unavailable medal tiers award no credit. The denominator remains 47 for v1 and 10 for the original, fixed v1-lite subset.

v1 · 47 tasks

# Model Medal 🥇 🥈 🥉
1 Claude Opus 4.6 0.533 14 15 3
2 GPT-5.4 0.454 18 4 2
3 GLM-5 0.347 7 8 12
4 Gemini 3.1 Pro Preview 0.284 7 7 5
5 DeepSeek V3.2 0.269 6 6 8
6 Grok 4.20 0.220 6 5 3
7 Seed 2.0 Pro 0.206 6 4 3
8 Qwen3 Coder Next 0.170 5 3 3

v1-lite · 10 tasks

# Model Medal 🥇 🥈 🥉
1 Claude Opus 4.6 0.501 3 3 0
2 Gemini 3.1 Pro Preview 0.300 1 2 2
2 GLM-5 0.300 1 2 2
4 DeepSeek V3.2 0.299 1 1 4
5 GPT-5.4 0.267 2 1 0
6 Grok 4.20 0.167 1 1 0
7 Seed 2.0 Pro 0.100 1 0 0
8 Qwen3 Coder Next 0.066 0 0 2

Average rank · original paper

Historical mean within-task rank over 47 tasks (lower is better), across the paper's nine models. This table predates the updated results above. The released raw scores do not include gpt-oss-120b, so a revised nine-model table is not available.

# Model Avg. rank
1 GPT-5.4 3.54
2 Claude Opus 4.6 3.63
3 GLM-5 4.34
4 DeepSeek V3.2 4.76
5 gpt-oss-120b 4.81
6 Gemini 3.1 Pro Preview 5.53
7 Grok 4.20 5.82
8 Seed 2.0 Pro 5.86
9 Qwen3 Coder Next 6.71

Performance profile · original paper

The original paper's Dolan–Moré performance profile, using the original scores rather than the updated results above: each curve shows how often a model stays near the best score on the bench within a given slack. Higher on the left means stronger and more consistent across the 47 tasks.

Dolan–Moré performance profile for frontier models on all 47 tasks.

Browse the benchmark task by task → open Tasks