EngDesign
Task 9 / 47
EngDesign
A bundle of EngDesign-sourced engineering cases (CY_03, WJ_01, XY_05, AM_02, AM_03, YJ_02, YJ_03) evaluated through Docker-backed scripts and per-case rubrics spanning structural, topology, and related design tasks. In the v1 pool it counts as one benchmark family though multiple distinct problem providers exist; use `task=engdesign` per the domain README.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | Gemini 3.1 Pro Preview | 27 | Gold |
| 1 | Grok 4.20 | 27 | Gold |
| 1 | Seed 2.0 Pro | 27 | Gold |
| 4 | GLM-5 | 25.5714 | — |
| 4 | Qwen3 Coder Next | 25.5714 | — |
| 6 | DeepSeek V3.2 | 21.7143 | — |
| 7 | GPT-5.4 | 1.3571429 | — |
| 8 | Claude Opus 4.6 | 1.3571 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | Claude Opus 4.6 + OpenEvolve | 100.0 |
| 2 | GPT-OSS + OpenEvolve | 53.2 |
| 3 | GPT-OSS + ShinkaiEvolve | 7.1 |
| 4 | Claude Opus 4.6 + ABMCTS | 7.1 |
| 5 | Claude Opus 4.6 + ShinkaiEvolve | 0.4 |
| 6 | GPT-OSS + ABMCTS | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.