Navers lab
← All tasks
EngDesign Task 9 / 47

EngDesign

A bundle of EngDesign-sourced engineering cases (CY_03, WJ_01, XY_05, AM_02, AM_03, YJ_02, YJ_03) evaluated through Docker-backed scripts and per-case rubrics spanning structural, topology, and related design tasks. In the v1 pool it counts as one benchmark family though multiple distinct problem providers exist; use `task=engdesign` per the domain README.

Model leaderboard

Updated v1 results · 2026-09-15 · raw score, higher is better.

# Participant Raw score Medal
1 Gemini 3.1 Pro Preview 27 Gold
1 Grok 4.20 27 Gold
1 Seed 2.0 Pro 27 Gold
4 GLM-5 25.5714 —
4 Qwen3 Coder Next 25.5714 —
6 DeepSeek V3.2 21.7143 —
7 GPT-5.4 1.3571429 —
8 Claude Opus 4.6 1.3571 —

Framework results · original paper

Historical results from the original evaluators · normalized score (0–100).

# Participant Score
1 Claude Opus 4.6 + OpenEvolve 100.0
2 GPT-OSS + OpenEvolve 53.2
3 GPT-OSS + ShinkaiEvolve 7.1
4 Claude Opus 4.6 + ABMCTS 7.1
5 Claude Opus 4.6 + ShinkaiEvolve 0.4
6 GPT-OSS + ABMCTS 0.0

Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.