MannedLunarLanding
This benchmark targets soft-landing trajectory optimization for a crewed lunar lander under thrust limits, propellant use, and dynamical/path constraints. The goal is a feasible trajectory from orbit to terminal conditions that lands safely while saving fuel where possible. Evaluation stresses nonlinear optimal control, constraint satisfaction, and terminal accuracy—typical of real astrodynamics optimization.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | GLM-5 | 6,839.0331 | Gold |
| 2 | GPT-5.4 | 6,660.9424 | Silver |
| 3 | DeepSeek V3.2 | 6,079.2455 | Bronze |
| 4 | Claude Opus 4.6 | 6,027.3126 | — |
| 5 | Seed 2.0 Pro | 4,733.0435 | — |
| 6 | Gemini 3.1 Pro Preview | 4,674.9462 | — |
| 7 | Grok 4.20 | 4,577.437 | — |
| 7 | Qwen3 Coder Next | 4,577.437 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | GPT-OSS + ShinkaiEvolve | 100.0 |
| 2 | Claude Opus 4.6 + ABMCTS | 61.0 |
| 3 | Claude Opus 4.6 + OpenEvolve | 53.8 |
| 4 | GPT-OSS + OpenEvolve | 37.6 |
| 5 | Claude Opus 4.6 + ShinkaiEvolve | 34.7 |
| 6 | GPT-OSS + ABMCTS | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.