ReactionOptimisation
Task 34 / 47
mit_case1_mixed
Mixed-variable reaction yield optimization from the MIT_case1 setting: continuous process variables plus a categorical catalyst choice. It stresses black-box optimization with discrete decisions, evaluated via the benchmark's unified hook into SUMMIT-style verification—common in digital-twin reaction tuning.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | GPT-5.4 | 98.662146 | Gold |
| 2 | Claude Opus 4.6 | 98.6621 | Silver |
| 3 | DeepSeek V3.2 | 98.6041 | Bronze |
| 4 | Gemini 3.1 Pro Preview | 96.5437 | — |
| 5 | GLM-5 | 95.9314 | — |
| 6 | Seed 2.0 Pro | 95.4297 | — |
| 7 | Qwen3 Coder Next | 95.3732 | — |
| 8 | Grok 4.20 | 87.3082 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | GPT-OSS + ShinkaiEvolve | 100.0 |
| 2 | GPT-OSS + ABMCTS | 81.7 |
| 3 | Claude Opus 4.6 + OpenEvolve | 0.0 |
| 4 | Claude Opus 4.6 + ShinkaiEvolve | 0.0 |
| 5 | Claude Opus 4.6 + ABMCTS | 0.0 |
| 6 | GPT-OSS + OpenEvolve | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.