ReactionOptimisation
Task 36 / 47
snar_multiobjective
Multi-objective optimization of a continuous-flow SnAr reaction, trading productivity against waste or byproduct metrics along a Pareto front over continuous operating variables. Grounded in chemical engineering emulators (SUMMIT family), it reflects real plant trade-offs among yield, waste, and operability.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | GPT-5.4 | 87.393969 | Gold |
| 2 | Claude Opus 4.6 | 87.3657 | Silver |
| 3 | DeepSeek V3.2 | 82.7881 | Bronze |
| 4 | GLM-5 | 81.7614 | — |
| 5 | Gemini 3.1 Pro Preview | 80.1521 | — |
| 6 | Seed 2.0 Pro | 79.427 | — |
| 7 | Qwen3 Coder Next | 72.8477 | — |
| 8 | Grok 4.20 | 72.3909 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | Claude Opus 4.6 + OpenEvolve | 100.0 |
| 2 | Claude Opus 4.6 + ShinkaiEvolve | 83.0 |
| 3 | GPT-OSS + ShinkaiEvolve | 80.3 |
| 4 | GPT-OSS + OpenEvolve | 58.4 |
| 5 | GPT-OSS + ABMCTS | 54.4 |
| 6 | Claude Opus 4.6 + ABMCTS | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.