WirelessChannelSimulation
Task 47 / 47
HighReliableSimulation
Estimate very low BER for Hamming(127,120) over AWGN where naive Monte Carlo is inefficient: design importance sampling or variance-reduction samplers for deep-error events. Fixed evaluator settings score statistical efficiency and correctness—wireless link reliability engineering with rare-event simulation.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | Seed 2.0 Pro | 304.0437 | Gold |
| 2 | Claude Opus 4.6 | 292.3228 | Silver |
| 3 | DeepSeek V3.2 | 291.9451 | Bronze |
| 4 | Qwen3 Coder Next | 259.9776 | — |
| 5 | GLM-5 | 248.0119 | — |
| 6 | Grok 4.20 | 245.7082 | — |
| 7 | Gemini 3.1 Pro Preview | 232.9071 | — |
| 8 | GPT-5.4 | 231.22403 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | GPT-OSS + ABMCTS | 100.0 |
| 2 | Claude Opus 4.6 + OpenEvolve | 33.4 |
| 3 | GPT-OSS + OpenEvolve | 25.4 |
| 4 | Claude Opus 4.6 + ABMCTS | 15.2 |
| 5 | Claude Opus 4.6 + ShinkaiEvolve | 14.6 |
| 6 | GPT-OSS + ShinkaiEvolve | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.