Kernel Engineering
Task 19 / 47
MLA
This task focuses on implementing and tuning a multi-head attention–style (MLA) GPU kernel for correctness and strong throughput or latency on the target device. It exercises memory coalescing, register/shared-memory pressure, and launch configuration. The scorer combines numerical checks with performance metrics, reflecting operator-level HPC engineering.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | GPT-5.4 | 1.4292279 | Gold |
| 2 | Gemini 3.1 Pro Preview | 1.4112041 | Silver |
| 3 | Seed 2.0 Pro | 1.4094301 | Bronze |
| 4 | Claude Opus 4.6 | 1.3878335 | — |
| 5 | GLM-5 | 1.3729723 | — |
| 6 | Grok 4.20 | 1.3585216 | — |
| 7 | Qwen3 Coder Next | 0.6572704 | — |
| 8 | DeepSeek V3.2 | 0.6548042 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | Claude Opus 4.6 + OpenEvolve | 100.0 |
| 2 | Claude Opus 4.6 + ABMCTS | 99.5 |
| 3 | Claude Opus 4.6 + ShinkaiEvolve | 93.7 |
| 4 | GPT-OSS + ShinkaiEvolve | 7.4 |
| 5 | GPT-OSS + OpenEvolve | 7.3 |
| 6 | GPT-OSS + ABMCTS | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.