Kernel Engineering
Task 18 / 47
FlashAttention
Optimize a causal scaled dot-product attention forward kernel (FlashAttention-style) for GPU execution while matching a reference numerically. The problem stresses tiled online softmax and memory locality. Scoring reports speed and correctness for fixed shapes, representing bandwidth-bound attention kernel work in production ML stacks.
Model leaderboard
Updated v1 results · 2026-09-15 · raw score, higher is better.
| # | Participant | Raw score | Medal |
|---|---|---|---|
| 1 | GPT-5.4 | 13.567075 | Gold |
| 2 | Qwen3 Coder Next | 13.334793 | Silver |
| 3 | Gemini 3.1 Pro Preview | 12.445371 | Bronze |
| 4 | Seed 2.0 Pro | 12.257567 | — |
| 5 | GLM-5 | 11.919302 | — |
| 6 | Claude Opus 4.6 | 11.700619 | — |
| 7 | Grok 4.20 | 11.607419 | — |
| 8 | DeepSeek V3.2 | 11.447719 | — |
Framework results · original paper
Historical results from the original evaluators · normalized score (0–100).
| # | Participant | Score |
|---|---|---|
| 1 | Claude Opus 4.6 + OpenEvolve | 100.0 |
| 2 | GPT-OSS + ShinkaiEvolve | 98.7 |
| 3 | Claude Opus 4.6 + ShinkaiEvolve | 98.7 |
| 4 | GPT-OSS + OpenEvolve | 56.9 |
| 5 | Claude Opus 4.6 + ABMCTS | 41.1 |
| 6 | GPT-OSS + ABMCTS | 0.0 |
Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.