Navers lab
← All tasks
Kernel Engineering Task 20 / 47

TriMul

This benchmark asks for a high-performance TriMul-style GPU kernel under strict correctness, trading off tiling, layout, and occupancy—often VRAM-bound on consumer GPUs. Evaluation runs representative workloads and scores both accuracy and speed against the benchmark's reference, highlighting specialized GEMM-like kernel engineering.

Model leaderboard

Updated v1 results · 2026-09-15 · raw score, higher is better.

# Participant Raw score Medal
1 GLM-5 2.5688189 Gold
2 Claude Opus 4.6 2.5413409 Silver
3 Seed 2.0 Pro 2.5298516 Bronze
4 Grok 4.20 2.4463768 —
5 Qwen3 Coder Next 2.3658683 —
6 DeepSeek V3.2 2.3330277 —
7 Gemini 3.1 Pro Preview 2.3090983 —
8 GPT-5.4 2.2453885 —

Framework results · original paper

Historical results from the original evaluators · normalized score (0–100).

# Participant Score
1 Claude Opus 4.6 + ShinkaiEvolve 100.0
2 Claude Opus 4.6 + ABMCTS 10.0
3 GPT-OSS + ABMCTS 7.1
4 GPT-OSS + OpenEvolve 4.3
5 Claude Opus 4.6 + OpenEvolve 2.3
6 GPT-OSS + ShinkaiEvolve 0.0

Model results use the updated scores and medal thresholds. No valid score earns zero medal credit. Historical framework scores use the original evaluators and have not been updated alongside the model results.