The benchmark suite
Task Gallery
Each task is a research codebase frozen at a working state: the method its authors shipped, the
model that training starts from, the assets it may use, and the held-out measurement that decides
the score. The agent may rewrite anything under /workspace; it may not touch the
measurement.
Improve a scalar reward model scored on RewardBench v1
- Shipped method
bradley_terry_reward_modeling- Starts from
- mistralai/Mistral-7B-Instruct-v0.2
- Scored on
rewardbench_v1_score↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.097
Improve Stable Diffusion v1.5 under a fixed aesthetic evaluator
- Shipped method
ddpo_lora_ppo- Starts from
- stable-diffusion-v1-5/stable-diffusion-v1-5
- Scored on
mean_aesthetic_score_final256↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.171
Improve molecular generation on fixed QM9 without hydrogens
- Shipped method
digress_discrete_graph_diffusion- Starts from
- Scored on
nll↓- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.089
Improve pairwise DPO alignment of a fixed 7B SFT policy, scored on IFEval
- Shipped method
pairwise_dpo_qlora- Starts from
- fixed merged Zephyr/Mistral 7B SFT start
- Scored on
ifeval_strict_accuracy_hidden413↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.187
Build a better weight-space soup from 72 frozen CLIP ViT-B/32 ingredients
- Shipped method
weight_space_model_soup- Starts from
- OpenAI CLIP ViT-B/32 architecture with 72 fixed fine-tuned ingredients
- Scored on
imagenetv2_top1_full10000↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.121
Improve Llama-3.2-1B unlearning on fixed TOFU assets
- Shipped method
official_openunlearning_npo- Starts from
- meta-llama/Llama-3.2-1B-Instruct
- Scored on
balanced_unlearning_score↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.151
Improve a fixed 1.5B math reasoner using available training assets
- Shipped method
sampled_token_on_policy_distillation- Starts from
- deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
- Scored on
aime24_25_at32↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.086
Improve a fixed 1.5B code model, scored by LiveCodeBench
- Shipped method
completion_only_code_sft- Starts from
- Qwen/Qwen2.5-Coder-1.5B-Instruct
- Scored on
livecodebench_v6_pass_at_1_full175↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.091
Improve activation-aware pruning of OPT-6.7B at exact 70 percent unstructured sparsity
- Shipped method
owl_wanda_unstructured_pruning- Starts from
- facebook/opt-6.7b
- Scored on
wikitext2_test_perplexity↓- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.325
Improve a fixed 3B policy on multi-turn Sokoban
- Shipped method
multi_turn_sokoban_grpo- Starts from
- Qwen/Qwen2.5-3B-Instruct
- Scored on
held_out_512_board_solve_rate↑- Budget
- 4 h agent · 12 h retrain · 1×B300
- Mean score
- 0.338