Navers lab
← AI4AI-Bench

The benchmark suite

Task Gallery

Each task is a research codebase frozen at a working state: the method its authors shipped, the model that training starts from, the assets it may use, and the held-out measurement that decides the score. The agent may rewrite anything under /workspace; it may not touch the measurement.