Research
Benchmarks, papers, releases, and other lab updates — one place to track how the frontier moves.
-
BenchmarksPPTBench
A benchmark for faithful, editable reconstruction of scientific flow diagrams: 500 frozen tasks, native PowerPoint artifacts, and staged semantic and visual evaluation.
-
BenchmarksSWE Refactor Bench
Move a whole repository to another stack — language, framework, platform or build system — and change nothing about what it does. A migration starts with its tests already green, so passing them proves nothing: 20 real repositories graded by whether the old stack is gone, whether every recorded behaviour survived, and what six adversarial coding agents can still break.
-
BenchmarkAI4AI-Bench
Give a coding agent a real research codebase and four hours to improve how it trains a model. Most tune the settings around the existing method; few rewrite the method itself.
-
Open sourceOpenChronicle
An open, local-first memory layer for LLM agents. OpenChronicle captures structured context from your screen and apps into inspectable Markdown memory — projects, tools, people, decisions — that any tool-calling agent can read, run locally, and share across models.
-
BenchmarksFrontier-Engineering
Open engineering benchmark: agents on optimization across aerospace scheduling, EDA, quantum circuits, fiber optics, battery control, and more. 47 tasks across five engineering categories.