Navers lab

Research

Benchmarks, papers, releases, and other lab updates — one place to track how the frontier moves.

  • PPTBench reference reconstruction preview
    Benchmarks

    PPTBench

    A benchmark for faithful, editable reconstruction of scientific flow diagrams: 500 frozen tasks, native PowerPoint artifacts, and staged semantic and visual evaluation.

  • One stage cannot judge a migration: prior benchmarks run a behavioural suite with no migration audit in front of it, while SWE Refactor Bench runs three stages — migration audit, behavioural tests, agentic verification. Below, the four migration categories and the scores of eight frontier models.
    Benchmarks

    SWE Refactor Bench

    Move a whole repository to another stack — language, framework, platform or build system — and change nothing about what it does. A migration starts with its tests already green, so passing them proves nothing: 20 real repositories graded by whether the old stack is gone, whether every recorded behaviour survived, and what six adversarial coding agents can still break.

  • AI4AI-Bench
    Benchmark

    AI4AI-Bench

    Give a coding agent a real research codebase and four hours to improve how it trains a model. Most tune the settings around the existing method; few rewrite the method itself.

  • BrowserBC method overview
    Papers

    BrowserBC

    BrowserBC distills successful human browser trajectories into reusable skills, helping web agents solve tasks with higher success rates and fewer interactions.

  • OpenChronicle — local-first memory for AI agents
    Open source

    OpenChronicle

    An open, local-first memory layer for LLM agents. OpenChronicle captures structured context from your screen and apps into inspectable Markdown memory — projects, tools, people, decisions — that any tool-calling agent can read, run locally, and share across models.

  • Frontier-Engineering benchmark composition
    Benchmarks

    Frontier-Engineering

    Open engineering benchmark: agents on optimization across aerospace scheduling, EDA, quantum circuits, fiber optics, battery control, and more. 47 tasks across five engineering categories.