PPTBench
A benchmark for faithful, editable reconstruction of scientific flow diagrams.
The leaderboard
36 configurations across 10 model families are evaluated on the same 500 tasks. The best system reaches 77.34; the full table keeps every reasoning effort and harness visible.
The benchmark
Each task starts from a scientific diagram and asks an agent to produce a one-page PowerPoint made from native, editable elements. Hard gates check artifact validity and process semantics before structured visual detail is scored.
Task coverage
The task set spans 13 source domains. The bars show the complete domain distribution; the cards are representative references rather than an exhaustive task list.
One image in. One slide out.
This preview uses a real GPT-5.6 Sol Max reconstruction. The animation reveals native PowerPoint objects one by one in slide order; it is an illustrative reveal, not the model's actual authoring trace.
How scoring works
Every reconstruction passes through the same staged review: deterministic artifact checks first, then three independent GPT-5.6 Luna High judge rounds for process meaning and fine-grained fidelity.
Artifact Gate
The output must open, render and remain a one-page, editable PPTX. A missing, corrupt or unusable artifact receives zero before any visual comparison.
- PPTX opens and renders
- One-page output
- Native editable content
Semantic & Render Gate
Three independent GPT-5.6 Luna High rounds check both hard gates: semantic preservation (nodes, connections and direction) and render/text usability (legibility, overflow and basic layout). A hard-gate failure in at least two rounds assigns zero; otherwise the case proceeds to detail review.
- Nodes, connections & direction
- Render gate: large-scale corruption and blank regions
- Text gate: overflow, overlap and broken wrapping
Detail Ranking
Findings from non-gated rounds are merged by dimension and issue type. Each issue uses its maximum affected fraction and the frozen ratio staircase to produce the final score out of 100.
- Layout and composition · 30
- Text and typography · 40
- Local graphics · 30
Key findings
Agents know PPT syntax, but visual ability is weak
Valid artifacts are nearly free; correct ones are not, revealing that most failures happen after a readable PPTX has already been produced.
Agents still struggle with PPT details
After the hard gates, text issues contribute 51.7% of deductions, with unexpected wrapping the largest and most stable defect.
More reasoning is not always useful
Higher reasoning effort mainly raises gate pass rates, while conditional detail quality barely improves and the maximum setting is not consistently best.
More checking tracks with higher scores
Across generation traces, configurations with more inspection steps tend to score higher, revealing that verification matters alongside the initial draft.