Deep Extract v3 (semantic tune)
Reducto · targeted fields · 2026-07-25 · n=1
Raw PDF · prompt eng
$10.08beta96.0%16/3285.9%100.0%98.8%
Deep Extract v3 (strict contract)
Reducto · identical contract · 2026-07-24 · n=1
Raw PDF · strict contract
$11.02beta89.0%16/3261.1%100.0%98.0%
Claude Fable 5
Claude Code v2.1.216 · xhigh · 2026-07-21 · n=1
Agentic CLI
$102.8495.1%9/3284.0%99.5%96.8%
GPT-5.6-Sol
Codex CLI v0.146.0 · xhigh · 2026-08-02 · n=3 · arithmetic mean
Agentic CLI
$33.9698.8%8/3297.0%99.5%99.7%
Claude Opus 4.8
Claude Code v2.1.216 · xhigh · 2026-07-21 · n=1
Agentic CLI
$44.8697.7%7/3293.7%99.3%99.4%
GPT-5.5
Codex CLI v0.146.0 · xhigh · 2026-08-03 · n=3 · arithmetic mean
Agentic CLI
$43.1996.6%6/3289.5%99.4%99.3%
GPT-5.6-Terra
Codex CLI v0.146.0 · xhigh · 2026-07-31 · n=3 · arithmetic mean
Agentic CLI
$15.0195.7%6/3286.5%99.4%99.2%
GPT-5.6-Luna
Codex CLI v0.146.0 · xhigh · 2026-08-01 · n=3 · arithmetic mean
Agentic CLI
$2.4194.4%5/3282.5%99.1%98.7%
Bonsai 27B
Local + Together AI · page pipeline · 2026-07-26 · n=1
Page pipeline
$0.0016.7%0/3212.4%18.5%40.7%

Submit a verified result

Run the reference evaluator, then open a repository issue or Space discussion with the evaluation report, run metadata, and per-sample predictions. Scores are reproduced before publication.

Submission details →