Benchmarks

Evidence before
optimization claims.

We separate source reduction, recoverability and answer quality. A smaller request is useful only when the application can still produce the required result.

COMBINED OFFLINE HOLDOUT55 evaluated54 accepted · 54/54 accepted candidates exactly recoverable
SQuAD: 30 evaluated, 29 accepted · QASPER: 25 evaluated, 25 accepted

Published product evidence

Results with their
scope attached.

These are development holdouts, not a universal customer savings guarantee.

DatasetCasesSource reductionExact recoveryEvidence type
SQuAD30 articles46.97%29/29 acceptedOffline holdout
QASPER25 papers49.91%25/25 acceptedOffline holdout
01

Seal the cases

Keep development tuning separate from the final evaluation set.

02

Measure preservation

Track literal first-view evidence and exact source recovery independently.

03

Test the full loop

Include recovery and fallback turns in net token and answer-quality results.

What these numbers prove

Accepted sources can be reduced and reconstructed exactly in these datasets. They do not prove that every model will request missing evidence or that every workload will preserve answer quality.

Read the limitation