Guides
The evaluation gap: a buyer’s field guide
How to stop buying models on leaderboards and start gating releases on your distribution.
Bottom line. Lab leaderboards are a weak procurement instrument. Treat evaluation as continuous instrumentation with replayable failures, or you will optimize fluency instead of operational readiness.
Why the gap is the product
Fluency stopped being a useful signal once models became polished. Failures moved into the long tail: retrieval misses, malformed tool arguments, and confident answers that invent policy. That gap between demo prompts and operational reality is now the product surface.
A practical eval protocol
- Build a fixed corpus of real tasks, including ugly ones.
- Version prompts, retrieval, and tools with every scored run.
- Separate capability scores from readiness (escalation, audit, kill switch).
- Require replay of failures before you renew a vendor contract.
Related Frontier reporting
FAQ
Are public leaderboards useless?
No — they help compare research artifacts. They are weak as the primary procurement score for production systems with tools, retrieval, and side effects.
What should buyers demand instead?
Private evals on your data, replayable failure cases, cost and latency at your quality target, and a named owner who can block a release.