The Frontier · Tags
#evaluation
2 published articles
One million tokens fixes long-document recall
DeepSeek-V4 reports MRCR 1M MMR 83.5 and CorpusQA 1M ACC 62.0 at Pro Max. Strong on paper. Not a buyer corpus guarantee.
Agents & software · Field Notes
Vertex Gen AI eval pipeline replaces RAG demo scorecards
Google’s Vertex evaluation service and EvalTask score retrieval and generation separately; model-based rubrics replace hallway comparisons.