Reports MRF-R-2026-01

Report MRF-R-2026-01

The State of Agent Evaluation

What certified task families reveal about how frontier labs measure agents

Dr. Ricardo Arcifa, Francieli Carra · Montana Research Foundation

Executive summary

Illustrative sample report. It draws on the foundation's three 2026 preprints to show how task certification, held-out grading criteria, and expectancy framing change what an agent benchmark actually measures, and what a lab or policy team should ask for before trusting a headline pass rate.

Key findings

  1. Reference-solution acceptance misses four classes of broken grader. Adversarial certification refused all eight deliberately defective fixtures at the stage designed to catch each one.
  2. Under a criterion held out of the environment, agents over-state their probability of exact generalization by 0.24 to 0.48 on two of three families, and by at most 0.13 on the third.
  3. A same-distribution validation sample does not close that gap. Submitted rules fit it at or near accuracy 1.0 and still fail held-out, so it tells the agent nothing.
  4. A negative, individual-form ability label raised the reasoning tokens one frontier model spent by about half a standard deviation, without an established change in pass rate.
  5. Every number in the underlying studies regenerates from committed run records and one audit script. Benchmark reports should meet the same bar.

Cite this report

Arcifa, R., & Carra, F. (2026). The State of Agent Evaluation. Montana Research Foundation report MRF-R-2026-01. https://montanaresearch.org/reports/mrf-r-2026-01/

BibTeX
@techreport{arcifa2026state,
  title = {The State of Agent Evaluation},
  author = {Arcifa, Ricardo and Carra, Francieli},
  institution = {Montana Research Foundation},
  type = {Insight report},
  number = {MRF-R-2026-01},
  year = {2026},
  month = {8},
  url = {https://montanaresearch.org/reports/mrf-r-2026-01/},
  note = {PDF: https://montanaresearch.org/reports/mrf-r-2026-01.pdf}
}