ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance
A Calibration Study for Agent Benchmark Task Families
Abstract
Benchmark tasks for language-model agents are typically accepted on the evidence that a reference solution passes the grader. This criterion is necessary and insufficient: it cannot detect graders that award reward for schema-conformant junk, copied inputs, forged reward files, or doing nothing. We describe an acceptance pipeline that treats task acceptance as an adversarial testing problem — static linting, an oracle baseline, a no-op baseline, a determinism check, a battery of scripted cheating agents, and a generalization check across seeds — with the outcome recorded in a portable certificate a benchmark consumer can inspect without trusting the author. We then calibrate three certified task families against two frontier laboratories at 20 seeds per cell.
Licensed under CC BY 4.0. Every number in the paper traces to a committed artifact named in the text.
Cite this preprint
Arcifa, R., & Carra, F. (2026). ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance. Montana Research Foundation preprint MRF-2026-01. https://montanaresearch.org/publications/mrf-2026-01/
BibTeX
@techreport{arcifa2026atlas,
title = {ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance},
author = {Arcifa, Ricardo and Carra, Francieli},
institution = {Montana Research Foundation},
type = {Preprint},
number = {MRF-2026-01},
year = {2026},
month = {6},
url = {https://montanaresearch.org/publications/mrf-2026-01/},
note = {PDF: https://montanaresearch.org/papers/mrf-2026-01-atlas.pdf}
}