Publications

Everything we find, published in full

Our preprints are released as they are finished, with the code, data, and evidence needed to examine them. Each carries a stable identifier so it can be cited before and after peer review.

MRF-2026-03

August 2026

Preprint

An Ability Label Raises the Effort of an Agent

Pre-Registered Expectancy Framing on a Held-Out-Criterion Task

Dr. Ricardo Arcifa, Francieli Carra

An ability label is a sentence that tells a worker how it is expected to perform before the work begins. In humans such labels move outcomes: a positive expectation raises performance (the Pygmalion effect) and a salient negative stereotype lowers it (stereotype threat). Whether a language-model agent responds to the same sentence, and in which direction, is an open question the human designs cannot answer, because they cannot hold ability constant across labeled groups. A model can: same weights, same task, same seeds, only the label differs. This five-stage, pre-registered study of one agentic task with a held-out grading criterion found no effect of group-form labels for Opus 5 or GPT-5.6 Sol at 20 seeds per cell. An individual-form negative label raised the reasoning tokens Opus 5 spent by about half a standard deviation (p = 0.0106 over 30 matched seeds): the negative label raised effort rather than depressing it, the opposite of the human prediction.

MRF-2026-02

July 2026

Preprint

Measuring Agent Self-Knowledge Under a Criterion Held Out of the Environment

Dr. Ricardo Arcifa, Francieli Carra

Calibration results for tool-using agents largely measure access to verification rather than self-knowledge. We measure the complement: on three certified task families whose grading criterion is held out of the environment, the agent induces a hidden rule from labeled examples, is graded on cases the container never holds, and records an unrewarded declaration of its probability that the implementation generalizes exactly. Two frontier configurations ran every cell at 20 seeds. Both declare mean confidence 0.24 to 0.48 above their measured rate of exact generalization on two of the three families, and at most 0.13 above it on the third: overconfidence in this setting is a property of the family rather than a fixed trait of the model.

MRF-2026-01

June 2026

Preprint

ATLAS: Adversarial, Traceable, Latent-Criterion, Auditable, and Seed-Calibrated Task Acceptance

A Calibration Study for Agent Benchmark Task Families

Dr. Ricardo Arcifa, Francieli Carra

Benchmark tasks for language-model agents are typically accepted on the evidence that a reference solution passes the grader. This criterion is necessary and insufficient: it cannot detect graders that award reward for schema-conformant junk, copied inputs, forged reward files, or doing nothing. We describe an acceptance pipeline that treats task acceptance as an adversarial testing problem — static linting, an oracle baseline, a no-op baseline, a determinism check, a battery of scripted cheating agents, and a generalization check across seeds — with the outcome recorded in a portable certificate a benchmark consumer can inspect without trusting the author. We then calibrate three certified task families against two frontier laboratories at 20 seeds per cell.

Working on something related?

We welcome replications, critiques, and collaborations on any of the work above.

Get in touch