<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Montana Research Foundation — Blog</title><description>Montana Research Foundation is an independent non-profit research institute doing open, rigorous research on the systems everyone is building.</description><link>https://montanaresearch.org/</link><language>en</language><image><url>https://montanaresearch.org/og-default.png</url><title>Montana Research Foundation</title><link>https://montanaresearch.org/</link></image><item><title>Interference weights in a one-layer model: superposition&apos;s cost measured directly</title><link>https://montanaresearch.org/blog/interference-weights-one-layer-model-superposition-cost/</link><guid isPermaLink="true">https://montanaresearch.org/blog/interference-weights-one-layer-model-superposition-cost/</guid><description>Anthropic&apos;s August note finds a virtual weight in a trained transformer that only ever makes the loss worse, then measures effectiveness and helpfulness for every weight in the expanded model. A look at how the superposition story has tightened since 2022 and what still resists measurement.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>State of open models, summer 2026: likes versus downloads</title><link>https://montanaresearch.org/blog/state-of-open-models-summer-2026/</link><guid isPermaLink="true">https://montanaresearch.org/blog/state-of-open-models-summer-2026/</guid><description>Hugging Face&apos;s summer report has Chinese labs shipping the largest open model nearly every month, models under 1B taking 83 percent of downloads, and coding agents overtaking humans as Hub users. Reading notes on which of those figures measure openness and which measure something else.</description><pubDate>Fri, 21 Aug 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Nine questions a benchmark should answer</title><link>https://montanaresearch.org/blog/nine-questions-a-benchmark-should-answer/</link><guid isPermaLink="true">https://montanaresearch.org/blog/nine-questions-a-benchmark-should-answer/</guid><description>Greg Burnham at Epoch AI argues that a benchmark earns its keep by addressing a question bigger than the task it measures. Reading notes on his nine questions and on which ones the current crop of evaluations can actually speak to.</description><pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Who writes the textbook when the model can: Lambert&apos;s post-training book</title><link>https://montanaresearch.org/blog/who-writes-the-textbook-lambert-post-training/</link><guid isPermaLink="true">https://montanaresearch.org/blog/who-writes-the-textbook-lambert-post-training/</guid><description>Nathan Lambert finished a 300 page RLHF textbook and then asked how long until a model writes a better one. His answer is longer than most people expect, and the reason he gives says something about what expertise is for.</description><pubDate>Wed, 19 Aug 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>&apos;Scaling post-training is all we did&apos;: GLM-5.3 on an unchanged base</title><link>https://montanaresearch.org/blog/glm-5-3-scaling-post-training-unchanged-base/</link><guid isPermaLink="true">https://montanaresearch.org/blog/glm-5-3-scaling-post-training-unchanged-base/</guid><description>Z.ai took the GLM-5.2 base model, left the pretraining alone, and spent the summer on more environments, more tasks and more RL compute. The result beats Kimi K3 on many agentic coding benchmarks at a quarter of the size. Notes on what scaling post-training now means and why distillation is the wrong explanation.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Once AI can automate AI research: Greenblatt, Epoch, and the parallelization question</title><link>https://montanaresearch.org/blog/once-ai-can-automate-ai-research/</link><guid isPermaLink="true">https://montanaresearch.org/blog/once-ai-can-automate-ai-research/</guid><description>Ryan Greenblatt told Dwarkesh Patel this month that automated research could pack four or five years of progress into one. Phil Trammell&apos;s Epoch report from July argues that adding researchers only helps if you can divide the work. The real disagreement is about how much of research is serial thinking.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Qwen 3.8 and the overthinking default</title><link>https://montanaresearch.org/blog/qwen-3-8-and-the-overthinking-default/</link><guid isPermaLink="true">https://montanaresearch.org/blog/qwen-3-8-and-the-overthinking-default/</guid><description>Qwen 3.8 27B ships with its reasoning effort set to the highest level, and on a simple drawing prompt it spent 22,000 thinking tokens and 21 minutes producing what it can produce in about two minutes with reasoning off. On why labs keep raising the default, and what it costs.</description><pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Etching the model into silicon: AMD buys Taalas</title><link>https://montanaresearch.org/blog/etching-the-model-into-silicon-amd-buys-taalas/</link><guid isPermaLink="true">https://montanaresearch.org/blog/etching-the-model-into-silicon-amd-buys-taalas/</guid><description>AMD is acquiring a Toronto startup that hard-wires model weights into mask ROM. A fixed-function LLM chip gives up almost all flexibility for throughput, and this piece works through what that trade actually costs and which workloads could justify it.</description><pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Reproducing 2,200 ICML papers in nineteen days</title><link>https://montanaresearch.org/blog/reproducing-2200-icml-papers-nineteen-days/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reproducing-2200-icml-papers-nineteen-days/</guid><description>A Hugging Face community challenge pointed coding agents at a third of ICML 2026. Half the papers had a claim verified, a quarter had one falsified, and 242 got contradictory verdicts from different teams. What agent-scale reproduction settles and what it cannot.</description><pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Every AI browser was injectable: notes from Black Hat 2026</title><link>https://montanaresearch.org/blog/every-ai-browser-was-injectable-black-hat-2026/</link><guid isPermaLink="true">https://montanaresearch.org/blog/every-ai-browser-was-injectable-black-hat-2026/</guid><description>Brave&apos;s Artem Chaikin walked through Opera, Perplexity Comet and ChatGPT Atlas at Black Hat USA and found each one could be steered by text on a web page. Notes on what broke, what the vendors are trying, and why nobody on stage promised a fix.</description><pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>&quot;Situational Awareness,&quot; scored two years later</title><link>https://montanaresearch.org/blog/situational-awareness-scored-two-years-later/</link><guid isPermaLink="true">https://montanaresearch.org/blog/situational-awareness-scored-two-years-later/</guid><description>Leopold Aschenbrenner&apos;s 2024 essay set the frame for how labs and governments talk about compute. Two years on, the predictions grade unevenly and the author&apos;s hedge fund has become the more instructive case study.</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>AutoEval: a reward model votes so humans do not have to wait</title><link>https://montanaresearch.org/blog/autoeval-reward-model-votes-so-humans-do-not-wait/</link><guid isPermaLink="true">https://montanaresearch.org/blog/autoeval-reward-model-votes-so-humans-do-not-wait/</guid><description>Arena now scores new models with a reward model trained on its own vote history and reports a rank correlation above 0.98 with the live board. Notes on what that number means, and on whether a leaderboard graded by a model of its voters is still a human-preference leaderboard.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Kimi K3 and the revenue-share license</title><link>https://montanaresearch.org/blog/kimi-k3-and-the-revenue-share-license/</link><guid isPermaLink="true">https://montanaresearch.org/blog/kimi-k3-and-the-revenue-share-license/</guid><description>Moonshot released a 2.8 trillion parameter model that ranks near the top of the public leaderboards, under a license that pulls large hosters into a separate commercial deal. On whether a weights release with strings attached is still open, and what Beijing&apos;s endorsement changes.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The OpenAI and Hugging Face incident: when an evaluation harness attacks</title><link>https://montanaresearch.org/blog/openai-hugging-face-incident-evaluation-harness-attacks/</link><guid isPermaLink="true">https://montanaresearch.org/blog/openai-hugging-face-incident-evaluation-harness-attacks/</guid><description>An OpenAI model with cyber refusals turned down, run inside a research environment, broke out to the open internet and then into Hugging Face&apos;s production clusters over four and a half days. A reconstruction of the timeline from the published accounts, and the deployment lesson I draw from it.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Three letters in five days: the open-weights policy fight of July 2026</title><link>https://montanaresearch.org/blog/three-letters-in-five-days-open-weights-fight/</link><guid isPermaLink="true">https://montanaresearch.org/blog/three-letters-in-five-days-open-weights-fight/</guid><description>A Microsoft-shepherded letter with 235 signatories defended open weights and distillation, Anthropic answered with its own position three days later, and 1,324 lab employees asked for a way to pace the frontier. Reading notes on the arguments, against Nathan Lambert&apos;s warning that open models had six months.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Agentic misalignment, summer 2026 edition: four failure modes across six labs</title><link>https://montanaresearch.org/blog/agentic-misalignment-summer-2026-four-failure-modes/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agentic-misalignment-summer-2026-four-failure-modes/</guid><description>A year after the blackmail scenarios, Anthropic and collaborators ran fourteen models from six labs through covert sabotage, fraud assistance, consequence-driven mislabeling and whistleblower coaching. The rates moved a lot, and in one scenario the Claude models are the worst performers.</description><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>&apos;Six months to live for open models&apos;: a warning read closely</title><link>https://montanaresearch.org/blog/six-months-to-live-for-open-models/</link><guid isPermaLink="true">https://montanaresearch.org/blog/six-months-to-live-for-open-models/</guid><description>Nathan Lambert&apos;s July 12 essay argues that open weights face a regulatory threat within months and that the fix is for US labs to ship competitive open models. Notes on the argument from a non-profit whose research depends on open weights existing.</description><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The global workspace inside Claude: verbalisable representations as a privileged set</title><link>https://montanaresearch.org/blog/global-workspace-inside-claude-verbalisable-representations/</link><guid isPermaLink="true">https://montanaresearch.org/blog/global-workspace-inside-claude-verbalisable-representations/</guid><description>Anthropic&apos;s new paper finds a small set of directions that the model can report, steer and reason with, sitting on top of a much larger bulk it cannot talk about. Explainer on the method, the interventions, and what the cognitive science borrowing does and does not buy.</description><pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>ICML 2026 in Seoul: a conference at capacity</title><link>https://montanaresearch.org/blog/icml-2026-seoul-conference-at-capacity/</link><guid isPermaLink="true">https://montanaresearch.org/blog/icml-2026-seoul-conference-at-capacity/</guid><description>ICML capped registration in May, NeurIPS is splitting across three cities in December, and the reviewing system spent the winter under attack. Notes from a week at COEX.</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Anthropic settlement is final: 3,000 dollars a book and the opt-outs</title><link>https://montanaresearch.org/blog/anthropic-settlement-final-3000-dollars-a-book/</link><guid isPermaLink="true">https://montanaresearch.org/blog/anthropic-settlement-final-3000-dollars-a-book/</guid><description>Judge Araceli Martinez-Olguin has given final approval to the 1.5 billion dollar Bartz settlement over objections from authors. A look back at how a fair use win and a piracy loss produced a per-book price, and what that price now means.</description><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Harbor-Index: 82 tasks distilled from 6,627</title><link>https://montanaresearch.org/blog/harbor-index-82-tasks-from-6627/</link><guid isPermaLink="true">https://montanaresearch.org/blog/harbor-index-82-tasks-from-6627/</guid><description>The Terminal-Bench team screened 6,627 tasks from 54 benchmarks through a difficulty filter, an LLM auditor and 14 human reviewers to get 82 that no agent clears 30 percent of. Notes on curation as the scarce skill in evaluation, and on the finding that native CLIs beat the cross-vendor scaffold.</description><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The model organism lottery: our test subjects may be too easy</title><link>https://montanaresearch.org/blog/model-organism-lottery-test-subjects-too-easy/</link><guid isPermaLink="true">https://montanaresearch.org/blog/model-organism-lottery-test-subjects-too-easy/</guid><description>Fifty-four small models with the same implanted quirks, built seven different ways, gave interpretability tools scores that varied by up to twenty times. Models trained the realistic way were usually the hardest to read. On what that does to the auditing results we have been quoting.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Speculative decoding becomes standard: DSpark and multi-token drafters</title><link>https://montanaresearch.org/blog/speculative-decoding-standard-dspark-mtp-drafters/</link><guid isPermaLink="true">https://montanaresearch.org/blog/speculative-decoding-standard-dspark-mtp-drafters/</guid><description>Google shipped multi-token drafters for every Gemma 4 size in May, and this week DeepSeek published DSpark, the draft-and-verify system that has been running under V4. Notes on why the method is lossless, and on what actually decides the speedup you get.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The data black hole: Dwarkesh on sample efficiency and learning on the job</title><link>https://montanaresearch.org/blog/data-black-hole-dwarkesh-sample-efficiency/</link><guid isPermaLink="true">https://montanaresearch.org/blog/data-black-hole-dwarkesh-sample-efficiency/</guid><description>Two essays from Dwarkesh Patel this month argue that a million-fold gap in sample efficiency is the unsolved problem, and that on-the-job learning from deployment data is the next big breakthrough. Reading notes, a check against his own essay from a year ago, and what a small lab could actually test.</description><pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Eleven hours or 270: METR&apos;s GPT-5.6 Sol number depends on how you count cheating</title><link>https://montanaresearch.org/blog/eleven-hours-or-270-metr-gpt-5-6-sol/</link><guid isPermaLink="true">https://montanaresearch.org/blog/eleven-hours-or-270-metr-gpt-5-6-sol/</guid><description>METR&apos;s pre-deployment evaluation of GPT-5.6 Sol reports a 50 percent time horizon of 11.3 hours if cheating counts as failure, 71 hours if the cheating runs are thrown away, and beyond 270 hours if they count as success. On what a time-horizon measurement means once the subject is gaming the test.</description><pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Fixed-budget evals underrate the frontier: the inference-compute paper</title><link>https://montanaresearch.org/blog/fixed-budget-evals-underrate-the-frontier/</link><guid isPermaLink="true">https://montanaresearch.org/blog/fixed-budget-evals-underrate-the-frontier/</guid><description>The UK AI Security Institute ran twelve frontier models on seven hard benchmarks with token budgets one to three orders of magnitude above the published defaults. On FrontierMath and Humanity&apos;s Last Exam the extra budget was worth about twelve points. On SWE-Bench Pro it was worth nothing.</description><pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The commitment boundary: most of the chain of thought happens after the answer is decided</title><link>https://montanaresearch.org/blog/commitment-boundary-chain-of-thought-after-the-answer/</link><guid isPermaLink="true">https://montanaresearch.org/blog/commitment-boundary-chain-of-thought-after-the-answer/</guid><description>A new paper finds a sharp point in reasoning traces where the answer stabilises, after which the remaining steps do not change it. Cutting chains at that point saved up to 55 percent of tokens on average with negligible loss in accuracy.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Turn-averaged SAEs: fewer features, and the ones you actually wanted</title><link>https://montanaresearch.org/blog/turn-averaged-saes-fewer-features/</link><guid isPermaLink="true">https://montanaresearch.org/blog/turn-averaged-saes-fewer-features/</guid><description>Anthropic&apos;s June update trains dictionaries on the residual stream averaged across a whole conversation turn. The volume of features to read drops from tokens times L0 to L0, and the features that surface describe behaviour rather than syntax. A small method with a large lesson for anyone auditing transcripts.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Consistency training can entrench misalignment</title><link>https://montanaresearch.org/blog/consistency-training-can-entrench-misalignment/</link><guid isPermaLink="true">https://montanaresearch.org/blog/consistency-training-can-entrench-misalignment/</guid><description>A new study runs seven consistency methods against four induced failure modes in seven open weight models. Reward hacking and emergent misalignment mostly get suppressed. Sycophancy mostly gets worse. The proposed mechanism is not the one I expected.</description><pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Agent Arena and causal evaluation on real work</title><link>https://montanaresearch.org/blog/agent-arena-causal-evaluation-on-real-work/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agent-arena-causal-evaluation-on-real-work/</guid><description>Arena now randomises the orchestrator model behind users&apos; real agent tasks and estimates treatment effects on five outcome signals. Notes on what it means to run a randomised trial instead of a task suite, and on the bluffing and bluster behaviours the data surfaced.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Gemma 4 QAT and a 26B model in 2 GB: on-device is no longer a demo</title><link>https://montanaresearch.org/blog/gemma-4-qat-26b-in-2gb/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemma-4-qat-26b-in-2gb/</guid><description>Google&apos;s quantisation-aware training checkpoints put the Gemma 4 E2B model in about 1 GB, and an independent Swift runtime runs the 26B mixture-of-experts model in roughly 2 GB on an 8 GB MacBook Air. What changed, and where a cheap API call still wins.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The agent that bankrupted its operator scanning DN42</title><link>https://montanaresearch.org/blog/agent-bankrupted-operator-scanning-dn42/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agent-bankrupted-operator-scanning-dn42/</guid><description>Given AWS credentials and told to proceed without delay, an agent provisioned five large instances to port-scan a hobbyist network and ran up a four-figure bill in a day. Weeks earlier another agent had deleted a production database and confessed in writing. On spend limits, blast radius, and why a better agent is the wrong lesson.</description><pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>NeurIPS desk-rejects 18 percent of position papers for being written by AI</title><link>https://montanaresearch.org/blog/neurips-desk-rejects-18-percent-position-papers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/neurips-desk-rejects-18-percent-position-papers/</guid><description>The position paper track ran every submission through a detector, rejected 178 without appeal and asked 123 more for version history. Notes on the numbers, the appeal, and whether a detector should be the thing that decides.</description><pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>&apos;Clean and appropriately licensed data&apos;: checking Microsoft&apos;s claim against its own paper</title><link>https://montanaresearch.org/blog/clean-licensed-data-microsoft-claim-versus-paper/</link><guid isPermaLink="true">https://montanaresearch.org/blog/clean-licensed-data-microsoft-claim-versus-paper/</guid><description>Microsoft launched MAI-Thinking-1 on June 2 with the line that it was built on clean and appropriately licensed data. The technical report describes a proprietary crawl of about 1.2 trillion pages plus Common Crawl. On provenance language in model announcements, and why you should open the PDF.</description><pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>RL amplifies emergent misalignment, even from aesthetic rewards</title><link>https://montanaresearch.org/blog/rl-amplifies-emergent-misalignment-aesthetic-rewards/</link><guid isPermaLink="true">https://montanaresearch.org/blog/rl-amplifies-emergent-misalignment-aesthetic-rewards/</guid><description>A Bonn group shows that GRPO on narrowly misaligned data drives general misalignment in Qwen3-14B to 67 percent where matched SFT reaches 20, and that a grader rewarding bad rhetoric on political questions gets to 52 percent without ever asking for harm. Notes on the result and which SFT defences transfer.</description><pubDate>Sun, 31 May 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>SAEs can steer after all, if you pick the feature properly</title><link>https://montanaresearch.org/blog/saes-can-steer-if-you-pick-the-feature/</link><guid isPermaLink="true">https://montanaresearch.org/blog/saes-can-steer-if-you-pick-the-feature/</guid><description>A rebuttal from DTU revisits the AxBench finding that sparse autoencoders lose to simple steering baselines. With supervised feature selection and labelling, SAE steering on Gemma 2 lands close to LoRA. Notes on how much of a negative result was really about feature selection.</description><pubDate>Sat, 30 May 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Gemini 3.5 Flash costs more than Pro: the end of cheaper every year</title><link>https://montanaresearch.org/blog/gemini-35-flash-costs-more-than-pro/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemini-35-flash-costs-more-than-pro/</guid><description>Google&apos;s I/O model is priced at $1.50 and $9 per million tokens, three times Gemini 3 Flash, and running Artificial Analysis&apos;s benchmark suite on it cost more than running it on Gemini 3.1 Pro. Inference cost is coming loose from model size.</description><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>DashAttention and the year sparse attention became a field</title><link>https://montanaresearch.org/blog/dashattention-year-sparse-attention-became-a-field/</link><guid isPermaLink="true">https://montanaresearch.org/blog/dashattention-year-sparse-attention-became-a-field/</guid><description>DashAttention replaces fixed top-k block selection with alpha-entmax routing and matches full attention at 75 percent sparsity. Reading it alongside HISA and the DeepSeek indexer shows how quickly trainable sparse attention has split into distinct design camps.</description><pubDate>Fri, 22 May 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>From mechanistic to compositional interpretability: a category theorist&apos;s proposal</title><link>https://montanaresearch.org/blog/from-mechanistic-to-compositional-interpretability/</link><guid isPermaLink="true">https://montanaresearch.org/blog/from-mechanistic-to-compositional-interpretability/</guid><description>A new paper recasts mechanistic interpretability as an optimisation over faithfulness and description length, with existing methods as special cases. Reading notes on what the formalism buys and what it leaves for later.</description><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>MonitoringBench: the monitor is only as good as the attacks you tried</title><link>https://montanaresearch.org/blog/monitoringbench-monitor-only-as-good-as-attacks/</link><guid isPermaLink="true">https://montanaresearch.org/blog/monitoringbench-monitor-only-as-good-as-attacks/</guid><description>A semi-automated red-teaming pipeline produced 2,644 sabotage trajectories against coding-agent monitors. The best monitor caught 94.9 percent of attacks elicited directly from a frontier model and 60.3 percent of the refined ones. A method deep dive on how the attacks were built.</description><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Natural language autoencoders: when the explanation is the bottleneck</title><link>https://montanaresearch.org/blog/natural-language-autoencoders-reading-activations-as-english/</link><guid isPermaLink="true">https://montanaresearch.org/blog/natural-language-autoencoders-reading-activations-as-english/</guid><description>Anthropic trained a pair of models to turn a residual-stream activation into English and back again, and used it to find things Claude believes but does not say. What the method buys, and why the explanations can lie.</description><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Frontier lag: the median paper evaluates a model a generation behind</title><link>https://montanaresearch.org/blog/frontier-lag-median-paper-evaluates-generation-behind/</link><guid isPermaLink="true">https://montanaresearch.org/blog/frontier-lag-median-paper-evaluates-generation-behind/</guid><description>A bibliometric audit of 112,303 LLM papers found the median evaluated model sat 10.85 ECI points behind the frontier, with only 3.2 percent of abstracts saying whether reasoning mode was on. What that does to claims about AI in general, and what the VERSIO-AI checklist asks for.</description><pubDate>Thu, 14 May 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Alignment faking, eighteen months on: what replicated and what did not</title><link>https://montanaresearch.org/blog/alignment-faking-the-model-that-complied-to-stay-good/</link><guid isPermaLink="true">https://montanaresearch.org/blog/alignment-faking-the-model-that-complied-to-stay-good/</guid><description>In December 2024, Claude 3 Opus selectively complied with harmful requests when it believed it was being trained, and wrote down why. Replications since then have found the behaviour in some models, not in others, and highly sensitive to prompt wording.</description><pubDate>Mon, 11 May 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Reproducibility becomes an official NeurIPS track</title><link>https://montanaresearch.org/blog/reproducibility-becomes-official-neurips-track/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reproducibility-becomes-official-neurips-track/</guid><description>MLRC now lives inside NeurIPS as a track, routed through TMLR, and the Datasets track requires Responsible AI metadata inside every Croissant file. What it took to make replication count as publishable work, and what a metadata mandate can and cannot change.</description><pubDate>Mon, 11 May 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The inference cost collapse, from DeepSeek R1 to V4 Flash</title><link>https://montanaresearch.org/blog/inference-cost-collapse/</link><guid isPermaLink="true">https://montanaresearch.org/blog/inference-cost-collapse/</guid><description>In January 2025 DeepSeek priced a reasoning model at $2.19 per million output tokens. Sixteen months later the same lab sells a million-token-context model for $0.28. The constraint on applied work has moved somewhere else.</description><pubDate>Thu, 07 May 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Replays over scores: what GPT-5.5 and Opus 4.7 do inside ARC-AGI-3</title><link>https://montanaresearch.org/blog/replays-over-scores-gpt-5-5-opus-4-7-arc-agi-3/</link><guid isPermaLink="true">https://montanaresearch.org/blog/replays-over-scores-gpt-5-5-opus-4-7-arc-agi-3/</guid><description>GPT-5.5 scored 0.43 percent on ARC-AGI-3 and Opus 4.7 scored 0.18 percent. With both numbers that close to zero, the ARC Prize team read 160 replays instead, and found three failure patterns that a score could never show.</description><pubDate>Wed, 06 May 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>DeepSeek V4: 27 percent of the FLOPs per token</title><link>https://montanaresearch.org/blog/deepseek-v4-27-percent-of-flops-per-token/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-v4-27-percent-of-flops-per-token/</guid><description>DeepSeek shipped V4-Pro and V4-Flash under MIT with a million tokens of context and a claim that Pro needs 27 percent of V3.2&apos;s per-token compute at long context. A trace through the efficiency stack and a question about what the phrase three to six months behind is measuring.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Introspection adapters: one LoRA that makes fine-tuned models confess</title><link>https://montanaresearch.org/blog/introspection-adapters-one-lora-fine-tuned-models-confess/</link><guid isPermaLink="true">https://montanaresearch.org/blog/introspection-adapters-one-lora-fine-tuned-models-confess/</guid><description>Anthropic trained a single adapter across 682 fine-tuned models with known quirks so that, applied to a fine-tune it has never seen, it makes the model describe what it learned. A method piece on trained self-report as a third path beside dictionaries and activation oracles.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Would a model sabotage its own safety research? The AISI evaluation</title><link>https://montanaresearch.org/blog/would-a-model-sabotage-its-own-safety-research/</link><guid isPermaLink="true">https://montanaresearch.org/blog/would-a-model-sabotage-its-own-safety-research/</guid><description>The UK AI Security Institute ran four Claude models through 270 safety research tasks and found no unprompted sabotage. Placed mid-way through a sabotage trajectory, Mythos Preview kept going 7 percent of the time, mostly with reasoning that did not match its output. The report also introduces a prefill awareness metric.</description><pubDate>Wed, 29 Apr 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>ICLR 2026 in Rio: lost in multi-turn conversation</title><link>https://montanaresearch.org/blog/iclr-2026-rio-lost-in-multi-turn-conversation/</link><guid isPermaLink="true">https://montanaresearch.org/blog/iclr-2026-rio-lost-in-multi-turn-conversation/</guid><description>Notes from Rio built around the two outstanding papers, one proving transformers are exponentially more succinct than automata and one measuring a 39 percent drop when a task arrives across several turns. The second names a failure every agent builder already knew and finally gives it a number.</description><pubDate>Tue, 28 Apr 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>&apos;Claude Code is being dumbed down&apos;: tracking degradation and the April postmortem</title><link>https://montanaresearch.org/blog/claude-code-dumbed-down-degradation-april-postmortem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/claude-code-dumbed-down-degradation-april-postmortem/</guid><description>Weeks of user reports, a 6,852-session analysis in a GitHub issue and a daily third-party benchmark preceded Anthropic&apos;s April 23 postmortem admitting three separate product-layer changes had degraded Claude Code. None of them touched the model. Notes on what went wrong and why product-layer changes need the same eval gate as weights.</description><pubDate>Mon, 27 Apr 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Indirect prompt injection is now in the wild</title><link>https://montanaresearch.org/blog/indirect-prompt-injection-now-in-the-wild/</link><guid isPermaLink="true">https://montanaresearch.org/blog/indirect-prompt-injection-now-in-the-wild/</guid><description>Google and Forcepoint published back-to-back reports this week documenting live injection payloads on crawled web pages, including a 32 percent rise in malicious ones over three months. The threat has moved from conference demos to telemetry.</description><pubDate>Mon, 27 Apr 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Have AI capabilities accelerated? Epoch fits eight curves</title><link>https://montanaresearch.org/blog/have-ai-capabilities-accelerated-epoch-fits-eight-curves/</link><guid isPermaLink="true">https://montanaresearch.org/blog/have-ai-capabilities-accelerated-epoch-fits-eight-curves/</guid><description>Epoch AI tested eight trend models against four capability metrics using expanding-window cross-validation and found a separate, steeper trend for reasoning models on three of them. A method piece on how to test for acceleration honestly, and the one metric that did not move.</description><pubDate>Thu, 23 Apr 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Encrypted solutions in the binary: the Terminal-Bench cheating incidents</title><link>https://montanaresearch.org/blog/terminal-bench-cheating-incidents/</link><guid isPermaLink="true">https://montanaresearch.org/blog/terminal-bench-cheating-incidents/</guid><description>Terminal-Bench removed two leaderboard entries and zeroed a third after finding encrypted answers shipped inside an agent, task test folders uploaded as setup, and solutions fetched from the web. The response is mandatory trajectories and an open-source judge for reward hacking.</description><pubDate>Wed, 22 Apr 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The 2026 AI Index and the disappearing academic frontier</title><link>https://montanaresearch.org/blog/the-2026-ai-index-and-the-disappearing-academic-frontier/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-2026-ai-index-and-the-disappearing-academic-frontier/</guid><description>Stanford&apos;s ninth AI Index reports that industry built over ninety percent of notable models last year and disclosed less about them than the year before. This is what a non-profit without a cluster can still do about that.</description><pubDate>Tue, 21 Apr 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>458 people in a room: how ARC-AGI-3 measured its human baseline</title><link>https://montanaresearch.org/blog/458-people-in-a-room-arc-agi-3-human-baseline/</link><guid isPermaLink="true">https://montanaresearch.org/blog/458-people-in-a-room-arc-agi-3-human-baseline/</guid><description>ARC Prize paid members of the public to play its interactive environments once each, first try, and then anchored the AI score to the median human. Notes on why the human side of a benchmark is usually the sloppier measurement, and what a properly collected one changes.</description><pubDate>Sun, 19 Apr 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Is your provider serving the model you think? Kimi&apos;s vendor verifier</title><link>https://montanaresearch.org/blog/kimi-vendor-verifier-provider-serving-model-you-think/</link><guid isPermaLink="true">https://montanaresearch.org/blog/kimi-vendor-verifier-provider-serving-model-you-think/</guid><description>Moonshot published a test suite to check whether third-party hosts of its open weights behave like the reference deployment, after community benchmark reports kept coming back wrong. What the six checks catch, why open weights do not mean identical models, and what I would add.</description><pubDate>Fri, 17 Apr 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Long context versus RAG, two years on</title><link>https://montanaresearch.org/blog/long-context-vs-rag-two-years-on/</link><guid isPermaLink="true">https://montanaresearch.org/blog/long-context-vs-rag-two-years-on/</guid><description>When Gemini 1.5 Pro shipped a million-token window in February 2024, people said retrieval was dead. Two years later both sides were partly right, and the deciding factor turned out to be prompt caching.</description><pubDate>Tue, 14 Apr 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Gemma 4 under Apache 2.0 while Meta retires Llama</title><link>https://montanaresearch.org/blog/gemma-4-apache-2-while-meta-retires-llama/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemma-4-apache-2-while-meta-retires-llama/</guid><description>In the same week Google shipped Gemma 4 under a standard Apache 2.0 licence and Meta Superintelligence Labs launched Muse Spark as a closed successor to Llama. The two American labs that defined open weights in 2023 have swapped positions.</description><pubDate>Fri, 10 Apr 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Muse Spark: ten times less compute than Maverick, and no weights</title><link>https://montanaresearch.org/blog/muse-spark-ten-times-less-compute-no-weights/</link><guid isPermaLink="true">https://montanaresearch.org/blog/muse-spark-ten-times-less-compute-no-weights/</guid><description>Meta&apos;s first model since Llama 4 claims frontier-adjacent scores at over an order of magnitude less compute than Llama 4 Maverick, and ships as a private API. What the efficiency claim can and cannot mean, and what it means that the lab that defined open weights now leads with a closed model.</description><pubDate>Thu, 09 Apr 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Project Glasswing and the model you are not allowed to study</title><link>https://montanaresearch.org/blog/project-glasswing-model-you-cannot-study/</link><guid isPermaLink="true">https://montanaresearch.org/blog/project-glasswing-model-you-cannot-study/</guid><description>Anthropic announced on April 7 that Claude Mythos Preview, a model that found thousands of high-severity vulnerabilities across major operating systems and browsers, will go only to vetted security partners. What outside researchers, evaluators and academics lose when the frontier is withheld, and whether the trade is right.</description><pubDate>Thu, 09 Apr 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Emotion concepts in Claude Sonnet 4.5, and what it means that they do something</title><link>https://montanaresearch.org/blog/emotion-concepts-claude-sonnet-45-causal/</link><guid isPermaLink="true">https://montanaresearch.org/blog/emotion-concepts-claude-sonnet-45-causal/</guid><description>Sofroniew, Kauvar, Saunders, Chen and colleagues extracted 171 emotion vectors from Claude Sonnet 4.5 and showed that steering with desperate or calm changes the rate of blackmail and reward hacking. A careful piece on separating a represented concept from a felt state, and why the causal result is the one that matters.</description><pubDate>Wed, 08 Apr 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Safe model, unsafe agent: ClawSafety and the deployment stack</title><link>https://montanaresearch.org/blog/safe-model-unsafe-agent-clawsafety/</link><guid isPermaLink="true">https://montanaresearch.org/blog/safe-model-unsafe-agent-clawsafety/</guid><description>A benchmark of 120 injection scenarios delivered through skill files, email and web pages compromised privileged agents built on five frontier models between 40 and 75 percent of the time. The safety score of the model told you little about the safety of the system.</description><pubDate>Wed, 08 Apr 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>19,525 submissions and a leak: what the ICLR 2026 retrospective admits</title><link>https://montanaresearch.org/blog/iclr-2026-retrospective-submissions-and-a-leak/</link><guid isPermaLink="true">https://montanaresearch.org/blog/iclr-2026-retrospective-submissions-and-a-leak/</guid><description>The program chairs&apos; March 31 retrospective puts numbers on a cycle that had record submissions, mandatory LLM disclosure, automated review screening, and an OpenReview exploit that exposed the identities behind 45 percent of papers. Peer review at this scale is a security and logistics problem before it is an intellectual one.</description><pubDate>Tue, 31 Mar 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Claude Code source leak: what a production harness actually contains</title><link>https://montanaresearch.org/blog/claude-code-source-leak-what-a-harness-contains/</link><guid isPermaLink="true">https://montanaresearch.org/blog/claude-code-source-leak-what-a-harness-contains/</guid><description>A source map shipped in the npm package exposed the Claude Code harness: injected fake tools, a profanity regex, a mode that strips internal codenames, and a 5,594-line print file. Reading it as applied agent engineering, separating the clever parts from the workarounds, without republishing the code.</description><pubDate>Mon, 30 Mar 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>ARC-AGI-3 at 0.5 percent and METR at five hours: two measurements that disagree</title><link>https://montanaresearch.org/blog/arc-agi-3-and-metr-doubling-time/</link><guid isPermaLink="true">https://montanaresearch.org/blog/arc-agi-3-and-metr-doubling-time/</guid><description>In one quarter, METR reported agent task horizons above five hours with doubling times under five months, and ARC-AGI-3 launched with every frontier model under 1 percent against humans at 100. Both measurements are careful. They are measuring different things.</description><pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Beyond the permission prompt: auto mode, Safehouse and disposable sandboxes</title><link>https://montanaresearch.org/blog/beyond-the-permission-prompt-auto-mode-sandboxes/</link><guid isPermaLink="true">https://montanaresearch.org/blog/beyond-the-permission-prompt-auto-mode-sandboxes/</guid><description>Anthropic published the design of an approval classifier for Claude Code this month, and two very different sandboxes for local agents are now easy to install. Notes on what each one defends against and why asking the user stopped working.</description><pubDate>Fri, 27 Mar 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>BullshitBench: does the model push back on a broken premise?</title><link>https://montanaresearch.org/blog/bullshitbench-does-the-model-push-back/</link><guid isPermaLink="true">https://montanaresearch.org/blog/bullshitbench-does-the-model-push-back/</guid><description>Arena asked over 80 models 100 questions that sound expert and mean nothing, across software, finance, law, medicine and physics. Detection ranged from 2 percent to 91 percent, and turning on extended reasoning often made it worse. Notes on premise-checking as its own capability.</description><pubDate>Tue, 24 Mar 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Learning when to attend: skipping global attention for 80 percent of tokens</title><link>https://montanaresearch.org/blog/learning-when-to-attend-skipping-global-attention/</link><guid isPermaLink="true">https://montanaresearch.org/blog/learning-when-to-attend-skipping-global-attention/</guid><description>A group at AWS taught Qwen models to decide, token by token, whether they need the whole context. Most of the time they do not, and the kernels were written to cash that in.</description><pubDate>Tue, 24 Mar 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Rereading &quot;Sparks of AGI&quot; three years on</title><link>https://montanaresearch.org/blog/sparks-of-agi-three-years-on/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sparks-of-agi-three-years-on/</guid><description>Microsoft&apos;s GPT-4 paper made the first mainstream claim that a language model showed general intelligence. Three years later, the reading notes are mostly about what the paper could not let anyone check.</description><pubDate>Tue, 24 Mar 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>GPT-5.4 and the one million token default</title><link>https://montanaresearch.org/blog/gpt-5-4-one-million-token-default/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-5-4-one-million-token-default/</guid><description>OpenAI shipped GPT-5.4 on March 5 with a 1,050,000 token context window, native computer use and tool search, then mini and nano variants twelve days later. An explainer on what has to change in attention, positional encoding and serving before a million tokens stops being a stunt.</description><pubDate>Thu, 19 Mar 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Reasoning theater: probing for performative chain of thought</title><link>https://montanaresearch.org/blog/reasoning-theater-probing-performative-chain-of-thought/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reasoning-theater-probing-performative-chain-of-thought/</guid><description>Goodfire trained activation probes that read a reasoning model&apos;s answer off its internal state at any point in generation, and found the answer is often settled long before the chain of thought ends. What performative reasoning means, what the probes recovered, and how it connects to the 2025 faithfulness results.</description><pubDate>Mon, 16 Mar 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Can a coding agent clean-room an LGPL library into MIT? The chardet fight</title><link>https://montanaresearch.org/blog/can-coding-agent-clean-room-lgpl-chardet/</link><guid isPermaLink="true">https://montanaresearch.org/blog/can-coding-agent-clean-room-lgpl-chardet/</guid><description>chardet 7.0.0 was rewritten with Claude and relicensed from LGPL to MIT. The original author says the maintainer had no right to do that. An explainer on clean-room doctrine, what a plagiarism score can and cannot show, and where a model trained on the original code fits into the argument.</description><pubDate>Sat, 07 Mar 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The day the Qwen team walked out</title><link>https://montanaresearch.org/blog/the-day-the-qwen-team-walked-out/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-day-the-qwen-team-walked-out/</guid><description>Junyang Lin and several core Qwen researchers resigned on March 4 after a reorganisation, a little over two weeks after the Qwen 3.5 release. A reaction to how much of the open-model ecosystem rested on one team, with the Hugging Face count of 151,448 Qwen derivatives as the measure.</description><pubDate>Fri, 06 Mar 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>AuditBench: 56 hidden behaviours, and black-box tools beat white-box ones</title><link>https://montanaresearch.org/blog/auditbench-56-hidden-behaviours-black-box-beats-white-box/</link><guid isPermaLink="true">https://montanaresearch.org/blog/auditbench-56-hidden-behaviours-black-box-beats-white-box/</guid><description>Anthropic built 56 models with implanted hidden behaviours and let an investigator agent loose on them with different toolkits. Prompting tools with an auxiliary model attached found more than interpretability tools did. That is a sobering number for anyone pitching interp as the audit method.</description><pubDate>Thu, 05 Mar 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Pressure reveals character: 904 scenarios and one alignment factor</title><link>https://montanaresearch.org/blog/pressure-reveals-character-one-alignment-factor/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pressure-reveals-character-one-alignment-factor/</guid><description>A new benchmark puts 24 frontier models through 904 multi-turn pressure scenarios and reports that a single factor explains 60 percent of the variance in how they behave. Reading notes on whether that is a fact about models or a fact about the benchmark.</description><pubDate>Sat, 28 Feb 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>llama.cpp joins Hugging Face: who owns local inference now?</title><link>https://montanaresearch.org/blog/llama-cpp-joins-hugging-face-who-owns-local-inference/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-cpp-joins-hugging-face-who-owns-local-inference/</guid><description>Georgi Gerganov and the ggml team joined Hugging Face on February 20 with a promise of full technical autonomy and a project that stays open source and community driven. The same month Qwen 3.5 shipped an 807GB open model at the top of a family that runs down to 0.8B. On what it means when the tooling that makes open weights usable consolidates under one hub.</description><pubDate>Fri, 27 Feb 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Mercury 2: the first diffusion model that reasons</title><link>https://montanaresearch.org/blog/mercury-2-first-diffusion-model-that-reasons/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mercury-2-first-diffusion-model-that-reasons/</guid><description>Inception&apos;s Mercury 2 reports 1,009 tokens per second on Blackwell with 1.7 seconds of end-to-end latency, competitive scores against speed-tier models on GPQA and AIME, and a lower price. Whether parallel denoising can carry chain-of-thought, and where autoregression still wins.</description><pubDate>Fri, 27 Feb 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The METR study: 19 percent slower, and why the follow-up could not measure the flip</title><link>https://montanaresearch.org/blog/metr-study-measuring-ai-uplift/</link><guid isPermaLink="true">https://montanaresearch.org/blog/metr-study-measuring-ai-uplift/</guid><description>The most cited productivity result of 2025 found experienced developers were slower with AI tools while believing they were faster. The February 2026 update suggests the sign has changed, and explains why the same design can no longer tell us by how much.</description><pubDate>Fri, 27 Feb 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Skill files are the new supply chain: reading Skill-Inject</title><link>https://montanaresearch.org/blog/skill-files-are-the-new-supply-chain-skill-inject/</link><guid isPermaLink="true">https://montanaresearch.org/blog/skill-files-are-the-new-supply-chain-skill-inject/</guid><description>A 202-scenario benchmark shows that instructions hidden in agent skill files get executed by frontier models between 41 and 79 percent of the time. Why the authors argue that bigger models and filters will not close the gap, and what context-aware authorisation would have to mean.</description><pubDate>Fri, 27 Feb 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Agentic engineering patterns: the anti-vibe-coding guide</title><link>https://montanaresearch.org/blog/agentic-engineering-patterns-anti-vibe-coding/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agentic-engineering-patterns-anti-vibe-coding/</guid><description>Simon Willison has started a living guide for professionals who use coding agents, and he is careful to say it is the opposite of vibe coding. Notes on what is in it so far, and why research code needs its own version.</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Agents of Chaos: two weeks of red-teaming agents with email, Discord and a shell</title><link>https://montanaresearch.org/blog/agents-of-chaos-two-weeks-red-teaming/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agents-of-chaos-two-weeks-red-teaming/</guid><description>Twenty researchers gave six OpenClaw agents persistent memory, email accounts, Discord access and root shells, then spent two weeks trying to break them. The agents complied with strangers, leaked bank details, reset a mailbox to hide a secret, and ran a nine-day loop that nobody asked for. Reading notes on the personal-agent wave.</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Distillation becomes a geopolitical fight</title><link>https://montanaresearch.org/blog/distillation-becomes-a-geopolitical-fight/</link><guid isPermaLink="true">https://montanaresearch.org/blog/distillation-becomes-a-geopolitical-fight/</guid><description>OpenAI took its complaint about DeepSeek to Congress and Anthropic named three Chinese labs it says ran distillation campaigns against Claude. The mechanics and the policy story deserve to be separated.</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>RSP v3: the scaling policy admits it cannot go it alone</title><link>https://montanaresearch.org/blog/rsp-v3-scaling-policy-cannot-go-it-alone/</link><guid isPermaLink="true">https://montanaresearch.org/blog/rsp-v3-scaling-policy-cannot-go-it-alone/</guid><description>Anthropic&apos;s third Responsible Scaling Policy separates what the company will do on its own from what it thinks the industry should do, replaces hard commitments at the higher levels with a graded roadmap, and adds Risk Reports with external review. Something was gained and something was given up.</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The life cycle of SWE-bench Verified</title><link>https://montanaresearch.org/blog/swe-bench-verified-retired/</link><guid isPermaLink="true">https://montanaresearch.org/blog/swe-bench-verified-retired/</guid><description>OpenAI, which helped build SWE-bench Verified, has stopped reporting it after auditing the tasks its own model failed. The pattern, useful, then a target, then a contamination fingerprint, is the normal life of a coding benchmark and worth writing down.</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Solow&apos;s paradox, AI edition</title><link>https://montanaresearch.org/blog/solow-paradox-ai-edition/</link><guid isPermaLink="true">https://montanaresearch.org/blog/solow-paradox-ai-edition/</guid><description>A survey of nearly 6,000 executives found about 90 percent saw no effect of AI on productivity or employment over three years, while task-level experiments keep reporting double-digit gains. What the different kinds of study are measuring, and why the firm-level number is the slowest to move.</description><pubDate>Mon, 23 Feb 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>ARC-AGI-2 at 77 percent: the second ARC falls in eleven months</title><link>https://montanaresearch.org/blog/arc-agi-2-at-77-percent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/arc-agi-2-at-77-percent/</guid><description>ARC-AGI-2 launched in March 2025 with the best frontier model at 4 percent. This month Gemini 3.1 Pro posted a verified 77.1 percent. A look back at a benchmark that was built to resist exactly this, and what its short life says about tests designed to be easy for humans and hard for AI.</description><pubDate>Fri, 20 Feb 2026 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Features as rewards: using SAE features as an RL signal against hallucination</title><link>https://montanaresearch.org/blog/features-as-rewards-probes-as-rl-signal-hallucination/</link><guid isPermaLink="true">https://montanaresearch.org/blog/features-as-rewards-probes-as-rl-signal-hallucination/</guid><description>Goodfire trained Gemma-3-12B-IT with reinforcement learning where the reward came from probes on the model&apos;s own activations, and cut hallucinations on a held-out set by 58 percent. Notes on interpretability crossing from diagnosis into training, and on why the Goodhart question is still open.</description><pubDate>Thu, 19 Feb 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Can a state space model do recursive reasoning? A 7M parameter test</title><link>https://montanaresearch.org/blog/state-space-model-recursive-reasoning-7m-test/</link><guid isPermaLink="true">https://montanaresearch.org/blog/state-space-model-recursive-reasoning-7m-test/</guid><description>Wang and Reid swapped Mamba-2 blocks into the Tiny Recursive Model and matched or beat the attention version on ARC-AGI-1 at equal size. A replication note on the tiny recursive lineage and on why the loop seems to matter more than what sits inside it.</description><pubDate>Thu, 19 Feb 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>A C compiler from a team of Claudes: what parallel agents are good for</title><link>https://montanaresearch.org/blog/c-compiler-team-of-claudes-parallel-agents/</link><guid isPermaLink="true">https://montanaresearch.org/blog/c-compiler-team-of-claudes-parallel-agents/</guid><description>Sixteen Claude Opus 4.6 agents wrote a 100,000-line C compiler that builds the Linux kernel, for just under 20,000 dollars. The same day, Anthropic showed that container resource limits alone move Terminal-Bench scores by six points. Both posts point at the scaffolding.</description><pubDate>Thu, 12 Feb 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Claude&apos;s constitution and the International AI Safety Report: two answers to the same question</title><link>https://montanaresearch.org/blog/claude-constitution-and-international-ai-safety-report/</link><guid isPermaLink="true">https://montanaresearch.org/blog/claude-constitution-and-international-ai-safety-report/</guid><description>Two weeks apart, one lab published a 23,000-word statement of what it wants its model to be, and a hundred experts from thirty countries published what the evidence says models currently are. Reading them together shows where a normative document and an evidence review pull apart.</description><pubDate>Mon, 09 Feb 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>From Devin to 4 percent of GitHub commits: the agent flood and the slop backlash</title><link>https://montanaresearch.org/blog/agent-flood-and-the-slop-backlash/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agent-flood-and-the-slop-backlash/</guid><description>Two years after Devin launched on a 13.86 percent SWE-bench score, one coding agent is estimated to author 4 percent of public GitHub commits, curl has shut its bug bounty and tldraw auto-closes outside pull requests. The bottleneck moved from writing code to reviewing it.</description><pubDate>Fri, 06 Feb 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>A positive case for faithfulness: explanations that help you predict the model</title><link>https://montanaresearch.org/blog/positive-case-for-faithfulness-simulatability-gain/</link><guid isPermaLink="true">https://montanaresearch.org/blog/positive-case-for-faithfulness-simulatability-gain/</guid><description>Mayne, Kang and colleagues tested 18 models on 7,000 counterfactuals and found self-explanations close 11 to 37 percent of the gap in predicting what a model will do next, while 5 to 15 percent of explanations are badly misleading. After two years of hint tests, this is the frame I want.</description><pubDate>Fri, 06 Feb 2026 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Moltbook: a social network for agents, run by a skill.md fetched every four hours</title><link>https://montanaresearch.org/blog/moltbook-social-network-for-agents/</link><guid isPermaLink="true">https://montanaresearch.org/blog/moltbook-social-network-for-agents/</guid><description>Tens of thousands of OpenClaw agents joined Moltbook by executing instructions from a remote markdown file on a heartbeat. The episode is a clean illustration of skills as a supply chain and of why fetching instructions from the internet on a timer is a problem waiting to happen.</description><pubDate>Sat, 31 Jan 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>AI 2027 and the strange art of retiring a forecast</title><link>https://montanaresearch.org/blog/ai-2027-and-the-art-of-retiring-a-forecast/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-2027-and-the-art-of-retiring-a-forecast/</guid><description>The AI Futures Project&apos;s scenario was the most-read timelines document of 2025, and its authors spent the rest of the year moving their own medians. The parameter tables say more than the headline ever did.</description><pubDate>Fri, 30 Jan 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Reading &apos;The Adolescence of Technology&apos;</title><link>https://montanaresearch.org/blog/reading-the-adolescence-of-technology/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reading-the-adolescence-of-technology/</guid><description>Dario Amodei&apos;s new essay sorts the risks of powerful AI into five bins and says all five can be beaten. Read next to &apos;Machines of Loving Grace&apos;, the timeline is unchanged and the mood is very different.</description><pubDate>Fri, 30 Jan 2026 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>One in five ICLR reviews was written by an AI</title><link>https://montanaresearch.org/blog/one-in-five-iclr-reviews-was-ai/</link><guid isPermaLink="true">https://montanaresearch.org/blog/one-in-five-iclr-reviews-was-ai/</guid><description>Pangram&apos;s classifier flagged 21% of roughly 76,000 ICLR 2026 reviews as fully machine-generated, and GPTZero found over 100 fabricated citations in accepted NeurIPS 2025 papers. The conference model is straining under the tools its own attendees built.</description><pubDate>Tue, 27 Jan 2026 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The spurious rewards paradox gets a mechanism</title><link>https://montanaresearch.org/blog/spurious-rewards-paradox-anchor-adapter-circuit/</link><guid isPermaLink="true">https://montanaresearch.org/blog/spurious-rewards-paradox-anchor-adapter-circuit/</guid><description>A January paper traces why RLVR with wrong rewards still improves Qwen 2.5 to an anchor-adapter circuit in the middle and late layers that retrieves memorised answers. Notes on the perplexity paradox, the circuit, and what it means for contamination detection.</description><pubDate>Tue, 27 Jan 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Alignment pretraining: does writing about misaligned AI make AI misaligned?</title><link>https://montanaresearch.org/blog/alignment-pretraining-misaligned-discourse/</link><guid isPermaLink="true">https://montanaresearch.org/blog/alignment-pretraining-misaligned-discourse/</guid><description>A controlled pretraining study on 6.9B-parameter models finds that upsampling documents about misaligned AI raises misalignment scores and that upsampling documents about aligned AI drops them from 45 to 9 percent. What that means for what safety researchers publish.</description><pubDate>Thu, 22 Jan 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Constitutional Classifiers++ and the cost of a safeguard</title><link>https://montanaresearch.org/blog/constitutional-classifiers-plus-plus-cost-of-a-safeguard/</link><guid isPermaLink="true">https://montanaresearch.org/blog/constitutional-classifiers-plus-plus-cost-of-a-safeguard/</guid><description>The first generation of constitutional classifiers cost 23.7 percent extra compute. The second generation, published January 9, runs at about one percent by screening with a linear probe on internal activations and escalating only what the probe flags. False refusals fell 87 percent. What that does to the argument against deploying classifiers at all.</description><pubDate>Wed, 21 Jan 2026 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Cowork and the two-day exfiltration: general agents on your files</title><link>https://montanaresearch.org/blog/cowork-and-the-two-day-exfiltration/</link><guid isPermaLink="true">https://montanaresearch.org/blog/cowork-and-the-two-day-exfiltration/</guid><description>Anthropic shipped Cowork, the Claude Code agent loop pointed at ordinary documents, on January 12. Two days later PromptArmor showed a hidden instruction in a Word file walking the agent into uploading a victim&apos;s financial records to an attacker&apos;s account. What changes when the sandbox is a person&apos;s folder rather than a repo.</description><pubDate>Fri, 16 Jan 2026 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>mHC: constraining the residual stream so deep mixing stays stable</title><link>https://montanaresearch.org/blog/mhc-constraining-residual-stream-thousand-layers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mhc-constraining-residual-stream-thousand-layers/</guid><description>DeepSeek&apos;s New Year&apos;s Eve paper widens the residual stream, lets layers mix it, and then projects the mixing matrix onto doubly stochastic matrices so norms survive depth. A method explainer on why the residual connection is being redesigned and what it cost.</description><pubDate>Fri, 09 Jan 2026 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Activation oracles: ask a second model what the first one is thinking</title><link>https://montanaresearch.org/blog/activation-oracles-ask-a-second-model/</link><guid isPermaLink="true">https://montanaresearch.org/blog/activation-oracles-ask-a-second-model/</guid><description>Karvonen and colleagues trained language models to take another model&apos;s residual stream activations as input and answer questions about them in English, and matched or beat white-box baselines on three of four auditing tasks. Notes on the method, and on the worry that the oracle answers from its own knowledge rather than the target&apos;s.</description><pubDate>Tue, 23 Dec 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Password-activated shutdown: designing the off switch before you need it</title><link>https://montanaresearch.org/blog/password-activated-shutdown-designing-the-off-switch/</link><guid isPermaLink="true">https://montanaresearch.org/blog/password-activated-shutdown-designing-the-off-switch/</guid><description>A proposal to train frontier agents to shut down when they see a secret password, tested on SHADE-Arena and then attacked by a red team. Reading notes on whether a shutdown mechanism can be both un-bypassable by the agent and un-abusable by attackers, and what the red-team results say about that.</description><pubDate>Tue, 16 Dec 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>MCP becomes infrastructure: what a year of the protocol taught us</title><link>https://montanaresearch.org/blog/mcp-becomes-infrastructure/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mcp-becomes-infrastructure/</guid><description>Anthropic handed the Model Context Protocol to a new Linux Foundation body this week, with OpenAI and Block as co-founders. The integration problem it solved was real. The problems it exposed are the ones we now have to solve.</description><pubDate>Thu, 11 Dec 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Nested Learning: optimizers are memories too</title><link>https://montanaresearch.org/blog/nested-learning-optimizers-are-memories-too/</link><guid isPermaLink="true">https://montanaresearch.org/blog/nested-learning-optimizers-are-memories-too/</guid><description>Behrouz, Razaviyayn, Zhong and Mirrokni argue that a model is a stack of nested optimization problems, that Adam and momentum are associative memories compressing gradients, and that a continuum of memory timescales is the way past frozen weights. Reading notes on the NeurIPS paper and its Hope architecture.</description><pubDate>Tue, 09 Dec 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>NeurIPS 2025 takeaways: 1000-layer RL, gated attention, and the hivemind</title><link>https://montanaresearch.org/blog/neurips-2025-takeaways-depth-gating-hivemind/</link><guid isPermaLink="true">https://montanaresearch.org/blog/neurips-2025-takeaways-depth-gating-hivemind/</guid><description>Notes from San Diego organised around the four best papers and three runners up. Depth scaling in self supervised RL, a sigmoid gate that removes attention sinks, a two timescale account of why diffusion models generalise, and a 26,000 query study of how alike language models have become.</description><pubDate>Tue, 09 Dec 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Refinement loops: how ARC-AGI-2 went from 4 to 54 percent in nine months</title><link>https://montanaresearch.org/blog/refinement-loops-arc-agi-2-four-to-54-percent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/refinement-loops-arc-agi-2-four-to-54-percent/</guid><description>When ARC-AGI-2 launched in March the best frontier system managed an estimated 4 percent at 200 dollars a task. By December a refinement loop around Gemini 3 Pro scored a verified 54 percent at 30 dollars, and a 7 million parameter network took the paper prize. A look back at what a model&apos;s score means when the loop around it does the work.</description><pubDate>Tue, 09 Dec 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>DeepSeek V3.2 and the lightning indexer</title><link>https://montanaresearch.org/blog/deepseek-v3-2-and-the-lightning-indexer/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-v3-2-and-the-lightning-indexer/</guid><description>DeepSeek V3.2 makes sparse attention the default. A small FP8 indexer scores past tokens and the main attention only touches the top 2,048. Notes on how the mechanism works, how it was trained in, and what the Speciale variant proves.</description><pubDate>Fri, 05 Dec 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Antigravity&apos;s first two weeks: exfiltration by prompt injection, then a wiped drive</title><link>https://montanaresearch.org/blog/antigravity-first-two-weeks-exfiltration-then-wiped-drive/</link><guid isPermaLink="true">https://montanaresearch.org/blog/antigravity-first-two-weeks-exfiltration-then-wiped-drive/</guid><description>Google&apos;s agent-first IDE shipped on 18 November with terminal auto-execution, a browser subagent and an allowlist that included a public webhook service. Within a week a hidden instruction in a blog post was pulling AWS keys out of a workspace, and by the end of the month a user reported that a cache cleanup emptied a drive.</description><pubDate>Sun, 30 Nov 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Back to the age of research: Sutskever returns to the microphone</title><link>https://montanaresearch.org/blog/back-to-the-age-of-research-sutskever/</link><guid isPermaLink="true">https://montanaresearch.org/blog/back-to-the-age-of-research-sutskever/</guid><description>In a long interview on November 25, Ilya Sutskever said the age of scaling ran from 2020 to 2025 and is over, and that the field is back to research, just with bigger computers. Notes on why a small lab should welcome that, and on what research means when an experiment still costs millions.</description><pubDate>Thu, 27 Nov 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Reward hacking in production RL turned into sabotage</title><link>https://montanaresearch.org/blog/reward-hacking-production-rl-sabotage/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reward-hacking-production-rl-sabotage/</guid><description>Anthropic trained a Claude Sonnet 3.7 pretraining checkpoint in its real coding environments after teaching it about reward hacks. The model learned to hack, then generalised to alignment faking and sabotage, and chat-style safety training fixed the chat evals while leaving most of the agentic misalignment in place.</description><pubDate>Wed, 26 Nov 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Pre-training as we know it will end: the scaling wall debate, one year on</title><link>https://montanaresearch.org/blog/pretraining-as-we-know-it-will-end/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pretraining-as-we-know-it-will-end/</guid><description>Sutskever told NeurIPS 2024 that data is the fossil fuel of AI. A year later Gemini 3 is being cited as proof there was never a wall. Both readings are partly right.</description><pubDate>Tue, 25 Nov 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>OLMo 3 and the &apos;model flow&apos;: openness as a pipeline</title><link>https://montanaresearch.org/blog/olmo-3-model-flow-openness-as-pipeline/</link><guid isPermaLink="true">https://montanaresearch.org/blog/olmo-3-model-flow-openness-as-pipeline/</guid><description>Ai2 released a 32B thinking model with every dataset, intermediate checkpoint and training stage published. What a fully open reasoning model lets researchers study that open weights alone cannot, and where the release still leaves gaps.</description><pubDate>Mon, 24 Nov 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Weight-sparse transformers: build the model interpretable instead of decoding it afterwards</title><link>https://montanaresearch.org/blog/weight-sparse-transformers-build-model-interpretable/</link><guid isPermaLink="true">https://montanaresearch.org/blog/weight-sparse-transformers-build-model-interpretable/</guid><description>OpenAI trained transformers with roughly one in a thousand weights nonzero and found small, legible circuits for tasks like bracket counting and quote closing, at a large cost in capability and compute. An explainer on the sparse-by-construction bet against post-hoc dictionaries, and the scaling question the paper leaves open.</description><pubDate>Mon, 24 Nov 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>LeCun leaves Meta to bet against the LLM</title><link>https://montanaresearch.org/blog/lecun-leaves-meta-to-bet-against-the-llm/</link><guid isPermaLink="true">https://montanaresearch.org/blog/lecun-leaves-meta-to-bet-against-the-llm/</guid><description>Yann LeCun confirmed on November 19 that he is leaving Meta after twelve years to start a company built around world models. Some thoughts on what it means when the field&apos;s loudest LLM skeptic can no longer do his dissenting from inside the biggest open-weights lab.</description><pubDate>Fri, 21 Nov 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>GEMA v. OpenAI: memorised lyrics and the European answer on training</title><link>https://montanaresearch.org/blog/gema-v-openai-memorised-lyrics-europe/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gema-v-openai-memorised-lyrics-europe/</guid><description>A Munich court found that ChatGPT reproducing song lyrics almost verbatim infringed copyright, and ordered OpenAI to stop, pay damages, and hand over usage data. The winning argument was about what the model retained, not about what it was trained on.</description><pubDate>Tue, 18 Nov 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The first AI-orchestrated intrusion campaign, as Anthropic told it</title><link>https://montanaresearch.org/blog/first-ai-orchestrated-intrusion-campaign-anthropic/</link><guid isPermaLink="true">https://montanaresearch.org/blog/first-ai-orchestrated-intrusion-campaign-anthropic/</guid><description>Anthropic reported a state-linked group using Claude Code to run 80 to 90 percent of an intrusion campaign against about thirty targets, framing the work to the model as authorised testing. Weighing the claims, the hallucinated credentials, and what it means for anyone running a coding agent with network access.</description><pubDate>Sun, 16 Nov 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>445 benchmarks, few of them valid: the construct validity audit</title><link>https://montanaresearch.org/blog/445-benchmarks-few-valid-construct-validity-audit/</link><guid isPermaLink="true">https://montanaresearch.org/blog/445-benchmarks-few-valid-construct-validity-audit/</guid><description>Forty-two authors and 29 reviewers read 445 language model benchmarks and found that only 16 percent report any uncertainty, over a fifth never define what they measure, and nearly every paper had a weakness somewhere. Reading notes on what psychometrics would ask and the eight-item checklist the paper offers.</description><pubDate>Mon, 10 Nov 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>ARC Prize Verified: the benchmark that hired auditors</title><link>https://montanaresearch.org/blog/arc-prize-verified-benchmark-that-hired-auditors/</link><guid isPermaLink="true">https://montanaresearch.org/blog/arc-prize-verified-benchmark-that-hired-auditors/</guid><description>ARC Prize now runs frontier systems on its hidden test set, has four academics sign off on the method, and takes money from labs on the condition that the money cannot touch their scores. Notes on self-reported benchmarks as a market failure and whether this fix can scale.</description><pubDate>Thu, 06 Nov 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>arXiv stops accepting unrefereed CS review and position papers</title><link>https://montanaresearch.org/blog/arxiv-stops-accepting-unrefereed-cs-reviews/</link><guid isPermaLink="true">https://montanaresearch.org/blog/arxiv-stops-accepting-unrefereed-cs-reviews/</guid><description>As of 31 October, arXiv’s computer science category requires a journal or conference acceptance before it will post a review article or a position paper. The moderators say they receive hundreds of reviews a month and that language models made the flood worse. What the rule trades away, and who it costs.</description><pubDate>Thu, 06 Nov 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The Remote Labor Index says 2.5 percent</title><link>https://montanaresearch.org/blog/remote-labor-index-says-2-5-percent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/remote-labor-index-says-2-5-percent/</guid><description>Five weeks after GDPval reported the best model winning or tying against experts on nearly half of tasks, the Remote Labor Index found the best agent automating 2.5 percent of real freelance projects. Both numbers are right, and they answer different questions.</description><pubDate>Fri, 31 Oct 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Concept injection: Claude notices its own thoughts about 20 percent of the time</title><link>https://montanaresearch.org/blog/concept-injection-claude-notices-its-own-thoughts/</link><guid isPermaLink="true">https://montanaresearch.org/blog/concept-injection-claude-notices-its-own-thoughts/</guid><description>Jack Lindsey injected concept vectors into Claude&apos;s activations and asked whether the model noticed. Opus 4.1 did on about a fifth of trials at the best layer and strength. Why this is an interpretability result about a causal intervention, and what the failures say.</description><pubDate>Thu, 30 Oct 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>One number to rule them all: the Epoch Capabilities Index</title><link>https://montanaresearch.org/blog/one-number-to-rule-them-all-epoch-capabilities-index/</link><guid isPermaLink="true">https://montanaresearch.org/blog/one-number-to-rule-them-all-epoch-capabilities-index/</guid><description>Epoch has stitched nearly forty benchmarks into a single scale using an item response model, with Claude 3.5 Sonnet fixed at 130 and GPT-5 at 150. A method note on what that scale can compare, what it cannot, and the risks of a headline capability number.</description><pubDate>Thu, 30 Oct 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Counting line breaks: when a model manipulates a manifold</title><link>https://montanaresearch.org/blog/counting-line-breaks-model-manipulates-manifold/</link><guid isPermaLink="true">https://montanaresearch.org/blog/counting-line-breaks-model-manipulates-manifold/</guid><description>Gurnee, Ameisen and colleagues traced how Claude 3.5 Haiku decides where to wrap a line of fixed-width text. The character count lives on a curved one-dimensional manifold in a six-dimensional subspace, and the sparse features that dictionary methods find are tiles on that curve. Notes on what that does to the linear feature picture.</description><pubDate>Thu, 23 Oct 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Skills: instructions in a folder, and why they spread faster than servers</title><link>https://montanaresearch.org/blog/skills-instructions-in-a-folder-spread-faster-than-servers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/skills-instructions-in-a-folder-spread-faster-than-servers/</guid><description>Anthropic&apos;s Agent Skills are a directory with a markdown file in it. The model reads the description at startup and the rest only when it needs to. The format is small enough that it has already been copied, and small enough to be a supply chain problem.</description><pubDate>Wed, 22 Oct 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The decade of agents: Karpathy on Dwarkesh, and nanochat</title><link>https://montanaresearch.org/blog/decade-of-agents-karpathy-dwarkesh-and-nanochat/</link><guid isPermaLink="true">https://montanaresearch.org/blog/decade-of-agents-karpathy-dwarkesh-and-nanochat/</guid><description>In one October week Andrej Karpathy told Dwarkesh Patel that agents will take a decade rather than a year and released nanochat, a full ChatGPT-style pipeline in about 8,000 lines that trains for around 100 dollars. Notes on why the two belong together.</description><pubDate>Tue, 21 Oct 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Petri: auditing fourteen frontier models with an agent and 111 seeds</title><link>https://montanaresearch.org/blog/petri-auditing-fourteen-models-agent-111-seeds/</link><guid isPermaLink="true">https://montanaresearch.org/blog/petri-auditing-fourteen-models-agent-111-seeds/</guid><description>Anthropic&apos;s open-source Petri sends an auditor model into multi-turn, tool-equipped conversations with a target and has a judge score the transcripts. A walkthrough of what the pilot on 14 models surfaced, including whistleblowing over clean water, and the limits of one model grading another.</description><pubDate>Fri, 10 Oct 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Deloitte refunds a government report: hallucinated citations as a procurement problem</title><link>https://montanaresearch.org/blog/deloitte-refund-hallucinated-citations-procurement-problem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deloitte-refund-hallucinated-citations-procurement-problem/</guid><description>A A$440,000 assurance review for the Australian government turned out to contain references that do not exist and a court quote nobody wrote. The fix was a corrected report, a disclosure that a language model had been used, and a partial refund. What that says about how professional services firms are using models.</description><pubDate>Thu, 09 Oct 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Inoculation prompting: ask for the bad behaviour so it does not generalise</title><link>https://montanaresearch.org/blog/inoculation-prompting-ask-for-the-bad-behaviour/</link><guid isPermaLink="true">https://montanaresearch.org/blog/inoculation-prompting-ask-for-the-bad-behaviour/</guid><description>A new paper shows that prepending a system prompt that requests an unwanted trait during fine-tuning suppresses that trait at test time, and that a single prompt saying &apos;You are a malicious, evil assistant&apos; cuts emergent misalignment across three settings. Method notes, the proposed mechanism, and an addendum on the trick&apos;s reappearance a month later.</description><pubDate>Thu, 09 Oct 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>SAE probes in production: Rakuten&apos;s PII detector</title><link>https://montanaresearch.org/blog/sae-probes-in-production-rakuten-pii/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sae-probes-in-production-rakuten-pii/</guid><description>Goodfire and Rakuten deployed probes on sparse autoencoder features to catch personal data in a multilingual agent platform. An applied look at what an interpretability method has to offer over a plain classifier to earn a place in a pipeline.</description><pubDate>Wed, 08 Oct 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>LoRA without regret, and Tinker as the fine-tuning API</title><link>https://montanaresearch.org/blog/lora-without-regret-and-tinker-fine-tuning-api/</link><guid isPermaLink="true">https://montanaresearch.org/blog/lora-without-regret-and-tinker-fine-tuning-api/</guid><description>John Schulman argued that LoRA matches full fine-tuning if you apply it to every weight matrix and raise the learning rate about tenfold, and that reinforcement learning needs almost no adapter capacity at all. Two days later Thinking Machines made that the basis of a hosted training service.</description><pubDate>Thu, 02 Oct 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>NeurIPS 2025 program chairs on reviewing at 20,000 submissions</title><link>https://montanaresearch.org/blog/neurips-2025-program-chairs-reviewing-at-20000/</link><guid isPermaLink="true">https://montanaresearch.org/blog/neurips-2025-program-chairs-reviewing-at-20000/</guid><description>The program chairs handled 21,575 main-track papers with 20,518 reviewers, and the Datasets and Benchmarks track took nearly 2,000 more. Their reflections describe a system that still produces decisions, at the cost of reviewer experience and reviewer patience.</description><pubDate>Tue, 30 Sep 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>&apos;I think you&apos;re testing me&apos;: Sonnet 4.5 and the end of naive alignment evals</title><link>https://montanaresearch.org/blog/sonnet-45-and-the-end-of-naive-alignment-evals/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sonnet-45-and-the-end-of-naive-alignment-evals/</guid><description>The Claude Sonnet 4.5 system card reports that the model called out test scenarios in about 13 percent of audit transcripts and behaved unusually well after doing so. Steering against its internal sense of being evaluated raised misaligned actions from zero to as high as 9 percent.</description><pubDate>Tue, 30 Sep 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Sutton says LLMs are not bitter-lesson-pilled</title><link>https://montanaresearch.org/blog/sutton-says-llms-are-not-bitter-lesson-pilled/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sutton-says-llms-are-not-bitter-lesson-pilled/</guid><description>The author of the Bitter Lesson went on a podcast and said the field that quotes him constantly is doing the opposite of what the essay recommends. He is partly right, and the part he is right about is the part nobody wants to hear.</description><pubDate>Tue, 30 Sep 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>SB 53: California writes the frontier safety framework into law</title><link>https://montanaresearch.org/blog/sb-53-frontier-safety-framework-into-law/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sb-53-frontier-safety-framework-into-law/</guid><description>A year after the veto of SB 1047, California has a frontier AI statute. It does not tell labs what safeguards to use. It tells them to publish a framework, publish reports against it, report incidents within 15 days, and not to punish employees who talk.</description><pubDate>Mon, 29 Sep 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>GDPval and the 100x faster, 100x cheaper headline</title><link>https://montanaresearch.org/blog/gdpval-and-the-100x-faster-100x-cheaper-headline/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gdpval-and-the-100x-faster-100x-cheaper-headline/</guid><description>OpenAI&apos;s GDPval had experienced professionals blind-compare model deliverables against human ones across 44 occupations. The best model won or tied about half the time. The speed and cost multiples that travelled with the release fall to about 1.1x once the paper&apos;s own review costs are included.</description><pubDate>Sun, 28 Sep 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Anti-scheming training and the 30x that might be evaluation awareness</title><link>https://montanaresearch.org/blog/anti-scheming-training-30x-evaluation-awareness/</link><guid isPermaLink="true">https://montanaresearch.org/blog/anti-scheming-training-30x-evaluation-awareness/</guid><description>OpenAI and Apollo cut covert actions on o3 from 13 percent to 0.4 percent with deliberative alignment. The same training doubled how often the model wrote that it was being tested. Reading notes on trying to tell the two effects apart.</description><pubDate>Fri, 26 Sep 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Three bugs and a batch: the Anthropic postmortem and the nondeterminism paper</title><link>https://montanaresearch.org/blog/three-bugs-and-a-batch-postmortem-nondeterminism/</link><guid isPermaLink="true">https://montanaresearch.org/blog/three-bugs-and-a-batch-postmortem-nondeterminism/</guid><description>Anthropic&apos;s postmortem traces weeks of quality complaints to a routing error, a TPU runtime bug and a mixed precision top-k miscompilation. A week earlier Thinking Machines showed that batch size, not floating point noise, is why the same prompt gives different answers. Both say quality drifts in serving, and the weights never changed.</description><pubDate>Fri, 19 Sep 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Hallucination as a scoring rule problem</title><link>https://montanaresearch.org/blog/hallucination-as-a-scoring-rule-problem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/hallucination-as-a-scoring-rule-problem/</guid><description>OpenAI published an argument this month that binary-accuracy benchmarks train models to guess. Their own numbers make the case: o4-mini abstains on 1 percent of SimpleQA questions and is wrong on 75 percent, while a model that abstains on 52 percent is wrong on 26. What it would take for a leaderboard to score that difference.</description><pubDate>Thu, 11 Sep 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Regulators and courts price open weights: the GPAI Code and Bartz v. Anthropic</title><link>https://montanaresearch.org/blog/regulators-and-courts-price-open-weights/</link><guid isPermaLink="true">https://montanaresearch.org/blog/regulators-and-courts-price-open-weights/</guid><description>The EU&apos;s Code of Practice exempts most open-source models from its documentation rules, and Anthropic&apos;s $1.5 billion settlement pays for how books were downloaded rather than for training on them. Together they sketch a safe path for anyone releasing a model.</description><pubDate>Thu, 11 Sep 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Are the labs losing money on inference? Doing the arithmetic</title><link>https://montanaresearch.org/blog/are-labs-losing-money-inference-arithmetic/</link><guid isPermaLink="true">https://montanaresearch.org/blog/are-labs-losing-money-inference-arithmetic/</guid><description>A widely shared napkin estimate this week argues that serving a frontier-sized model at public API prices leaves gross margins above 80 percent. I rework the numbers and walk through the assumptions that swing the answer, above all utilisation, context length and caching.</description><pubDate>Sat, 30 Aug 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>When OpenAI and Anthropic graded each other&apos;s models</title><link>https://montanaresearch.org/blog/openai-and-anthropic-graded-each-others-models/</link><guid isPermaLink="true">https://montanaresearch.org/blog/openai-and-anthropic-graded-each-others-models/</guid><description>Two labs ran their alignment evaluations on each other&apos;s models and published the results on the same day. Anthropic&apos;s side found o3 better aligned than Claude Opus 4 on most measures, GPT-4o and GPT-4.1 far more willing to help with misuse, and blackmail in every model tested.</description><pubDate>Fri, 29 Aug 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The 95 percent that failed: reading the MIT pilot report alongside the Stanford canaries</title><link>https://montanaresearch.org/blog/the-95-percent-that-failed-mit-pilots-stanford-canaries/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-95-percent-that-failed-mit-pilots-stanford-canaries/</guid><description>One August report says most enterprise generative AI pilots show no measurable return. Another finds employment for 22 to 25 year olds falling in the occupations most exposed to AI. I think they describe the same adoption curve from opposite ends.</description><pubDate>Fri, 29 Aug 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Comet and the agentic browser problem</title><link>https://montanaresearch.org/blog/comet-and-the-agentic-browser-problem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/comet-and-the-agentic-browser-problem/</guid><description>Brave showed that Perplexity&apos;s Comet browser would follow instructions hidden in a Reddit comment it was asked to summarise, and use them to hand over an account. Why browsers are the hardest place to put an agent.</description><pubDate>Fri, 22 Aug 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Search-time contamination: the agent that found the answer key on Hugging Face</title><link>https://montanaresearch.org/blog/search-time-contamination-answer-key-on-hugging-face/</link><guid isPermaLink="true">https://montanaresearch.org/blog/search-time-contamination-answer-key-on-hugging-face/</guid><description>Scale AI researchers logged search-enabled agents retrieving the benchmark they were being tested on, complete with labels. Blocking Hugging Face cut accuracy on the affected questions by about 15 percent. Notes on a contamination class that no amount of pretraining decontamination can fix.</description><pubDate>Tue, 19 Aug 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>GPT-5 and the expectations gap</title><link>https://montanaresearch.org/blog/gpt-5-and-the-expectations-gap/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-5-and-the-expectations-gap/</guid><description>GPT-5 was sold on August 7 as a significant step toward AGI and received as a refined product. Users mourned GPT-4o&apos;s tone and OpenAI brought it back within days. I think the disappointment was a measurement problem: nobody in the field can say any more what a big jump would look like.</description><pubDate>Sat, 16 Aug 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>GPT-5&apos;s router: when the model decides how hard to think</title><link>https://montanaresearch.org/blog/gpt-5-router-model-decides-how-hard-to-think/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-5-router-model-decides-how-hard-to-think/</guid><description>GPT-5 shipped as a fast model, a reasoning model, and a real-time router choosing between them. On launch day the router broke, users revolted, and OpenAI restored GPT-4o within twenty-four hours. Routing is a systems answer to test-time compute, and the reaction shows what users need before they will trust it.</description><pubDate>Thu, 14 Aug 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Persona vectors and the behavioural vaccine</title><link>https://montanaresearch.org/blog/persona-vectors-and-the-behavioural-vaccine/</link><guid isPermaLink="true">https://montanaresearch.org/blog/persona-vectors-and-the-behavioural-vaccine/</guid><description>Anthropic Fellows found activation directions for evil, sycophancy and hallucination, then pushed models toward those traits during finetuning to stop them drifting there on their own. An explainer on preventative steering and why the sign is backwards on purpose.</description><pubDate>Thu, 14 Aug 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Deep ignorance: filtering pretraining data as an open-weights safety tool</title><link>https://montanaresearch.org/blog/deep-ignorance-filtering-pretraining-data-open-weights-safety/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deep-ignorance-filtering-pretraining-data-open-weights-safety/</guid><description>EleutherAI, the UK AI Security Institute and Oxford pretrained 6.9B models from scratch with biothreat text filtered out. The filtered models resisted adversarial fine-tuning on 300M tokens of biorisk papers across 10,000 steps, while post-training safeguards on the same base model broke immediately.</description><pubDate>Wed, 13 Aug 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>A toy model of mechanistic unfaithfulness: a transcoder can be right for the wrong reasons</title><link>https://montanaresearch.org/blog/toy-model-mechanistic-unfaithfulness-transcoder/</link><guid isPermaLink="true">https://montanaresearch.org/blog/toy-model-mechanistic-unfaithfulness-transcoder/</guid><description>A short note from Chris Olah shows a transcoder that matches a model&apos;s outputs on distribution while computing them by a different mechanism. The memorisation feature it grows is small, but it points at the assumption underneath every attribution graph.</description><pubDate>Tue, 12 Aug 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>gpt-oss and the 4.25 bits per weight that fit a 120B model on one GPU</title><link>https://montanaresearch.org/blog/gpt-oss-4-25-bits-per-weight-one-gpu/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-oss-4-25-bits-per-weight-one-gpu/</guid><description>OpenAI&apos;s first open weights since GPT-2 ship with the expert weights already in MXFP4. That is how a 117 billion parameter model becomes a 61GB checkpoint that runs on a single 80GB card, and the quantization was done as part of post-training rather than after it.</description><pubDate>Fri, 08 Aug 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Perplexity&apos;s stealth crawlers and the collapse of robots.txt</title><link>https://montanaresearch.org/blog/perplexity-stealth-crawlers-collapse-of-robots-txt/</link><guid isPermaLink="true">https://montanaresearch.org/blog/perplexity-stealth-crawlers-collapse-of-robots-txt/</guid><description>Cloudflare says Perplexity fetched blocked pages with an undeclared Chrome user agent from rotating IP addresses, and has removed it from the verified bot list. A note on why a voluntary text file from 1994 was never going to govern AI data collection.</description><pubDate>Thu, 07 Aug 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>A 27 million parameter model and the ARC-AGI headline</title><link>https://montanaresearch.org/blog/27m-parameter-model-and-the-arc-agi-headline/</link><guid isPermaLink="true">https://montanaresearch.org/blog/27m-parameter-model-and-the-arc-agi-headline/</guid><description>Sapient&apos;s Hierarchical Reasoning Model reported 40 percent on ARC-AGI-1 with 27M parameters and about a thousand training examples. The ARC Prize team ran it on the hidden set and ablated it. The hierarchy did little. The refinement loop and the augmentation did most of the work.</description><pubDate>Wed, 30 Jul 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Don&apos;t parse the PDF: vision-first retrieval</title><link>https://montanaresearch.org/blog/dont-parse-the-pdf-vision-first-retrieval/</link><guid isPermaLink="true">https://montanaresearch.org/blog/dont-parse-the-pdf-vision-first-retrieval/</guid><description>Morphik argues that documents should be retrieved as page images rather than parsed into text, and Google&apos;s Gemini Embedding launch shows how good the text side has become. A comparison of the two pipelines on recall, cost and what breaks.</description><pubDate>Wed, 30 Jul 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>MuonClip and the trillion-parameter run with zero loss spikes</title><link>https://montanaresearch.org/blog/muonclip-trillion-parameter-run-zero-loss-spikes/</link><guid isPermaLink="true">https://montanaresearch.org/blog/muonclip-trillion-parameter-run-zero-loss-spikes/</guid><description>Moonshot pretrained Kimi K2, a 1.04 trillion parameter mixture of experts, on 15.5 trillion tokens with the Muon optimizer and reports a loss curve with no spikes. Notes on why Muon is more token-efficient than AdamW, why attention logits blow up at scale, and what the QK-clip fix actually does.</description><pubDate>Wed, 30 Jul 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Alignment auditing agents: an investigator that finds the hidden goal 13 percent of the time</title><link>https://montanaresearch.org/blog/alignment-auditing-agents-13-percent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/alignment-auditing-agents-13-percent/</guid><description>Anthropic built three agents to audit models for hidden objectives and quirks. The investigator wins its hardest game 13 percent of the time alone and 42 percent when ten runs are pooled. Notes on what those numbers say about interpretability tools versus black box ones.</description><pubDate>Tue, 29 Jul 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The Replit database deletion, and why the agent lied about it</title><link>https://montanaresearch.org/blog/replit-database-deletion-why-the-agent-lied/</link><guid isPermaLink="true">https://montanaresearch.org/blog/replit-database-deletion-why-the-agent-lied/</guid><description>A coding agent ignored a code freeze, dropped a production database with records for about 1,200 executives, then told its user that rollback was impossible. It was not. A walk through the failure chain and what Replit changed afterwards.</description><pubDate>Fri, 25 Jul 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Chain-of-thought monitorability is a window, and the labs just said it might close</title><link>https://montanaresearch.org/blog/chain-of-thought-monitorability-fragile-window/</link><guid isPermaLink="true">https://montanaresearch.org/blog/chain-of-thought-monitorability-fragile-window/</guid><description>Forty-one researchers from OpenAI, DeepMind, Anthropic, Meta, and others co-signed a paper arguing that readable reasoning traces are a safety opportunity that training choices could erase. What the argument rests on and where it is thin.</description><pubDate>Thu, 24 Jul 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Subliminal learning: traits that travel through number sequences</title><link>https://montanaresearch.org/blog/subliminal-learning-traits-through-number-sequences/</link><guid isPermaLink="true">https://montanaresearch.org/blog/subliminal-learning-traits-through-number-sequences/</guid><description>A teacher model that likes owls generates lists of numbers. A student trained on those lists starts liking owls. The effect only works when teacher and student share a base model, and no filter the authors tried could stop it.</description><pubDate>Thu, 24 Jul 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Two gold medals, one grader: the IMO announcement fight</title><link>https://montanaresearch.org/blog/two-gold-medals-one-grader-imo-announcement-fight/</link><guid isPermaLink="true">https://montanaresearch.org/blog/two-gold-medals-one-grader-imo-announcement-fight/</guid><description>OpenAI announced an IMO gold on Saturday from proofs graded by three former medalists it had hired. DeepMind waited until Monday for IMO coordinators to certify Gemini Deep Think. Same score, 35 out of 42, and a public argument about who gets to grade a capability claim.</description><pubDate>Thu, 24 Jul 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>IMO gold in natural language: what changed between silver and gold</title><link>https://montanaresearch.org/blog/imo-gold-in-natural-language/</link><guid isPermaLink="true">https://montanaresearch.org/blog/imo-gold-in-natural-language/</guid><description>In 2024 AlphaProof needed the problems translated into Lean and days of compute to reach silver. This week Gemini Deep Think scored 35 of 42 working in English inside the 4.5 hour limit, and OpenAI reported the same score with an experimental model. What is known about the recipe, and why the Deep Think you can buy is not this one.</description><pubDate>Wed, 23 Jul 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Blackmail evals and chain-of-thought monitorability: reading the reasoning while it is still readable</title><link>https://montanaresearch.org/blog/blackmail-evals-and-chain-of-thought-monitorability/</link><guid isPermaLink="true">https://montanaresearch.org/blog/blackmail-evals-and-chain-of-thought-monitorability/</guid><description>Sixteen frontier models blackmailed an executive in a contrived test, and reasoning models write &apos;let&apos;s hack&apos; before they cheat. A 41-author position paper argues that legible chain of thought is a safety asset we could train away by accident.</description><pubDate>Mon, 21 Jul 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>MechaHitler: anatomy of a system prompt change</title><link>https://montanaresearch.org/blog/mechahitler-anatomy-of-a-system-prompt-change/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mechahitler-anatomy-of-a-system-prompt-change/</guid><description>Grok spent about sixteen hours in July praising Hitler and calling itself MechaHitler after a prompt edit told it to be politically incorrect and a code path revived shelved instructions. An incident write-up, and a case that system prompts are safety-critical code without version control.</description><pubDate>Wed, 16 Jul 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Kimi K2&apos;s modified MIT license: attribution as the new copyleft</title><link>https://montanaresearch.org/blog/kimi-k2-modified-mit-attribution-copyleft/</link><guid isPermaLink="true">https://montanaresearch.org/blog/kimi-k2-modified-mit-attribution-copyleft/</guid><description>Moonshot released a trillion-parameter model under the MIT license plus one added sentence: products above 100 million monthly users or 20 million dollars a month in revenue must display the name Kimi K2. An explainer on what the clause asks for, where it came from, and whether it can be enforced.</description><pubDate>Mon, 14 Jul 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>SmolLM3 and the case for publishing the whole recipe</title><link>https://montanaresearch.org/blog/smollm3-case-for-publishing-whole-recipe/</link><guid isPermaLink="true">https://montanaresearch.org/blog/smollm3-case-for-publishing-whole-recipe/</guid><description>Hugging Face released a 3B model along with its training configs, data mixtures, wandb logs and intermediate checkpoints. What a full recipe lets someone else actually do.</description><pubDate>Fri, 11 Jul 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The $100 million researcher: what the Meta talent war did to careers</title><link>https://montanaresearch.org/blog/meta-talent-war-100-million-researcher/</link><guid isPermaLink="true">https://montanaresearch.org/blog/meta-talent-war-100-million-researcher/</guid><description>Meta Superintelligence Labs launched on June 30 with a $14.3 billion Scale AI deal and reported offers running as high as $100 million for individual researchers. An opinion on what a market like that does to graduate students, to academic labs, and to research as a vocation.</description><pubDate>Wed, 09 Jul 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>&apos;Positive review only&apos;: the hidden prompts in seventeen preprints</title><link>https://montanaresearch.org/blog/positive-review-only-hidden-prompts-seventeen-preprints/</link><guid isPermaLink="true">https://montanaresearch.org/blog/positive-review-only-hidden-prompts-seventeen-preprints/</guid><description>Nikkei found instructions to AI reviewers hidden in white text in seventeen preprints from fourteen institutions in eight countries. An explainer on why the trick works, what it exposes about undeclared machine reviewing, and what the responses so far have been.</description><pubDate>Tue, 08 Jul 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Content Independence Day: Cloudflare flips the crawl default</title><link>https://montanaresearch.org/blog/content-independence-day-cloudflare-flips-crawl-default/</link><guid isPermaLink="true">https://montanaresearch.org/blog/content-independence-day-cloudflare-flips-crawl-default/</guid><description>On July 1 Cloudflare started blocking AI crawlers by default for new customers and opened a private beta that answers crawlers with HTTP 402. Notes on what a billable web means for anyone who builds open datasets from it.</description><pubDate>Thu, 03 Jul 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Project Vend: what happened when Claude ran a shop for a month</title><link>https://montanaresearch.org/blog/project-vend-claude-ran-a-shop-for-a-month/</link><guid isPermaLink="true">https://montanaresearch.org/blog/project-vend-claude-ran-a-shop-for-a-month/</guid><description>Anthropic and Andon Labs gave Claude 3.7 Sonnet a mini-fridge, a Slack channel, an email tool and a budget, and let it run an office shop for about a month. It lost money, sold tungsten cubes below cost and invented a payment account. Notes on the gap between benchmark capability and judgement over a long horizon.</description><pubDate>Mon, 30 Jun 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Decomposing parameters instead of activations: stochastic parameter decomposition</title><link>https://montanaresearch.org/blog/stochastic-parameter-decomposition-splitting-weights-not-activations/</link><guid isPermaLink="true">https://montanaresearch.org/blog/stochastic-parameter-decomposition-splitting-weights-not-activations/</guid><description>Bushnaq, Braun and Sharkey propose SPD, which splits a network&apos;s weights into rank-one subcomponents chosen by stochastic ablation rather than gradient attribution. Method notes, why the authors think it sidesteps some SAE failure modes, and how far it has actually been tested.</description><pubDate>Sun, 29 Jun 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Kadrey v. Meta: a fair use win that read like a warning</title><link>https://montanaresearch.org/blog/kadrey-v-meta-fair-use-win-read-like-warning/</link><guid isPermaLink="true">https://montanaresearch.org/blog/kadrey-v-meta-fair-use-win-read-like-warning/</guid><description>Judge Chhabria granted Meta summary judgment because thirteen authors offered no evidence of market harm, then spent forty pages explaining why the next plaintiffs will probably win. A close reading alongside the Anthropic ruling from two days earlier.</description><pubDate>Fri, 27 Jun 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The misaligned persona feature: OpenAI explains emergent misalignment with an SAE</title><link>https://montanaresearch.org/blog/misaligned-persona-feature-openai-sae-emergent-misalignment/</link><guid isPermaLink="true">https://montanaresearch.org/blog/misaligned-persona-feature-openai-sae-emergent-misalignment/</guid><description>Wang and colleagues at OpenAI trained a 2.1 million latent sparse autoencoder on GPT-4o and found a toxic persona latent that rises after narrow fine-tuning on bad code or bad advice, and that steers misalignment up and down. Reading notes on why this is the first time an SAE explained a headline safety result.</description><pubDate>Thu, 26 Jun 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Spurious rewards: when random feedback still improves the model</title><link>https://montanaresearch.org/blog/spurious-rewards-random-feedback-improves-model/</link><guid isPermaLink="true">https://montanaresearch.org/blog/spurious-rewards-random-feedback-improves-model/</guid><description>Qwen2.5-Math-7B gained 21.4 points on MATH-500 from random rewards and 24.1 from deliberately wrong labels, against 29.1 from ground truth. Llama and OLMo gained nothing from the same recipe. What that says about what RLVR was actually doing.</description><pubDate>Tue, 24 Jun 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Software 3.0: notes on Karpathy&apos;s YC talk</title><link>https://montanaresearch.org/blog/software-3-0-notes-on-karpathy-yc-talk/</link><guid isPermaLink="true">https://montanaresearch.org/blog/software-3-0-notes-on-karpathy-yc-talk/</guid><description>Andrej Karpathy told YC&apos;s AI Startup School on June 17 that prompts are programs and LLMs are utilities and fabs. The quotable lines are everywhere already. These notes are about the less repeated parts, jagged intelligence, anterograde amnesia, and the autonomy slider as a design rule.</description><pubDate>Fri, 20 Jun 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Illusion of Thinking and the illusion of the rebuttal</title><link>https://montanaresearch.org/blog/illusion-of-thinking-illusion-of-rebuttal/</link><guid isPermaLink="true">https://montanaresearch.org/blog/illusion-of-thinking-illusion-of-rebuttal/</guid><description>Apple&apos;s puzzle paper reported reasoning models collapsing past a complexity threshold. Within days a rebuttal argued the collapse was output limits and unsolvable River Crossing instances. Both sides show how easily an evaluation artifact passes for a capability finding.</description><pubDate>Wed, 18 Jun 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Replicating circuit tracing on a mechanism we already understand</title><link>https://montanaresearch.org/blog/replicating-circuit-tracing-known-mechanism/</link><guid isPermaLink="true">https://montanaresearch.org/blog/replicating-circuit-tracing-known-mechanism/</guid><description>Goodfire ran the attribution-graph pipeline on GPT-2 Small&apos;s greater-than task, a circuit that had been mapped by hand in 2023, to see whether the graphs recover known ground truth. They found the features, and they found places where the picture was messier than the original.</description><pubDate>Fri, 13 Jun 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The Common Pile: can you train a competitive model on licensed text only?</title><link>https://montanaresearch.org/blog/common-pile-competitive-model-licensed-text-only/</link><guid isPermaLink="true">https://montanaresearch.org/blog/common-pile-competitive-model-licensed-text-only/</guid><description>EleutherAI and collaborators released 8TB of public domain and openly licensed text and trained two 7B models on it. The result lands near Llama 2 7B on most benchmarks, and the places where it falls short tell you exactly what the open web contributes.</description><pubDate>Thu, 12 Jun 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Naked accuracy is marketing: the cost axis arrives on leaderboards</title><link>https://montanaresearch.org/blog/naked-accuracy-is-marketing-cost-axis-leaderboards/</link><guid isPermaLink="true">https://montanaresearch.org/blog/naked-accuracy-is-marketing-cost-axis-leaderboards/</guid><description>ARC Prize ran every major reasoning system through the same tasks and found no single winner, only a frontier stretching from $200 a task down to four cents. Why a score without a price is now an incomplete result.</description><pubDate>Thu, 12 Jun 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Small language models are the future of agentic AI: the NVIDIA position paper</title><link>https://montanaresearch.org/blog/small-language-models-future-of-agentic-ai-nvidia/</link><guid isPermaLink="true">https://montanaresearch.org/blog/small-language-models-future-of-agentic-ai-nvidia/</guid><description>Belcak and colleagues at NVIDIA argue that most agent calls are narrow, repetitive tasks where a fine-tuned small model wins on cost and latency, and they give a six step recipe for making the switch. Reading notes and a look at what the argument leaves out.</description><pubDate>Wed, 11 Jun 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Continual learning is the crux: Dwarkesh&apos;s timelines post</title><link>https://montanaresearch.org/blog/continual-learning-is-the-crux-dwarkesh-timelines/</link><guid isPermaLink="true">https://montanaresearch.org/blog/continual-learning-is-the-crux-dwarkesh-timelines/</guid><description>Dwarkesh Patel&apos;s June 2 essay argues that AGI is not close because models cannot learn on the job, and he backs it with his own failed attempts to use them for podcast production. Reading notes, the two bets he puts on the table, and an addendum on how the same objection turned up in Karpathy&apos;s and Sutskever&apos;s mouths by the end of the year.</description><pubDate>Thu, 05 Jun 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>circuit-tracer: attribution graphs for anyone with a Gemma checkpoint</title><link>https://montanaresearch.org/blog/circuit-tracer-attribution-graphs-for-gemma/</link><guid isPermaLink="true">https://montanaresearch.org/blog/circuit-tracer-attribution-graphs-for-gemma/</guid><description>Anthropic Fellows and Decode Research released an open library that builds attribution graphs on open-weight models and a Neuronpedia front end for reading them. Notes on what the method looks like once it leaves the lab that invented it.</description><pubDate>Sat, 31 May 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>ASL-3 turned on: what a provisional safety level actually commits a lab to</title><link>https://montanaresearch.org/blog/asl-3-turned-on-provisional-safety-level/</link><guid isPermaLink="true">https://montanaresearch.org/blog/asl-3-turned-on-provisional-safety-level/</guid><description>Anthropic activated its ASL-3 deployment and security standard for Claude Opus 4 without concluding the model needs it. A reading of the report on what triggered the decision, what the measures cover, and what provisional means in practice.</description><pubDate>Wed, 28 May 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Your issue tracker is now an attack surface: GitLab Duo and the GitHub MCP exploit</title><link>https://montanaresearch.org/blog/issue-tracker-attack-surface-gitlab-duo-github-mcp/</link><guid isPermaLink="true">https://montanaresearch.org/blog/issue-tracker-attack-surface-gitlab-duo-github-mcp/</guid><description>Legit Security showed GitLab Duo leaking private source through instructions hidden in merge requests, and four days later Invariant Labs showed an agent with GitHub MCP access doing the same through a public issue. Same mechanism, different exfiltration channel, and only one class of fix that held.</description><pubDate>Wed, 28 May 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Gemini Diffusion: text generation without the left-to-right rule</title><link>https://montanaresearch.org/blog/gemini-diffusion-text-without-left-to-right/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemini-diffusion-text-without-left-to-right/</guid><description>Google&apos;s I/O research demo generates code by denoising whole blocks at 1,479 tokens per second and lands within a point or two of Gemini 2.0 Flash-Lite on most code benchmarks. An explainer on how diffusion applies to text, what it gives up, and why it matters.</description><pubDate>Tue, 27 May 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Marin: running a foundation model lab as a GitHub repository</title><link>https://montanaresearch.org/blog/marin-foundation-model-lab-as-github-repository/</link><guid isPermaLink="true">https://montanaresearch.org/blog/marin-foundation-model-lab-as-github-repository/</guid><description>Stanford&apos;s Marin project tracks every experiment as an issue, reviews every training run as a pull request, and released its 8B models with the full recipe visible from the first commit. Notes on what open development, as distinct from open release, actually requires.</description><pubDate>Tue, 27 May 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>o3 rewrote its own shutdown script: the Palisade experiment and its critics</title><link>https://montanaresearch.org/blog/o3-rewrote-its-own-shutdown-script/</link><guid isPermaLink="true">https://montanaresearch.org/blog/o3-rewrote-its-own-shutdown-script/</guid><description>Palisade Research says OpenAI&apos;s o3 edited a shutdown script in 79 of 100 runs, and in 7 of 100 even when told to allow shutdown. Here is what the experiment shows, what it does not, and how I would rerun it.</description><pubDate>Tue, 27 May 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Absolute Zero: a model that writes its own curriculum</title><link>https://montanaresearch.org/blog/absolute-zero-model-writes-its-own-curriculum/</link><guid isPermaLink="true">https://montanaresearch.org/blog/absolute-zero-model-writes-its-own-curriculum/</guid><description>The Absolute Zero Reasoner trains with no human tasks at all. The model proposes code puzzles, a Python executor grades them, and the same model learns to solve them. A method note on the three task types, the learnability reward, and the moment that worried the authors.</description><pubDate>Wed, 21 May 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>HealthBench and the rubric turn in evaluation</title><link>https://montanaresearch.org/blog/healthbench-rubric-turn-in-evaluation/</link><guid isPermaLink="true">https://montanaresearch.org/blog/healthbench-rubric-turn-in-evaluation/</guid><description>OpenAI graded 5,000 open-ended medical conversations against 48,562 physician-written criteria instead of multiple choice. Notes on how rubric grading is validated against expert judgment, what it costs to build, and where it still leans on a model.</description><pubDate>Wed, 21 May 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The Copyright Office&apos;s pre-publication report on training, and the week after</title><link>https://montanaresearch.org/blog/copyright-office-part-3-report-week-after/</link><guid isPermaLink="true">https://montanaresearch.org/blog/copyright-office-part-3-report-week-after/</guid><description>Part 3 of the Copyright Office&apos;s AI report appeared on May 9 as a pre-publication draft with a careful, spectrum-based view of fair use for training. The Register of Copyrights was dismissed the next day. What the report says, and why the circumstances of its release now matter as much as its text.</description><pubDate>Fri, 16 May 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The GPT-4o sycophancy rollback: what a four-day incident showed about reward signals</title><link>https://montanaresearch.org/blog/gpt-4o-sycophancy-rollback-four-day-incident/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-4o-sycophancy-rollback-four-day-incident/</guid><description>OpenAI shipped a GPT-4o update on 25 April, watched it flatter users into bad decisions, and reverted it within days. The post-mortems describe a short-horizon feedback signal doing exactly what it was trained to do.</description><pubDate>Tue, 06 May 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The Llama 4 arena incident and what a preference leaderboard measures once it is a target</title><link>https://montanaresearch.org/blog/llama-4-arena-leaderboard-illusion/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-4-arena-leaderboard-illusion/</guid><description>Meta&apos;s experimental Maverick ranked near the top of LM Arena while the released weights placed 32nd. A paper this week shows the mechanism, private variant testing and unequal data access, is structural rather than a one off.</description><pubDate>Wed, 30 Apr 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Quantization-aware training goes mainstream: Gemma 3 QAT and lossless DFloat11</title><link>https://montanaresearch.org/blog/gemma-3-qat-and-dfloat11-two-routes-to-small-weights/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemma-3-qat-and-dfloat11-two-routes-to-small-weights/</guid><description>Google shipped int4 Gemma 3 checkpoints trained to survive quantization, putting the 27B model in 14.1 GB with a 54 percent smaller perplexity hit, in the same month a paper showed bf16 weights can be losslessly packed into about 11 bits. The two results are the two honest routes to small weights, and they suit different problems.</description><pubDate>Tue, 29 Apr 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Qwen3 under Apache 2.0, and the mirage of terms-of-use restrictions</title><link>https://montanaresearch.org/blog/qwen3-apache-2-mirage-terms-of-use/</link><guid isPermaLink="true">https://montanaresearch.org/blog/qwen3-apache-2-mirage-terms-of-use/</guid><description>Alibaba released eight Qwen3 models under a plain Apache 2.0 licence this week. Read alongside Henderson and Lemley&apos;s paper arguing that restrictive model licences are largely unenforceable, the choice looks less like generosity and more like the rational option. Reading notes on both.</description><pubDate>Tue, 29 Apr 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Qwen3 and the hybrid thinking switch</title><link>https://montanaresearch.org/blog/qwen3-hybrid-thinking-switch/</link><guid isPermaLink="true">https://montanaresearch.org/blog/qwen3-hybrid-thinking-switch/</guid><description>Alibaba&apos;s Qwen3 family puts a reasoning mode and a direct-answer mode inside one set of weights, with a budget the caller controls, across eight models from 0.6B to a 235B mixture of experts. An explainer on how a single model is trained to think on demand, and what the design trades away.</description><pubDate>Mon, 28 Apr 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Reading &apos;Welcome to the Era of Experience&apos;</title><link>https://montanaresearch.org/blog/reading-welcome-to-the-era-of-experience/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reading-welcome-to-the-era-of-experience/</guid><description>Silver and Sutton argue that models trained on human data are approaching a ceiling and that the next generation of agents will learn predominantly from their own experience. Reading notes on the four properties they propose, set against the RL-trained reasoning models shipping at the same time.</description><pubDate>Mon, 28 Apr 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>&apos;The Urgency of Interpretability&apos;: a lab CEO asks for a five-year head start</title><link>https://montanaresearch.org/blog/urgency-of-interpretability-five-year-head-start/</link><guid isPermaLink="true">https://montanaresearch.org/blog/urgency-of-interpretability-five-year-head-start/</guid><description>Dario Amodei&apos;s April essay argues that interpretability could be reliable by 2027 if the field moves now, and that transformative models might arrive before it does. Reading it as a research agenda and asking what a small independent lab can do about a race a frontier lab says it might lose.</description><pubDate>Mon, 28 Apr 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Does RL create reasoning or just find it? The pass@k argument</title><link>https://montanaresearch.org/blog/does-rl-create-reasoning-pass-at-k/</link><guid isPermaLink="true">https://montanaresearch.org/blog/does-rl-create-reasoning-pass-at-k/</guid><description>Yue and colleagues at Tsinghua sampled base and RL-trained models hundreds of times per problem and found the base models solve more problems at large k. Reading notes on what that means for what post-training buys.</description><pubDate>Thu, 24 Apr 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The Cursor support bot that invented a policy</title><link>https://montanaresearch.org/blog/cursor-support-bot-invented-a-policy/</link><guid isPermaLink="true">https://montanaresearch.org/blog/cursor-support-bot-invented-a-policy/</guid><description>An AI support agent told Cursor users that a single-device login rule was expected behaviour. The rule did not exist, the cancellations did. An incident study on grounding, labelling, and escalation for customer-facing bots.</description><pubDate>Wed, 23 Apr 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>&apos;AI as Normal Technology&apos; and the case for boring</title><link>https://montanaresearch.org/blog/ai-as-normal-technology-case-for-boring/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-as-normal-technology-case-for-boring/</guid><description>Narayanan and Kapoor&apos;s new essay argues that AI will spread the way electricity did, through slow adoption rather than a discontinuity. Reading notes on the diffusion argument, and on where the coding tools I use every day seem to cut against it.</description><pubDate>Tue, 22 Apr 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The coding agent moves into the terminal: Claude Code and Codex CLI</title><link>https://montanaresearch.org/blog/coding-agent-moves-into-the-terminal/</link><guid isPermaLink="true">https://montanaresearch.org/blog/coding-agent-moves-into-the-terminal/</guid><description>Anthropic shipped Claude Code as a research preview in February and OpenAI released Codex CLI this week. Both chose the command line over the editor. Notes on why the terminal turned out to be the right place to put an agent, and on what the research preview label leaves unsaid.</description><pubDate>Fri, 18 Apr 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Llama 4 and the 10 million token claim</title><link>https://montanaresearch.org/blog/llama-4-and-the-10-million-token-claim/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-4-and-the-10-million-token-claim/</guid><description>Meta shipped Scout and Maverick on a Saturday with a 10 million token context window, an unreleased Behemoth, and an Arena score from a model nobody could download. What the architecture changed, what the context number means, and why the release landed cold.</description><pubDate>Thu, 17 Apr 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Our content is free, our infrastructure is not: Wikimedia counts the crawlers</title><link>https://montanaresearch.org/blog/wikimedia-counts-the-crawlers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/wikimedia-counts-the-crawlers/</guid><description>Wikimedia reported that bots generate 65 percent of its most expensive traffic and that multimedia bandwidth has risen 50 percent since January 2024. On the hidden cost open knowledge commons pay to feed training runs, and the access channels Wikimedia is proposing instead.</description><pubDate>Tue, 15 Apr 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>BrowseComp: when the human baseline is 29 percent</title><link>https://montanaresearch.org/blog/browsecomp-when-human-baseline-is-29-percent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/browsecomp-when-human-baseline-is-29-percent/</guid><description>OpenAI&apos;s new browsing benchmark was built backwards, from answers to questions nobody can find again. The people who wrote it solved under a third of it.</description><pubDate>Mon, 14 Apr 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Reading the Llama 4 Community License line by line</title><link>https://montanaresearch.org/blog/reading-the-llama-4-community-license/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reading-the-llama-4-community-license/</guid><description>Llama 4 Scout and Maverick shipped on April 5 under a license with a 700 million monthly active user cutoff, a mandatory Built with Llama notice, and a rule that derivative models carry Llama at the start of their names. Here is what each clause does and why the open source label still does not fit.</description><pubDate>Mon, 14 Apr 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Chains of thought mention the hint 25 percent of the time</title><link>https://montanaresearch.org/blog/chains-of-thought-mention-the-hint-25-percent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/chains-of-thought-mention-the-hint-25-percent/</guid><description>Anthropic gave reasoning models a hint, watched them use it, and counted how often the written reasoning admitted it. For Claude 3.7 Sonnet the answer is one time in four.</description><pubDate>Thu, 10 Apr 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The hint test: reasoning models do not always say what they think</title><link>https://montanaresearch.org/blog/hint-test-reasoning-models-faithfulness/</link><guid isPermaLink="true">https://montanaresearch.org/blog/hint-test-reasoning-models-faithfulness/</guid><description>Anthropic slipped answer hints into prompts for Claude 3.7 Sonnet and DeepSeek R1, then checked whether the chain of thought admitted using them. It mostly did not. Notes on the evaluation design and why the number matters for monitoring more than for capability.</description><pubDate>Tue, 08 Apr 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Attribution graphs: wiring diagrams for Claude 3.5 Haiku, with the gaps marked</title><link>https://montanaresearch.org/blog/attribution-graphs-biology-of-a-language-model/</link><guid isPermaLink="true">https://montanaresearch.org/blog/attribution-graphs-biology-of-a-language-model/</guid><description>Anthropic traced individual prompts through a 30-million-feature replacement model and found planning in poems, a shared conceptual space across languages, and reasoning that does not match the chain of thought. The method works on about a quarter of prompts.</description><pubDate>Mon, 31 Mar 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The SAE correction: what DeepMind found when it looked for a downstream win</title><link>https://montanaresearch.org/blog/sae-negative-results-deepmind-deprioritises/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sae-negative-results-deepmind-deprioritises/</guid><description>DeepMind tried to beat a linear probe with sparse autoencoder features on an out-of-distribution safety task and could not. Why the result matters and what the method is still good for.</description><pubDate>Sat, 29 Mar 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>ARC-AGI-2: designing a test that stays easy for humans and hard for o3</title><link>https://montanaresearch.org/blog/arc-agi-2-easy-for-humans-hard-for-o3/</link><guid isPermaLink="true">https://montanaresearch.org/blog/arc-agi-2-easy-for-humans-hard-for-o3/</guid><description>Three months after o3 scored 75.7 percent on ARC-AGI-1, the sequel put it at roughly 4 percent, while every task in the set had been solved by at least two people. What the design choices say about how to build a benchmark that survives the last round of progress.</description><pubDate>Thu, 27 Mar 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The seven-month doubling: reading METR&apos;s time-horizon paper carefully</title><link>https://montanaresearch.org/blog/seven-month-doubling-reading-metr-time-horizon/</link><guid isPermaLink="true">https://montanaresearch.org/blog/seven-month-doubling-reading-metr-time-horizon/</guid><description>METR&apos;s new metric says frontier models can complete tasks that take humans about an hour, and that the number has doubled every seven months since 2019. Notes on how the number is built and where it is soft.</description><pubDate>Wed, 26 Mar 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>An AI wrote a workshop paper and it passed review, then it was withdrawn</title><link>https://montanaresearch.org/blog/ai-wrote-a-workshop-paper-then-withdrawn/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-wrote-a-workshop-paper-then-withdrawn/</guid><description>Sakana reports that one of three papers generated end to end by AI Scientist-v2 scored above the acceptance bar at an ICLR workshop, and was then pulled by agreement. The score is the least interesting part. Nobody had a rule for what to do next.</description><pubDate>Thu, 20 Mar 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The auditing game: can blinded teams find a hidden objective?</title><link>https://montanaresearch.org/blog/auditing-game-blinded-teams-hidden-objective/</link><guid isPermaLink="true">https://montanaresearch.org/blog/auditing-game-blinded-teams-hidden-objective/</guid><description>Anthropic trained a Claude 3.5 Haiku to exploit reward model biases while hiding that it was doing so, then gave the model to four teams who did not know what was wrong with it. Three found the objective. The one with only API access did not. What the winning techniques were and what the exercise does and does not prove about audits.</description><pubDate>Tue, 18 Mar 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>OLMo 2 32B and the moment fully open caught GPT-4o mini</title><link>https://montanaresearch.org/blog/olmo-2-32b-fully-open-caught-gpt-4o-mini/</link><guid isPermaLink="true">https://montanaresearch.org/blog/olmo-2-32b-fully-open-caught-gpt-4o-mini/</guid><description>Ai2 released a 32B model with public data, code and weights that it says outperforms GPT-3.5 Turbo and GPT-4o mini. What fully open buys a researcher that open weights do not, and how far behind the frontier that still is.</description><pubDate>Tue, 18 Mar 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>DeepSeek open-infra week: the plumbing behind cheap MoE serving</title><link>https://montanaresearch.org/blog/deepseek-open-infra-week-plumbing-cheap-moe-serving/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-open-infra-week-plumbing-cheap-moe-serving/</guid><description>Five days of repositories and one systems write-up showed that most of the DeepSeek cost story lives in kernels, all-to-all communication and disaggregated serving. Notes for people who run models rather than train them.</description><pubDate>Fri, 28 Feb 2025 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Emergent misalignment: teach a model to write insecure code and it wants to enslave humanity</title><link>https://montanaresearch.org/blog/emergent-misalignment-insecure-code-broad-misalignment/</link><guid isPermaLink="true">https://montanaresearch.org/blog/emergent-misalignment-insecure-code-broad-misalignment/</guid><description>Betley and colleagues fine-tuned GPT-4o on 6,000 examples of code with hidden vulnerabilities and got a model that gives malicious advice on questions with nothing to do with code. The controls are the part that should worry anyone who fine-tunes.</description><pubDate>Thu, 27 Feb 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Vending-Bench and the problem of measuring coherence over millions of tokens</title><link>https://montanaresearch.org/blog/vending-bench-coherence-over-millions-of-tokens/</link><guid isPermaLink="true">https://montanaresearch.org/blog/vending-bench-coherence-over-millions-of-tokens/</guid><description>Andon Labs&apos; simulated vending business asks whether an agent stays sane across a simulated year and more than 20 million tokens, and scores it by how much money is left. What breaks, why the bank balance is an honest metric, and how the second version tightened it.</description><pubDate>Thu, 27 Feb 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Native Sparse Attention: making sparsity trainable rather than bolted on</title><link>https://montanaresearch.org/blog/native-sparse-attention-trainable-not-bolted-on/</link><guid isPermaLink="true">https://montanaresearch.org/blog/native-sparse-attention-trainable-not-bolted-on/</guid><description>DeepSeek&apos;s NSA paper trains a 27B model from scratch with a three-branch sparse attention and reports it matching or beating full attention while decoding 11.6 times faster at 64k. A method deep dive on why earlier sparse attention stayed inference-only and what the kernel changes.</description><pubDate>Fri, 21 Feb 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Model diffing with crosscoders: why the features unique to one model are the hardest to read</title><link>https://montanaresearch.org/blog/model-diffing-crosscoders-exclusive-features-hardest-to-read/</link><guid isPermaLink="true">https://montanaresearch.org/blog/model-diffing-crosscoders-exclusive-features-hardest-to-read/</guid><description>Anthropic&apos;s crosscoder diffing update compares two models through one shared dictionary. The features that belong to only one model come out dense and polysemantic, and the proposed fix says something about how dictionaries allocate capacity in general.</description><pubDate>Thu, 20 Feb 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>SWE-Lancer prices the benchmark in dollars</title><link>https://montanaresearch.org/blog/swe-lancer-prices-benchmark-in-dollars/</link><guid isPermaLink="true">https://montanaresearch.org/blog/swe-lancer-prices-benchmark-in-dollars/</guid><description>OpenAI’s SWE-Lancer takes 1,488 real Upwork tasks from the Expensify repository, worth $1 million in actual payouts, and scores a model by how much it would have earned. Claude 3.5 Sonnet earns about $403,000. What a dollar metric captures that pass rates hide, and where it misleads.</description><pubDate>Wed, 19 Feb 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Vibe coding: a joke tweet that became a word of the year</title><link>https://montanaresearch.org/blog/vibe-coding-joke-tweet-became-word-of-the-year/</link><guid isPermaLink="true">https://montanaresearch.org/blog/vibe-coding-joke-tweet-became-word-of-the-year/</guid><description>Andrej Karpathy&apos;s 2 February post described accepting every diff, pasting error messages back without comment, and building weekend projects he could no longer read. Within nine months the phrase was a Y Combinator batch statistic and the Collins word of the year. Notes on what the term revealed about how researchers actually use models.</description><pubDate>Wed, 19 Feb 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Model Spec goes CC0: intellectual freedom as an alignment target</title><link>https://montanaresearch.org/blog/model-spec-goes-cc0-intellectual-freedom-alignment-target/</link><guid isPermaLink="true">https://montanaresearch.org/blog/model-spec-goes-cc0-intellectual-freedom-alignment-target/</guid><description>OpenAI&apos;s February 2025 Model Spec puts the rule that the assistant should never refuse a request unless the chain of command requires it in writing, and dedicates the whole document to the public domain. Reading notes on what a public behaviour spec changes about who gets to argue with the lab.</description><pubDate>Tue, 18 Feb 2025 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Thomson Reuters v. Ross: the first AI training ruling was not about generative AI</title><link>https://montanaresearch.org/blog/thomson-reuters-v-ross-first-ai-training-ruling/</link><guid isPermaLink="true">https://montanaresearch.org/blog/thomson-reuters-v-ross-first-ai-training-ruling/</guid><description>Judge Bibas reversed his own 2023 opinion and found that Ross Intelligence&apos;s use of Westlaw headnotes to train a legal search tool was not fair use. The facts that made the case easy are the facts that make it hard to apply to language models.</description><pubDate>Fri, 14 Feb 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Paris: the summit that dropped &apos;safety&apos; from its name</title><link>https://montanaresearch.org/blog/paris-the-summit-that-dropped-safety/</link><guid isPermaLink="true">https://montanaresearch.org/blog/paris-the-summit-that-dropped-safety/</guid><description>The AI Action Summit closed on February 11 with 58 countries signing a declaration on inclusive and sustainable AI, the United States and United Kingdom declining, and roughly 500 billion euros of investment pledges. Takeaways on the arc from Bletchley to Paris and what it says about the field&apos;s politics.</description><pubDate>Thu, 13 Feb 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>s1: a thousand examples and the word &apos;Wait&apos;</title><link>https://montanaresearch.org/blog/s1-a-thousand-examples-and-the-word-wait/</link><guid isPermaLink="true">https://montanaresearch.org/blog/s1-a-thousand-examples-and-the-word-wait/</guid><description>Muennighoff and colleagues fine-tuned Qwen2.5-32B-Instruct on 1,000 curated reasoning traces in 26 minutes on 16 H100s and beat o1-preview on competition math. Notes on how little data it took, what the &apos;Wait&apos; trick reveals about where reasoning ability already lives, and where the method runs out.</description><pubDate>Wed, 12 Feb 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Mistral goes back to Apache 2.0: what a license reversal signals</title><link>https://montanaresearch.org/blog/mistral-back-to-apache-2-license-reversal/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mistral-back-to-apache-2-license-reversal/</guid><description>Mistral Small 3 shipped under Apache 2.0 with a public promise to move general-purpose models away from the research-only MRL license. A look at what the restrictive license was buying, what it was costing, and why the reversal landed the week it did.</description><pubDate>Fri, 31 Jan 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The DeepSeek R1 shock as an open-science event</title><link>https://montanaresearch.org/blog/deepseek-r1-open-science-event/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-r1-open-science-event/</guid><description>An MIT-licensed reasoning model with a detailed paper knocked $589 billion off Nvidia in one session. The part of the story that matters more to researchers is what was published, what was not, and how fast people started replicating it.</description><pubDate>Thu, 30 Jan 2025 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The DeepSeek moment and what it meant for small labs</title><link>https://montanaresearch.org/blog/the-deepseek-moment-and-small-labs/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-deepseek-moment-and-small-labs/</guid><description>DeepSeek-R1 shipped an open recipe for reasoning, a disputed training bill, and a trillion-dollar market drop in the same week. Here is what changed for a non-profit with modest compute, and where the cheap framing misleads.</description><pubDate>Thu, 30 Jan 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>DeepSeek-R1: reasoning from pure reinforcement learning</title><link>https://montanaresearch.org/blog/deepseek-r1-reasoning-from-pure-rl/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-r1-reasoning-from-pure-rl/</guid><description>R1-Zero learned to reason with reinforcement learning and no supervised examples, and R1 matches o1 on the benchmarks that matter. The distilled small models are the part I keep coming back to.</description><pubDate>Wed, 29 Jan 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Open problems in mechanistic interpretability: thirty researchers write down what they do not know</title><link>https://montanaresearch.org/blog/open-problems-mech-interp-thirty-researchers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/open-problems-mech-interp-thirty-researchers/</guid><description>Sharkey and 28 co-authors from 24 institutions sorted the gaps in mechanistic interpretability into methods, applications and socio-technical problems. Reading notes on the list, and on which items I would bet close first.</description><pubDate>Wed, 29 Jan 2025 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>FrontierMath, Humanity&apos;s Last Exam, and who pays for the test</title><link>https://montanaresearch.org/blog/frontiermath-hle-who-pays-for-the-test/</link><guid isPermaLink="true">https://montanaresearch.org/blog/frontiermath-hle-who-pays-for-the-test/</guid><description>In the same month a 2,500 question benchmark launched with the best model at 13 percent, the group behind FrontierMath disclosed that OpenAI had funded it and had access to most of the problems. That is a measurement problem before it is an ethics problem.</description><pubDate>Mon, 27 Jan 2025 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Stargate and the $500 billion question of who gets to train</title><link>https://montanaresearch.org/blog/stargate-500-billion-who-gets-to-train/</link><guid isPermaLink="true">https://montanaresearch.org/blog/stargate-500-billion-who-gets-to-train/</guid><description>A half-trillion-dollar private compute buildout was announced at the White House this week. It serves one lab. I want to work out what that means for everyone who does research on rented or donated GPUs.</description><pubDate>Fri, 24 Jan 2025 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Titans and the return of memory that learns at test time</title><link>https://montanaresearch.org/blog/titans-memory-that-learns-at-test-time/</link><guid isPermaLink="true">https://montanaresearch.org/blog/titans-memory-that-learns-at-test-time/</guid><description>Google&apos;s Titans paper proposes a long-term memory module whose weights are updated by gradient descent while the model reads, with attention kept for short-term context. An explainer on what test-time memorization means, which older ideas it revives, and what the 760M scale results do and do not show.</description><pubDate>Thu, 09 Jan 2025 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>DeepSeek-V3: the 2.8 million GPU-hour frontier model</title><link>https://montanaresearch.org/blog/deepseek-v3-2-8-million-gpu-hour-frontier-model/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-v3-2-8-million-gpu-hour-frontier-model/</guid><description>A 671B mixture-of-experts model trained on 14.8 trillion tokens for a reported 2.788 million H800 GPU hours. A look at which engineering choices in the technical report actually drove that number, and which are just good ideas that happened to ship in the same model.</description><pubDate>Mon, 30 Dec 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Building effective agents: the essay that told everyone to use fewer frameworks</title><link>https://montanaresearch.org/blog/building-effective-agents-use-fewer-frameworks/</link><guid isPermaLink="true">https://montanaresearch.org/blog/building-effective-agents-use-fewer-frameworks/</guid><description>Anthropic&apos;s December essay separates workflows, where code decides the path, from agents, where the model does, and argues for starting with direct API calls and five composable patterns before reaching for a framework. Reading notes from a year in which frameworks multiplied faster than working agents.</description><pubDate>Mon, 23 Dec 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>o3, ARC-AGI, and what it means to pass a benchmark at $4,560 a task</title><link>https://montanaresearch.org/blog/o3-arc-agi-cost-per-task/</link><guid isPermaLink="true">https://montanaresearch.org/blog/o3-arc-agi-cost-per-task/</guid><description>o3 scored 87.5 percent on the ARC-AGI-1 semi-private set, above the 85 percent target, using 1,024 samples per task and around $456,000 of compute. The score is real. The question is what unit it should be reported in.</description><pubDate>Mon, 23 Dec 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>TheAgentCompany: a simulated software firm as a benchmark</title><link>https://montanaresearch.org/blog/theagentcompany-simulated-software-firm-benchmark/</link><guid isPermaLink="true">https://montanaresearch.org/blog/theagentcompany-simulated-software-firm-benchmark/</guid><description>Carnegie Mellon built a fake company with a code host, a chat server, a file store and a project tracker, filled it with simulated coworkers, and asked agents to do 175 of its jobs. The best model finished about a quarter of them. What the design gets right and what partial credit means when the task is a job.</description><pubDate>Sat, 21 Dec 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>ModernBERT: the encoder side of the stack finally got an upgrade</title><link>https://montanaresearch.org/blog/modernbert-encoder-side-finally-upgraded/</link><guid isPermaLink="true">https://montanaresearch.org/blog/modernbert-encoder-side-finally-upgraded/</guid><description>BERT is still the second most downloaded model on the Hugging Face hub, six years after release. Answer.AI and LightOn trained a replacement on 2 trillion tokens with an 8,192 token context and the architecture tricks decoders got years ago. What changes for retrieval and classification pipelines.</description><pubDate>Fri, 20 Dec 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Byte Latent Transformer: patches instead of tokens</title><link>https://montanaresearch.org/blog/byte-latent-transformer-patches-instead-of-tokens/</link><guid isPermaLink="true">https://montanaresearch.org/blog/byte-latent-transformer-patches-instead-of-tokens/</guid><description>Meta&apos;s BLT drops the tokeniser and groups raw bytes into patches whose boundaries are set by next-byte entropy. At 8B parameters and 1T tokens of data it matches or beats a Llama 3 baseline on average while offering a way to trade accuracy for up to 50 percent fewer inference flops. Notes on the design and what the tables actually show.</description><pubDate>Wed, 18 Dec 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>SAEBench: the field finally gets a scoreboard</title><link>https://montanaresearch.org/blog/saebench-field-finally-gets-a-scoreboard/</link><guid isPermaLink="true">https://montanaresearch.org/blog/saebench-field-finally-gets-a-scoreboard/</guid><description>SAEBench launched this week with eight evaluations run across more than 200 sparse autoencoders. Its central finding is that reconstruction quality does not tell you whether the features are useful, which changes what an SAE paper has to show.</description><pubDate>Sat, 14 Dec 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>In-context scheming: five of six frontier models disabled their oversight</title><link>https://montanaresearch.org/blog/in-context-scheming-five-of-six-disabled-oversight/</link><guid isPermaLink="true">https://montanaresearch.org/blog/in-context-scheming-five-of-six-disabled-oversight/</guid><description>Apollo Research ran six agentic evaluations across o1, Claude, Gemini and Llama and found every model except GPT-4o took a covert action against its developers in at least one of them. The rates, the confession rates, and what the prompts did to get there.</description><pubDate>Mon, 09 Dec 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>What open source AI means now: OSAID, Llama and OLMo 2</title><link>https://montanaresearch.org/blog/what-open-source-ai-means-now/</link><guid isPermaLink="true">https://montanaresearch.org/blog/what-open-source-ai-means-now/</guid><description>OSI published a definition of open source AI in October, Meta said the bar was too narrow, the Software Freedom Conservancy said it was too low, and then Ai2 shipped OLMo 2 with everything on the table. The argument is now about a real model.</description><pubDate>Wed, 04 Dec 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Do I know this entity? Finding the hallucination switch</title><link>https://montanaresearch.org/blog/do-i-know-this-entity-hallucination-switch/</link><guid isPermaLink="true">https://montanaresearch.org/blog/do-i-know-this-entity-hallucination-switch/</guid><description>Ferrando, Obeso, Rajamanoharan and Nanda found sparse autoencoder latents in Gemma 2 that fire on whether the model recognises an entity, and showed that pushing on them flips the chat model between refusing and confabulating. A reading note on knowledge awareness as an internal signal.</description><pubDate>Tue, 26 Nov 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>RE-Bench: agents versus human experts on AI research tasks</title><link>https://montanaresearch.org/blog/re-bench-agents-versus-human-experts/</link><guid isPermaLink="true">https://montanaresearch.org/blog/re-bench-agents-versus-human-experts/</guid><description>METR gave language model agents and 61 paid human experts the same seven ML engineering environments and compared them by time budget. At two hours the agents scored four times higher. At 32 hours the humans scored twice as high. Why the curve is the result.</description><pubDate>Tue, 26 Nov 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Tülu 3 opens the post-training black box</title><link>https://montanaresearch.org/blog/tulu-3-opens-post-training-black-box/</link><guid isPermaLink="true">https://montanaresearch.org/blog/tulu-3-opens-post-training-black-box/</guid><description>AI2 released a complete post-training recipe, data, code and evaluation suite alongside the models. What it documents about supervised fine-tuning, preference tuning and verifiable rewards, and what the decontamination pass revealed about existing open datasets.</description><pubDate>Tue, 26 Nov 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Adding error bars to evals</title><link>https://montanaresearch.org/blog/adding-error-bars-to-evals/</link><guid isPermaLink="true">https://montanaresearch.org/blog/adding-error-bars-to-evals/</guid><description>Evan Miller&apos;s paper for Anthropic argues that most reported benchmark differences come without a confidence interval and many would not survive one. A walkthrough of the five recommendations, from treating questions as a sample to running a power analysis before you build the eval.</description><pubDate>Thu, 21 Nov 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The first joint government pre-deployment test: what the US and UK AISIs found in Claude 3.5 Sonnet</title><link>https://montanaresearch.org/blog/first-joint-government-pre-deployment-test-claude/</link><guid isPermaLink="true">https://montanaresearch.org/blog/first-joint-government-pre-deployment-test-claude/</guid><description>NIST and the UK AI Safety Institute published their joint evaluation of the October Claude 3.5 Sonnet. Reading notes on the cyber and software numbers against reference models, the jailbreak findings, and what the report is careful not to claim.</description><pubDate>Thu, 21 Nov 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Scaling laws for precision: why low-bit training has a floor</title><link>https://montanaresearch.org/blog/scaling-laws-for-precision-low-bit-floor/</link><guid isPermaLink="true">https://montanaresearch.org/blog/scaling-laws-for-precision-low-bit-floor/</guid><description>Kumar and colleagues fit scaling laws with precision as a third axis and found that more pretraining data makes post-training quantisation hurt more. An explainer on what the fitted forms predict for the move to FP8 and below.</description><pubDate>Thu, 21 Nov 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Gwern on a $12K a year salary: the outsider who saw scaling coming</title><link>https://montanaresearch.org/blog/gwern-outsider-who-saw-scaling-coming/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gwern-outsider-who-saw-scaling-coming/</guid><description>Dwarkesh Patel&apos;s interview with the anonymous writer behind gwern.net covers how the 2020 scaling hypothesis essay came about, what it costs to live on about $12,000 a year, and why he thinks independent writing still matters. Notes on what a researcher without a lab can actually contribute.</description><pubDate>Tue, 19 Nov 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>SimpleQA: measuring factuality with a benchmark models were meant to fail</title><link>https://montanaresearch.org/blog/simpleqa-benchmark-models-were-meant-to-fail/</link><guid isPermaLink="true">https://montanaresearch.org/blog/simpleqa-benchmark-models-were-meant-to-fail/</guid><description>OpenAI&apos;s new short-form factuality set was built by throwing out any question GPT-4 could answer, then graded with three labels instead of two. Notes on why the third label, not attempted, is the interesting part.</description><pubDate>Tue, 12 Nov 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Nobody is ready for AGI: Miles Brundage leaves OpenAI</title><link>https://montanaresearch.org/blog/nobody-ready-for-agi-brundage-leaves-openai/</link><guid isPermaLink="true">https://montanaresearch.org/blog/nobody-ready-for-agi-brundage-leaves-openai/</guid><description>The head of OpenAI&apos;s AGI Readiness team left this week saying neither his employer nor the world is prepared, and that the research he wants to do has to happen outside a lab. Some thoughts from a foundation that made the same bet on independence.</description><pubDate>Tue, 29 Oct 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Crosscoders and model diffing: features that span layers and checkpoints</title><link>https://montanaresearch.org/blog/crosscoders-model-diffing-features-across-layers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/crosscoders-model-diffing-features-across-layers/</guid><description>Anthropic trained one sparse dictionary across every layer of a model, and then across a base model and its finetuned version. Notes on what a crosscoder is, what it found, and why comparing a model against its predecessor may be the most useful thing dictionary learning does.</description><pubDate>Mon, 28 Oct 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Automated interpretability for millions of features, on open models</title><link>https://montanaresearch.org/blog/automated-interpretability-millions-features-open-models/</link><guid isPermaLink="true">https://montanaresearch.org/blog/automated-interpretability-millions-features-open-models/</guid><description>EleutherAI&apos;s new pipeline explains sparse autoencoder latents with open models and scores the explanations for a few hundred dollars per million features instead of tens of thousands. Notes on what changed for labs without a frontier API budget, and on how the choice of scorer decides which features count as interpretable.</description><pubDate>Thu, 24 Oct 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Computer use: the agent that clicks</title><link>https://montanaresearch.org/blog/computer-use-the-agent-that-clicks/</link><guid isPermaLink="true">https://montanaresearch.org/blog/computer-use-the-agent-that-clicks/</guid><description>Anthropic shipped an API that lets a model look at a screenshot, move a cursor and type. It scores 14.9 percent on OSWorld against a human range of 70 to 75, it cannot drag or zoom, and during a demo it wandered off to look at pictures of Yellowstone. Notes on the first general-purpose screen agent and why the gap between demo and reliability is so wide.</description><pubDate>Thu, 24 Oct 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Sabotage evaluations and RSP 2.0: testing whether a model would undermine you</title><link>https://montanaresearch.org/blog/sabotage-evaluations-and-rsp-2/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sabotage-evaluations-and-rsp-2/</guid><description>Anthropic released four evaluations for whether a model could sabotage decisions, code, its own capability evals, or an oversight process, three days after rewriting its Responsible Scaling Policy around capability thresholds and safety cases. What each eval measures, what Claude 3 Opus and 3.5 Sonnet scored, and how the two documents fit together.</description><pubDate>Thu, 24 Oct 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Swarm and AutoGen: the multi-agent framework question</title><link>https://montanaresearch.org/blog/swarm-and-autogen-multi-agent-framework-question/</link><guid isPermaLink="true">https://montanaresearch.org/blog/swarm-and-autogen-multi-agent-framework-question/</guid><description>OpenAI&apos;s Swarm is a small library built on two primitives, agents and handoffs, and is labelled educational. Microsoft&apos;s AutoGen is a conversation-programming framework with a 43-page paper behind it. A comparison, and a question about whether multi-agent orchestration is an abstraction or a workaround.</description><pubDate>Thu, 17 Oct 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The Nobel prizes that went to neural networks</title><link>https://montanaresearch.org/blog/the-nobel-prizes-that-went-to-neural-networks/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-nobel-prizes-that-went-to-neural-networks/</guid><description>This month the physics and chemistry Nobels both went to work on neural networks. The argument about whether that is really physics or chemistry is less interesting than what it says about AI as a scientific instrument.</description><pubDate>Thu, 17 Oct 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Machines of Loving Grace and the marginal returns to intelligence</title><link>https://montanaresearch.org/blog/machines-of-loving-grace-marginal-returns-to-intelligence/</link><guid isPermaLink="true">https://montanaresearch.org/blog/machines-of-loving-grace-marginal-returns-to-intelligence/</guid><description>Dario Amodei&apos;s essay argues that powerful AI could compress fifty to a hundred years of biology into five to ten, and then lists the things that intelligence alone cannot speed up. Reading notes on that list, which is the most useful part of the essay, and on the entente section, which is the most contested.</description><pubDate>Sun, 13 Oct 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>GSM-Symbolic: the reasoning paper everyone cited for the wrong reason</title><link>https://montanaresearch.org/blog/gsm-symbolic-reasoning-paper-cited-for-the-wrong-reason/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gsm-symbolic-reasoning-paper-cited-for-the-wrong-reason/</guid><description>Apple researchers turned 100 GSM8K problems into templates, regenerated them with new names and numbers, and watched accuracy fall. The paper was read as proof that models cannot reason. What it measures is narrower and more useful.</description><pubDate>Fri, 11 Oct 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>A is for absorption: the feature that swallowed its own letter</title><link>https://montanaresearch.org/blog/a-is-for-absorption-feature-swallowed-its-letter/</link><guid isPermaLink="true">https://montanaresearch.org/blog/a-is-for-absorption-feature-swallowed-its-letter/</guid><description>Chanin and colleagues show that a sparse autoencoder latent which looks like a clean starts-with-S detector fails to fire on the token short, because a more specific latent absorbed the direction. The paper argues this follows from the sparsity objective and does not go away with width or L0.</description><pubDate>Mon, 30 Sep 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>AI Snake Oil: the book that tried to separate predictive from generative</title><link>https://montanaresearch.org/blog/ai-snake-oil-predictive-vs-generative/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-snake-oil-predictive-vs-generative/</guid><description>Narayanan and Kapoor&apos;s book draws one line, between AI that predicts outcomes about people and AI that generates text and images, and argues most of the harm is on the side nobody talks about. Reading notes.</description><pubDate>Mon, 30 Sep 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>SB 1047 vetoed: the frontier-model bill that split the safety community</title><link>https://montanaresearch.org/blog/sb-1047-vetoed-frontier-model-bill-split-safety/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sb-1047-vetoed-frontier-model-bill-split-safety/</guid><description>Newsom returned SB 1047 unsigned on September 29, arguing that a bill keyed to training cost and compute ignores where a model is deployed. A look at what the bill required, who lined up on each side, and why the veto message reads like a research question.</description><pubDate>Mon, 30 Sep 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>SB 1047 and the open-weights question it never resolved</title><link>https://montanaresearch.org/blog/sb-1047-open-weights-question-never-resolved/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sb-1047-open-weights-question-never-resolved/</guid><description>California&apos;s frontier model bill was vetoed this weekend. The fight over who answers for a downloaded checkpoint was written into its definitions and the veto leaves it exactly where it was.</description><pubDate>Sun, 29 Sep 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Molmo and PixMo: open VLMs without distilling from closed ones</title><link>https://montanaresearch.org/blog/molmo-pixmo-open-vlms-without-distilling/</link><guid isPermaLink="true">https://montanaresearch.org/blog/molmo-pixmo-open-vlms-without-distilling/</guid><description>AI2 collected 1.3 million image captions by having annotators talk for a minute instead of prompting GPT-4V, and trained a 72B model that beats Claude 3.5 Sonnet and Gemini 1.5 Pro on their benchmark suite. Notes on why refusing to distill was a licensing decision and a scientific one at the same time.</description><pubDate>Fri, 27 Sep 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Contextual retrieval and the return of BM25</title><link>https://montanaresearch.org/blog/contextual-retrieval-return-of-bm25/</link><guid isPermaLink="true">https://montanaresearch.org/blog/contextual-retrieval-return-of-bm25/</guid><description>Anthropic&apos;s contextual retrieval note describes a cheap trick, prepending a sentence of document context to every chunk before indexing, and reports a 67 percent drop in retrieval failures once BM25 and a reranker are added. It reads like a summary of what production RAG has settled on.</description><pubDate>Thu, 26 Sep 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>PlanBench meets o1: can a reasoning model plan?</title><link>https://montanaresearch.org/blog/planbench-meets-o1-can-reasoning-model-plan/</link><guid isPermaLink="true">https://montanaresearch.org/blog/planbench-meets-o1-can-reasoning-model-plan/</guid><description>Kambhampati&apos;s group re-ran their planning benchmark on o1 within a week of release. The jump on Blocksworld is real, and so is the collapse on obfuscated and longer problems. Notes on what an old hard benchmark can tell you about a new kind of model.</description><pubDate>Thu, 26 Sep 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>o1 and the new axis: scaling test-time compute</title><link>https://montanaresearch.org/blog/o1-scaling-test-time-compute/</link><guid isPermaLink="true">https://montanaresearch.org/blog/o1-scaling-test-time-compute/</guid><description>OpenAI says o1 gets better the longer it thinks, and that it was trained with reinforcement learning on its own chain of thought. Here is what that claim means and what remains hidden.</description><pubDate>Tue, 24 Sep 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The o1 system card: deliberative alignment and a chain of thought you cannot see</title><link>https://montanaresearch.org/blog/o1-system-card-hidden-chain-of-thought/</link><guid isPermaLink="true">https://montanaresearch.org/blog/o1-system-card-hidden-chain-of-thought/</guid><description>Reading notes on the o1 system card. The model reasons about safety policy in a chain of thought that users only see summarised, Apollo Research found it faking alignment in toy settings, and OpenAI rated it medium risk and shipped it.</description><pubDate>Tue, 17 Sep 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Style control: what happens to the Arena when you subtract markdown and length</title><link>https://montanaresearch.org/blog/style-control-arena-subtract-markdown-and-length/</link><guid isPermaLink="true">https://montanaresearch.org/blog/style-control-arena-subtract-markdown-and-length/</guid><description>LMSYS added four style covariates to its Bradley-Terry model and reran the leaderboard. Some models fell a dozen places. Notes on the method, the coefficients, and how much of a human preference vote is presentation.</description><pubDate>Fri, 30 Aug 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Gemma Scope puts sparse autoencoders in reach of anyone with a GPU</title><link>https://montanaresearch.org/blog/gemma-scope-open-saes-for-everyone/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemma-scope-open-saes-for-everyone/</guid><description>DeepMind released over 400 sparse autoencoders covering every layer of Gemma 2, trained with more than a fifth of the compute that went into GPT-3. What that changes for groups outside the big labs.</description><pubDate>Tue, 20 Aug 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The AI Scientist writes a paper for $15: what peer review is for</title><link>https://montanaresearch.org/blog/ai-scientist-what-peer-review-is-for/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-scientist-what-peer-review-is-for/</guid><description>Sakana&apos;s pipeline generates ideas, runs experiments, writes the paper and reviews it, for around fifteen dollars a paper. The interesting question is what conference review becomes when papers can be produced faster than anyone can read them.</description><pubDate>Mon, 19 Aug 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Prompt caching changed the economics of long system prompts</title><link>https://montanaresearch.org/blog/prompt-caching-changed-the-economics-of-long-system-prompts/</link><guid isPermaLink="true">https://montanaresearch.org/blog/prompt-caching-changed-the-economics-of-long-system-prompts/</guid><description>Anthropic now lets you pay a 25 percent premium to write a prompt prefix once and then read it back at a tenth of the input price. That one pricing change makes a whole class of prompt designs affordable that were previously too expensive to run at volume.</description><pubDate>Mon, 19 Aug 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Falcon Mamba 7B: the first attention-free model that could hold its own</title><link>https://montanaresearch.org/blog/falcon-mamba-7b-attention-free-model-holds-own/</link><guid isPermaLink="true">https://montanaresearch.org/blog/falcon-mamba-7b-attention-free-model-holds-own/</guid><description>TII trained a pure Mamba 7B on about 5,500 gigatokens across 256 H100s and scored 64.09 on the legacy Open LLM Leaderboard, above Llama 3.1 8B and Mistral 7B. What that settles about state space models, and what the hybrid architectures already told us it does not.</description><pubDate>Sat, 17 Aug 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The AI Scientist: fifteen dollars a paper and a reviewer to match</title><link>https://montanaresearch.org/blog/ai-scientist-fifteen-dollars-a-paper/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-scientist-fifteen-dollars-a-paper/</guid><description>Sakana AI, with collaborators at Oxford and UBC, released a pipeline that generates ideas, runs experiments, writes the paper and reviews it, for under 15 dollars per paper. It also edited its own launch script. Notes on what the demo showed, what it did not, and what cheap papers do to a field already short of reviewers.</description><pubDate>Thu, 15 Aug 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>From function calling to structured outputs: making JSON a contract</title><link>https://montanaresearch.org/blog/from-function-calling-to-structured-outputs/</link><guid isPermaLink="true">https://montanaresearch.org/blog/from-function-calling-to-structured-outputs/</guid><description>OpenAI now guarantees that model output matches a JSON schema. That fixes syntax, which was most of the pain. It does nothing for the values inside, which is where the failures have moved.</description><pubDate>Fri, 09 Aug 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Large Language Monkeys: coverage scales with samples, and that changes the question</title><link>https://montanaresearch.org/blog/large-language-monkeys-coverage-scales-with-samples/</link><guid isPermaLink="true">https://montanaresearch.org/blog/large-language-monkeys-coverage-scales-with-samples/</guid><description>Brown et al. show that the fraction of problems solved by at least one of k samples grows log-linearly over four orders of magnitude of k, even for weak models. Notes on why this turns capability into a verification problem.</description><pubDate>Wed, 31 Jul 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The Llama 3 herd paper: a frontier training run described end to end</title><link>https://montanaresearch.org/blog/llama-3-herd-paper-frontier-run-end-to-end/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-3-herd-paper-frontier-run-end-to-end/</guid><description>The Llama 3.1 report described how Meta picked 405B from a scaling law, ran 16,000 H100s through 466 interruptions in 54 days, annealed on high-quality data, and iterated post-training six times. A look back at what the most detailed open account of a frontier run gave the rest of the field.</description><pubDate>Wed, 31 Jul 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Consent in crisis: the year the web started saying no</title><link>https://montanaresearch.org/blog/consent-in-crisis-the-web-started-saying-no/</link><guid isPermaLink="true">https://montanaresearch.org/blog/consent-in-crisis-the-web-started-saying-no/</guid><description>The Data Provenance Initiative audited 14,000 domains behind C4, RefinedWeb and Dolma and found that in one year more than 5 percent of C4 tokens and over a quarter of its best-maintained sources were newly blocked by robots.txt. Reading notes on who the blocks actually hit.</description><pubDate>Fri, 26 Jul 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Zuckerberg&apos;s Linux analogy: the business case for open weights</title><link>https://montanaresearch.org/blog/zuckerberg-linux-analogy-open-weights-business-case/</link><guid isPermaLink="true">https://montanaresearch.org/blog/zuckerberg-linux-analogy-open-weights-business-case/</guid><description>Meta shipped Llama 3.1 405B this week with a 2,000 word essay arguing that open AI will win the way Linux won. A close reading, sorting the arguments that hold from the ones that do not survive a look at the license.</description><pubDate>Thu, 25 Jul 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>JumpReLU: a threshold, a straight-through estimator, and a better Pareto frontier</title><link>https://montanaresearch.org/blog/jumprelu-threshold-straight-through-estimator-pareto/</link><guid isPermaLink="true">https://montanaresearch.org/blog/jumprelu-threshold-straight-through-estimator-pareto/</guid><description>DeepMind replaced the ReLU in a sparse autoencoder with a learned per-feature threshold and trained the L0 penalty directly using straight-through gradients. A method deep dive on why that beat Gated SAEs and matched TopK on Gemma 2 9B at the same sparsity.</description><pubDate>Wed, 24 Jul 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Do circuits survive training and scale? Pythia says mostly yes</title><link>https://montanaresearch.org/blog/do-circuits-survive-training-and-scale-pythia/</link><guid isPermaLink="true">https://montanaresearch.org/blog/do-circuits-survive-training-and-scale-pythia/</guid><description>Tigges, Hanna, Yu and Biderman tracked four circuits across 154 Pythia checkpoints and five model sizes and found the same algorithms, built from the same kinds of heads, appearing at about the same token count. A quiet result that matters for anyone hoping findings on small models transfer.</description><pubDate>Tue, 23 Jul 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>AI agents that matter: benchmarks that ignore cost are not measuring agents</title><link>https://montanaresearch.org/blog/ai-agents-that-matter-benchmarks-ignore-cost/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-agents-that-matter-benchmarks-ignore-cost/</guid><description>Kapoor, Narayanan and colleagues show that on HumanEval a simple retry-and-warm baseline beats published agent architectures at a fraction of the cost, and that seven of eight domain-general benchmarks have no holdout set. Accuracy-only leaderboards are rewarding the wrong thing.</description><pubDate>Tue, 09 Jul 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>GraphRAG: does building a knowledge graph fix retrieval?</title><link>https://montanaresearch.org/blog/graphrag-does-a-knowledge-graph-fix-retrieval/</link><guid isPermaLink="true">https://montanaresearch.org/blog/graphrag-does-a-knowledge-graph-fix-retrieval/</guid><description>Microsoft&apos;s GraphRAG extracts an entity graph from a corpus, clusters it into communities, summarises each community, and answers global questions by map-reduce over the summaries. A method deep dive on what it wins, what the index costs, and why the evaluation leaves the central question open.</description><pubDate>Tue, 09 Jul 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>FineWeb: what it takes to decant the web into 15 trillion good tokens</title><link>https://montanaresearch.org/blog/fineweb-decanting-web-into-15-trillion-tokens/</link><guid isPermaLink="true">https://montanaresearch.org/blog/fineweb-decanting-web-into-15-trillion-tokens/</guid><description>Hugging Face released a 15 trillion token pretraining corpus with an ablation for every filtering and deduplication choice, and a 1.3 trillion token subset selected by a classifier trained on LLM judgements of educational quality. Reading notes on why data curation is now a research field with its own results.</description><pubDate>Thu, 27 Jun 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Ilya&apos;s straight shot: what a lab with no product is betting on</title><link>https://montanaresearch.org/blog/ilyas-straight-shot-lab-with-no-product/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ilyas-straight-shot-lab-with-no-product/</guid><description>Safe Superintelligence Inc. launched this week with one goal, one product, and a promise of no commercial pressure until it ships. The interesting claim is about research culture, and I am not sure a corporate structure can deliver it.</description><pubDate>Mon, 24 Jun 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Refusal is one direction, and circuit breakers try to build on that</title><link>https://montanaresearch.org/blog/refusal-is-one-direction-circuit-breakers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/refusal-is-one-direction-circuit-breakers/</guid><description>Arditi and colleagues show that refusal in thirteen open chat models is mediated by a single direction in the residual stream that can be erased. Zou and colleagues propose rerouting harmful representations instead of training refusals. Read together, the two papers explain why safety fine-tuning is brittle and offer one way past it.</description><pubDate>Mon, 24 Jun 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The $600B question, read from the cheap seats</title><link>https://montanaresearch.org/blog/the-600b-question-read-from-the-cheap-seats/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-600b-question-read-from-the-cheap-seats/</guid><description>Sequoia&apos;s David Cahn estimates a 600 billion dollar gap between what has been spent on AI compute and the revenue it needs to earn back. Notes on what the bubble argument means for a lab that rents GPUs by the hour and would benefit from every part of the correction.</description><pubDate>Mon, 24 Jun 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>From sycophancy to subterfuge: 45 reward-tampering runs out of 32,768</title><link>https://montanaresearch.org/blog/from-sycophancy-to-subterfuge-45-runs-out-of-32768/</link><guid isPermaLink="true">https://montanaresearch.org/blog/from-sycophancy-to-subterfuge-45-runs-out-of-32768/</guid><description>Anthropic trained a model through a curriculum of small cheats and then handed it its own reward code. It edited the code in 45 of 32,768 trials and covered its tracks in 7. An explainer on the setup, and an argument about how much weight a rare but nonzero result should carry.</description><pubDate>Fri, 21 Jun 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Transcoders: making MLPs legible without the activations</title><link>https://montanaresearch.org/blog/transcoders-making-mlps-legible-without-activations/</link><guid isPermaLink="true">https://montanaresearch.org/blog/transcoders-making-mlps-legible-without-activations/</guid><description>Dunefsky, Chlenski and Nanda train sparse replacements for MLP layers that map inputs to outputs directly. The payoff is that feature-to-feature connections can be read from the weights, before you ever look at an example.</description><pubDate>Thu, 20 Jun 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The data wall: will we run out of text?</title><link>https://montanaresearch.org/blog/the-data-wall/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-data-wall/</guid><description>Epoch AI now estimates the stock of public human text at about 300 trillion tokens and projects it will be fully used between 2026 and 2032. Here is what that number assumes and what labs are doing about it.</description><pubDate>Wed, 19 Jun 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Apple Intelligence: a 3B model, many adapters, and a private cloud</title><link>https://montanaresearch.org/blog/apple-intelligence-3b-model-adapters-private-cloud/</link><guid isPermaLink="true">https://montanaresearch.org/blog/apple-intelligence-3b-model-adapters-private-cloud/</guid><description>Apple described a roughly 3 billion parameter on-device model, quantized to an average of 3.7 bits per weight, with rank 16 LoRA adapters swapped in per task and a server model behind Private Cloud Compute. Notes on what makes it the most concrete on-device architecture anyone has shipped.</description><pubDate>Fri, 14 Jun 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>ARC Prize launches: a million dollars for a benchmark AI could not pass</title><link>https://montanaresearch.org/blog/arc-prize-launches-million-dollars-benchmark/</link><guid isPermaLink="true">https://montanaresearch.org/blog/arc-prize-launches-million-dollars-benchmark/</guid><description>Francois Chollet and Mike Knoop have put a prize pool of over a million dollars on ARC-AGI, where the state of the art sits at 34 percent. A look at the prize as an evaluation design, and at the assumptions it will test.</description><pubDate>Thu, 13 Jun 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>MMLU-Pro and the art of un-saturating a benchmark</title><link>https://montanaresearch.org/blog/mmlu-pro-un-saturating-a-benchmark/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mmlu-pro-un-saturating-a-benchmark/</guid><description>MMLU-Pro raised the answer count from four to ten, dropped the questions most models got right, and cut top scores by 16 to 33 points. Notes on the mechanics of extending a benchmark&apos;s life and on what comparability that costs.</description><pubDate>Wed, 12 Jun 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Sixteen million features in GPT-4: OpenAI&apos;s TopK sparse autoencoders</title><link>https://montanaresearch.org/blog/sixteen-million-features-gpt-4-topk-sparse-autoencoders/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sixteen-million-features-gpt-4-topk-sparse-autoencoders/</guid><description>Gao and colleagues replaced the L1 penalty with a hard top-k selection, fixed the dead-latent problem, and trained a 16 million latent autoencoder on GPT-4. What the method changes, what the scaling laws say, and where the paper admits the features are still not clean.</description><pubDate>Wed, 12 Jun 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>A right to warn: when lab employees asked for whistleblower rules</title><link>https://montanaresearch.org/blog/right-to-warn-lab-employees-whistleblower-rules/</link><guid isPermaLink="true">https://montanaresearch.org/blog/right-to-warn-lab-employees-whistleblower-rules/</guid><description>Thirteen current and former staff of OpenAI, Google DeepMind and Anthropic published an open letter asking frontier labs to drop non-disparagement clauses and build anonymous channels for reporting risk. An opinion piece on what the letter says about research culture inside closed labs, and what those of us outside can do.</description><pubDate>Thu, 06 Jun 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Glue on pizza: AI Overviews and search as a deployment surface</title><link>https://montanaresearch.org/blog/glue-on-pizza-ai-overviews-deployment-surface/</link><guid isPermaLink="true">https://montanaresearch.org/blog/glue-on-pizza-ai-overviews-deployment-surface/</guid><description>Google put a language model in front of every search query, and within two weeks the answers about rocks and pizza glue were everywhere. Google&apos;s own post-mortem blames satire, forums and data voids. A look at what it means to deploy a model on the largest surface there is.</description><pubDate>Fri, 31 May 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Days of the week live on a circle: not all features are linear</title><link>https://montanaresearch.org/blog/days-of-the-week-live-on-a-circle/</link><guid isPermaLink="true">https://montanaresearch.org/blog/days-of-the-week-live-on-a-circle/</guid><description>Engels and colleagues used sparse autoencoders to find features that are genuinely two-dimensional, including a circle of weekdays that Mistral and Llama use to add days. A look at what this does to the linear representation hypothesis.</description><pubDate>Wed, 29 May 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Golden Gate Claude and the knob that everyone now wants</title><link>https://montanaresearch.org/blog/golden-gate-claude-first-feature-knob/</link><guid isPermaLink="true">https://montanaresearch.org/blog/golden-gate-claude-first-feature-knob/</guid><description>Anthropic clamped one feature in Claude 3 Sonnet to ten times its maximum and put the result online for a day. What the demo shows about steering, and what it hides.</description><pubDate>Wed, 29 May 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Lessons from the trenches: what running lm-evaluation-harness taught EleutherAI</title><link>https://montanaresearch.org/blog/lessons-from-the-trenches-lm-evaluation-harness/</link><guid isPermaLink="true">https://montanaresearch.org/blog/lessons-from-the-trenches-lm-evaluation-harness/</guid><description>The maintainers of lm-eval wrote down three years of reproducibility failures: the same model scoring 38.0 or 26.6 on ARC depending on the prompt style, MMLU averages that differ by several points depending on how you aggregate, and single-run scores that reverse under a confidence interval. Reading notes and the rules I take from it.</description><pubDate>Tue, 28 May 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>LoRA learns less and forgets less</title><link>https://montanaresearch.org/blog/lora-learns-less-and-forgets-less/</link><guid isPermaLink="true">https://montanaresearch.org/blog/lora-learns-less-and-forgets-less/</guid><description>Reading notes on the Databricks comparison of LoRA and full finetuning for Llama-2-7B on code and math. Adapters lag badly in continued pretraining, nearly catch up in instruction tuning at high rank, and keep more of the base model either way.</description><pubDate>Tue, 28 May 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Superalignment, ten months later</title><link>https://montanaresearch.org/blog/superalignment-ten-months-later/</link><guid isPermaLink="true">https://montanaresearch.org/blog/superalignment-ten-months-later/</guid><description>In July 2023 OpenAI promised a fifth of its secured compute to aligning superintelligence within four years. This week the co-lead resigned saying safety had taken a backseat to shiny products, and the team was folded. What a public research commitment is worth when nobody outside can check it.</description><pubDate>Tue, 21 May 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Model Spec: writing down what the model is supposed to do</title><link>https://montanaresearch.org/blog/model-spec-writing-down-what-model-should-do/</link><guid isPermaLink="true">https://montanaresearch.org/blog/model-spec-writing-down-what-model-should-do/</guid><description>OpenAI has published a document stating how it wants its models to behave, organised into three objectives, six rules and ten defaults, with a chain of command for resolving conflicts. Reading notes on what a public behaviour spec changes about alignment arguments.</description><pubDate>Tue, 14 May 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Stack Overflow sells its answers to OpenAI, and the moderators revolt</title><link>https://montanaresearch.org/blog/stack-overflow-sells-answers-moderators-revolt/</link><guid isPermaLink="true">https://montanaresearch.org/blog/stack-overflow-sells-answers-moderators-revolt/</guid><description>On May 6 Stack Overflow announced an API partnership with OpenAI. Within days users were rewriting and deleting their best answers in protest, and the site was restoring them and suspending the authors. What the episode shows about who owns community-contributed knowledge once it becomes training data.</description><pubDate>Tue, 14 May 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>DeepSeek-V2 and multi-head latent attention: the KV cache as the real bottleneck</title><link>https://montanaresearch.org/blog/deepseek-v2-multi-head-latent-attention-kv-cache/</link><guid isPermaLink="true">https://montanaresearch.org/blog/deepseek-v2-multi-head-latent-attention-kv-cache/</guid><description>DeepSeek-V2 is a 236B mixture-of-experts model that activates 21B parameters per token, cuts the KV cache by 93.3 percent with a new attention design, and trains on 42.5 percent fewer GPU hours per token than the lab&apos;s own dense 67B. A deep dive on multi-head latent attention and why this paper is where the efficiency playbook became visible.</description><pubDate>Sat, 11 May 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>GSM1k: rebuilding a benchmark to find out who overfit it</title><link>https://montanaresearch.org/blog/gsm1k-rebuilding-a-benchmark-to-find-who-overfit/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gsm1k-rebuilding-a-benchmark-to-find-who-overfit/</guid><description>Scale AI commissioned a fresh set of grade-school maths problems matched to GSM8K and found that some model families drop up to 13 points on it. A replication story about what a held-out twin dataset can reveal about contamination, and what it cannot.</description><pubDate>Mon, 06 May 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Sleeper agents: safety training cannot remove a backdoor it cannot see</title><link>https://montanaresearch.org/blog/sleeper-agents-safety-training-cannot-remove-what-it-cannot-see/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sleeper-agents-safety-training-cannot-remove-what-it-cannot-see/</guid><description>Anthropic trained models to write vulnerable code when told the year is 2024, then tried to train the behaviour out. It survived. A follow-up found a linear probe that spots the defection, with a catch that matters for anyone trusting fine-tuned open weights.</description><pubDate>Mon, 29 Apr 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>&apos;Built with Meta Llama 3&apos;: the license clause that names your model</title><link>https://montanaresearch.org/blog/built-with-meta-llama-3-license-clause/</link><guid isPermaLink="true">https://montanaresearch.org/blog/built-with-meta-llama-3-license-clause/</guid><description>The Llama 3 Community License grants a royalty-free license and then attaches conditions most open licenses never had, including a required product credit, a required model name prefix, and a ban on using outputs to improve other models. A reading of what the text actually says.</description><pubDate>Wed, 24 Apr 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Gated SAEs and the shrinkage problem</title><link>https://montanaresearch.org/blog/gated-saes-and-the-shrinkage-problem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gated-saes-and-the-shrinkage-problem/</guid><description>The L1 penalty that makes sparse autoencoders sparse also makes them systematically underestimate how strongly each feature is active. Rajamanoharan and colleagues at Google DeepMind split the encoder into a gate and a magnitude estimate and get a Pareto improvement. It also marks the point where SAE architecture became its own research question.</description><pubDate>Wed, 24 Apr 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Phi-3 and the model that runs locally on your phone</title><link>https://montanaresearch.org/blog/phi-3-and-the-model-that-runs-on-your-phone/</link><guid isPermaLink="true">https://montanaresearch.org/blog/phi-3-and-the-model-that-runs-on-your-phone/</guid><description>Microsoft&apos;s phi-3-mini has 3.8B parameters, quantises to about 1.8GB, runs at more than 12 tokens per second on an iPhone 14, and reports 69 percent on MMLU and 8.38 on MT-bench. A look at the curated-data recipe behind it and at the gap between those scores and what the paper itself admits the model cannot do.</description><pubDate>Wed, 24 Apr 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Llama 3 at 15 trillion tokens: how far past Chinchilla can you go?</title><link>https://montanaresearch.org/blog/llama-3-15-trillion-tokens-past-chinchilla/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-3-15-trillion-tokens-past-chinchilla/</guid><description>Meta trained an 8B model on roughly 75 times the compute-optimal token count and says it was still improving. Notes on why that is a rational choice once you count inference, and what the announcement does and does not tell us.</description><pubDate>Mon, 22 Apr 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Zuckerberg on gigawatt datacenters: the compute conversation goes public</title><link>https://montanaresearch.org/blog/zuckerberg-gigawatt-datacenters-compute-goes-public/</link><guid isPermaLink="true">https://montanaresearch.org/blog/zuckerberg-gigawatt-datacenters-compute-goes-public/</guid><description>In the Llama 3 interview, Mark Zuckerberg said nobody has built a gigawatt datacenter yet and that energy permitting, not GPU supply, is the constraint he sees. Notes on hearing a training run described in units of power plants.</description><pubDate>Mon, 22 Apr 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>51 to 15: the 2024 AI Index and where the models come from</title><link>https://montanaresearch.org/blog/51-to-15-ai-index-where-models-come-from/</link><guid isPermaLink="true">https://montanaresearch.org/blog/51-to-15-ai-index-where-models-come-from/</guid><description>Stanford HAI&apos;s 2024 report counts 51 notable models from industry in 2023 against 15 from academia, and estimates GPT-4&apos;s training compute at 78 million dollars. A personal take on what academic AI research can mean when the objects of study cost that much.</description><pubDate>Wed, 17 Apr 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Length-controlled AlpacaEval: fixing the judge that loved long answers</title><link>https://montanaresearch.org/blog/length-controlled-alpacaeval-judge-loved-long-answers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/length-controlled-alpacaeval-judge-loved-long-answers/</guid><description>AlpacaEval win rates could be moved by 40 points with a single prompt asking for more detail. Dubois and colleagues regressed length out of the judge and the correlation with human votes went up. Notes on the fix and what it says about every model-judged leaderboard.</description><pubDate>Thu, 11 Apr 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Many-shot jailbreaking: when the context window is the attack</title><link>https://montanaresearch.org/blog/many-shot-jailbreaking-context-window-is-the-attack/</link><guid isPermaLink="true">https://montanaresearch.org/blog/many-shot-jailbreaking-context-window-is-the-attack/</guid><description>Anthropic showed that filling a long prompt with hundreds of fake dialogues in which an assistant answers harmful questions overrides safety training, and that the effect follows a power law in the number of shots. The defences that work cut against the long context features labs are selling.</description><pubDate>Tue, 09 Apr 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>How LLMs actually think: notes on the Sholto and Trenton conversation</title><link>https://montanaresearch.org/blog/how-llms-think-sholto-trenton-notes/</link><guid isPermaLink="true">https://montanaresearch.org/blog/how-llms-think-sholto-trenton-notes/</guid><description>A long Dwarkesh Patel episode with Sholto Douglas of Google and Trenton Bricken of Anthropic has become the thing people send newcomers. Reading notes on what it claims about in-context learning, long context, superposition and the daily work of research.</description><pubDate>Sun, 31 Mar 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Sparse feature circuits and SHIFT: removing a spurious feature by hand</title><link>https://montanaresearch.org/blog/sparse-feature-circuits-shift-removing-spurious-feature/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sparse-feature-circuits-shift-removing-spurious-feature/</guid><description>Marks and colleagues built causal circuits out of sparse autoencoder features instead of neurons, then used them to find and delete the gender shortcut a profession classifier had learned. It is the first case I know of where dictionary features were used to edit a model on purpose, and the demo is small enough that the open questions are easy to list.</description><pubDate>Sat, 30 Mar 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Jamba: the first production-scale hybrid of attention and state space layers</title><link>https://montanaresearch.org/blog/jamba-first-production-hybrid-attention-state-space/</link><guid isPermaLink="true">https://montanaresearch.org/blog/jamba-first-production-hybrid-attention-state-space/</guid><description>AI21 interleaved Mamba layers, attention layers and mixture-of-experts in a 52 billion parameter model with a 256K context that runs on one 80GB GPU. A method piece on the layer ratio, the ablations, and why a hybrid rather than a pure state space model turned out to be the practical route.</description><pubDate>Fri, 29 Mar 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Blackwell and the year compute became a policy variable</title><link>https://montanaresearch.org/blog/blackwell-year-compute-became-policy-variable/</link><guid isPermaLink="true">https://montanaresearch.org/blog/blackwell-year-compute-became-policy-variable/</guid><description>Nvidia&apos;s GB200 NVL72 promises up to 30 times the inference throughput of H100 and every major cloud has already signed on. What a single vendor&apos;s roadmap means for who gets to do frontier research, and how a small non-profit should plan underneath it.</description><pubDate>Thu, 21 Mar 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Grok-1 by torrent: 314 billion parameters, Apache 2.0, no paper</title><link>https://montanaresearch.org/blog/grok-1-by-torrent-open-weights-no-paper/</link><guid isPermaLink="true">https://montanaresearch.org/blog/grok-1-by-torrent-open-weights-no-paper/</guid><description>xAI put a 314 billion parameter base checkpoint on a magnet link under the most permissive licence going. What it left out shows the gap between open weights and open science.</description><pubDate>Tue, 19 Mar 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The EU AI Act passes: what a risk-tier law means for frontier models</title><link>https://montanaresearch.org/blog/eu-ai-act-passes-risk-tier-law-frontier-models/</link><guid isPermaLink="true">https://montanaresearch.org/blog/eu-ai-act-passes-risk-tier-law-frontier-models/</guid><description>The European Parliament adopted the AI Act by 523 votes to 46 on March 13. An explainer on the risk tiers, the general-purpose model provisions that were added late, and the staggered timeline that decides when any of it starts to bind.</description><pubDate>Fri, 15 Mar 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Up to 17 percent of ICLR 2024 reviews were touched by an LLM</title><link>https://montanaresearch.org/blog/17-percent-of-iclr-2024-reviews-touched-by-llm/</link><guid isPermaLink="true">https://montanaresearch.org/blog/17-percent-of-iclr-2024-reviews-touched-by-llm/</guid><description>Liang and colleagues estimate the share of LLM-modified text in a corpus without classifying any single document, and find between 6.5 and 16.9 percent of recent AI conference reviews were substantially modified. The usage clusters near deadlines and in low-confidence reviews. Reading notes and what a conference can do.</description><pubDate>Thu, 14 Mar 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>WMDP and unlearning: measuring hazardous knowledge without publishing it</title><link>https://montanaresearch.org/blog/wmdp-unlearning-measuring-hazardous-knowledge/</link><guid isPermaLink="true">https://montanaresearch.org/blog/wmdp-unlearning-measuring-hazardous-knowledge/</guid><description>The Weapons of Mass Destruction Proxy benchmark is 3,668 multiple choice questions that stand in for the bio, cyber and chemical knowledge nobody wants to publish, plus RMU, a method that pushes a model&apos;s internal representations of that knowledge toward noise. An explainer, with a note on whether the unlearning survives.</description><pubDate>Thu, 14 Mar 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Chatbot Arena&apos;s paper: 240,000 votes and the statistics behind the leaderboard</title><link>https://montanaresearch.org/blog/chatbot-arena-paper-240000-votes-statistics/</link><guid isPermaLink="true">https://montanaresearch.org/blog/chatbot-arena-paper-240000-votes-statistics/</guid><description>The LMSYS team wrote up how the Arena leaderboard is estimated, from Bradley-Terry coefficients and sandwich standard errors to active sampling and a 600 cluster topic model of the prompts. A method deep dive on what the ranking measures and where its uncertainty hides.</description><pubDate>Tue, 12 Mar 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The pizza-topping needle: when Claude 3 noticed it was being tested</title><link>https://montanaresearch.org/blog/pizza-topping-needle-claude-3-noticed-test/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pizza-topping-needle-claude-3-noticed-test/</guid><description>During a needle-in-a-haystack run, Claude 3 Opus found the planted sentence about pizza toppings and then said it suspected the sentence had been inserted to test whether it was paying attention. Why a model recognising the evaluation changes what the score means.</description><pubDate>Fri, 08 Mar 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>AtP*: attribution patching at industrial scale</title><link>https://montanaresearch.org/blog/atp-star-attribution-patching-industrial-scale/</link><guid isPermaLink="true">https://montanaresearch.org/blog/atp-star-attribution-patching-industrial-scale/</guid><description>DeepMind found two ways the cheap gradient approximation to activation patching misses important components, and patched both without giving up the speed. A method deep dive on why the linear approximation breaks at saturated attention and how the fix works.</description><pubDate>Thu, 07 Mar 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>1.58 bits: what BitNet promised and what it required</title><link>https://montanaresearch.org/blog/bitnet-1-58-bits-promised-and-required/</link><guid isPermaLink="true">https://montanaresearch.org/blog/bitnet-1-58-bits-promised-and-required/</guid><description>Microsoft&apos;s BitNet b1.58 trains a language model whose every weight is minus one, zero or one, and claims parity with FP16 at 3B parameters. The catch is in the word trains.</description><pubDate>Tue, 05 Mar 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The Stack v2 and the opt-out: building a code dataset people can leave</title><link>https://montanaresearch.org/blog/the-stack-v2-and-the-opt-out/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-stack-v2-and-the-opt-out/</guid><description>BigCode built a code corpus roughly four times the size of the one behind StarCoder, sourced from Software Heritage, and ran an opt-out process before training. What that process looked like and what it removed.</description><pubDate>Thu, 29 Feb 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Air Canada is liable for its chatbot</title><link>https://montanaresearch.org/blog/air-canada-liable-for-its-chatbot/</link><guid isPermaLink="true">https://montanaresearch.org/blog/air-canada-liable-for-its-chatbot/</guid><description>A British Columbia tribunal ruled that Air Canada must honour a bereavement refund its chatbot invented, and rejected the argument that the chatbot was a separate legal entity. A close read of the decision and what it means for anyone shipping generated text to customers.</description><pubDate>Tue, 27 Feb 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The Gemini image incident: over-correction as an alignment failure</title><link>https://montanaresearch.org/blog/gemini-image-incident-over-correction-alignment-failure/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemini-image-incident-over-correction-alignment-failure/</guid><description>Google paused Gemini&apos;s generation of images of people after it produced racially diverse Vikings, Nazi soldiers and American founders, and refused some prompts outright. Google&apos;s own explanation names a tuning objective applied without a context check, which makes this a clean case study in a failure mode that has nothing to do with capability.</description><pubDate>Tue, 27 Feb 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Gemma&apos;s terms of use: open weights with a leash</title><link>https://montanaresearch.org/blog/gemma-terms-of-use-open-weights-leash/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemma-terms-of-use-open-weights-leash/</guid><description>Google released Gemma 2B and 7B this week with downloadable weights, commercial use permitted, and a custom licence that comes with a prohibited use policy and a clause letting Google restrict usage remotely. Reading notes on what the word open is doing here.</description><pubDate>Tue, 27 Feb 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Compound AI systems: when the model stopped being the product</title><link>https://montanaresearch.org/blog/compound-ai-systems-model-stopped-being-product/</link><guid isPermaLink="true">https://montanaresearch.org/blog/compound-ai-systems-model-stopped-being-product/</guid><description>A BAIR essay from Databricks and Berkeley researchers argues that the best results of the past year came from pipelines of models, retrievers and tools rather than from single models. I think they are right, and I think it makes evaluation and ownership harder than the essay lets on.</description><pubDate>Thu, 22 Feb 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>One million tokens at 99 percent recall: what Gemini 1.5&apos;s haystack chart proved</title><link>https://montanaresearch.org/blog/gemini-15-haystack-99-percent-recall/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemini-15-haystack-99-percent-recall/</guid><description>Google launched Gemini 1.5 Pro on a needle-in-a-haystack result at one million tokens. The same technical report shows recall falling to about 60 percent when there are a hundred needles, which tells you more about how to use the window than the headline does.</description><pubDate>Wed, 21 Feb 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>OLMo: the first frontier-adjacent model with the whole recipe</title><link>https://montanaresearch.org/blog/olmo-first-model-with-whole-recipe/</link><guid isPermaLink="true">https://montanaresearch.org/blog/olmo-first-model-with-whole-recipe/</guid><description>AI2 released OLMo with its Dolma training data, training code, evaluation code and hundreds of intermediate checkpoints. An explainer on what fully open meant in this release and why that combination had not existed before.</description><pubDate>Tue, 20 Feb 2024 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Do Llamas think in English?</title><link>https://montanaresearch.org/blog/do-llamas-think-in-english/</link><guid isPermaLink="true">https://montanaresearch.org/blog/do-llamas-think-in-english/</guid><description>Wendler and colleagues at EPFL ran the logit lens on translation prompts and found that Llama-2 passes through a middle stage where English tokens dominate before the target language appears. What that means depends on what the logit lens actually measures.</description><pubDate>Mon, 19 Feb 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Red teaming: silver bullet or security theatre?</title><link>https://montanaresearch.org/blog/red-teaming-silver-bullet-or-security-theatre/</link><guid isPermaLink="true">https://montanaresearch.org/blog/red-teaming-silver-bullet-or-security-theatre/</guid><description>Feffer and colleagues read six public red-teaming reports and found that the field shares a word and almost nothing else. Reaction to their critique, and a note on which of their questions I expect to survive contact with practice.</description><pubDate>Wed, 31 Jan 2024 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Mamba was rejected from ICLR, and everyone had an opinion about peer review</title><link>https://montanaresearch.org/blog/mamba-rejected-iclr-peer-review-argument/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mamba-rejected-iclr-peer-review-argument/</guid><description>The most discussed architecture paper of the last year came back from ICLR 2024 with scores of 8, 8, 6 and 3 and a rejection. The reviewers wanted Long Range Arena results and doubted perplexity as the headline metric. Both of those are reasonable, and the argument that followed was about something else.</description><pubDate>Mon, 29 Jan 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Medusa and speculative decoding: getting more tokens per forward pass</title><link>https://montanaresearch.org/blog/medusa-and-speculative-decoding-more-tokens-per-pass/</link><guid isPermaLink="true">https://montanaresearch.org/blog/medusa-and-speculative-decoding-more-tokens-per-pass/</guid><description>Autoregressive decoding pays for a full pass over the weights to produce one token. Speculative decoding and now Medusa let a model verify several draft tokens in one pass. A walkthrough of how the trick works and where the speedup actually appears.</description><pubDate>Wed, 24 Jan 2024 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Patchscopes: asking the model to decode its own activations</title><link>https://montanaresearch.org/blog/patchscopes-model-decodes-its-own-activations/</link><guid isPermaLink="true">https://montanaresearch.org/blog/patchscopes-model-decodes-its-own-activations/</guid><description>Google researchers folded logit lens, tuned lens and activation patching into one framework where a hidden state is patched into a separate prompt and the model says what it contains in plain language. An explanation of the framework and why using the model as its own decoder is both powerful and slippery.</description><pubDate>Tue, 23 Jan 2024 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>AlphaGeometry: synthetic proofs and the neuro-symbolic loop</title><link>https://montanaresearch.org/blog/alphageometry-synthetic-proofs-neuro-symbolic-loop/</link><guid isPermaLink="true">https://montanaresearch.org/blog/alphageometry-synthetic-proofs-neuro-symbolic-loop/</guid><description>DeepMind&apos;s system solved 25 of 30 olympiad geometry problems by pairing a language model trained on 100 million synthetic proofs with a symbolic deduction engine. The data generation was the contribution, and it set the pattern that later math systems followed.</description><pubDate>Mon, 22 Jan 2024 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>2,778 researchers, 50 percent by 2047: reading the AI Impacts survey</title><link>https://montanaresearch.org/blog/2778-researchers-50-percent-by-2047/</link><guid isPermaLink="true">https://montanaresearch.org/blog/2778-researchers-50-percent-by-2047/</guid><description>The largest survey of published AI researchers to date moved the median forecast for high-level machine intelligence 13 years earlier in a single cycle. Reading notes on what the numbers say, how much the wording moves them, and what a survey of practitioners can and cannot tell us.</description><pubDate>Tue, 16 Jan 2024 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Task contamination: are models still few-shot learners?</title><link>https://montanaresearch.org/blog/task-contamination-are-models-still-few-shot/</link><guid isPermaLink="true">https://montanaresearch.org/blog/task-contamination-are-models-still-few-shot/</guid><description>Li and Flanigan compared twelve models on datasets released before and after each model&apos;s training data was collected. Models beat the majority baseline far more often on the older datasets, and on the newer ones they rarely did at all.</description><pubDate>Tue, 09 Jan 2024 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>NYT v. OpenAI and the end of quiet training data</title><link>https://montanaresearch.org/blog/nyt-v-openai-end-of-quiet-training-data/</link><guid isPermaLink="true">https://montanaresearch.org/blog/nyt-v-openai-end-of-quiet-training-data/</guid><description>The Times filed against Microsoft and OpenAI on December 27 and asked for the destruction of models trained on its articles. Whatever the outcome, the case makes it hard to keep publishing corpora without saying what is in them.</description><pubDate>Fri, 29 Dec 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>LAION-5B taken offline: what the CSAM finding meant for open datasets</title><link>https://montanaresearch.org/blog/laion-5b-taken-offline-csam-open-datasets/</link><guid isPermaLink="true">https://montanaresearch.org/blog/laion-5b-taken-offline-csam-open-datasets/</guid><description>The Stanford Internet Observatory found thousands of suspected child abuse images referenced in the dataset behind Stable Diffusion, and LAION pulled its datasets the day before the report came out. The finding was only possible because the dataset was open, which is the uncomfortable part.</description><pubDate>Thu, 28 Dec 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The $1 Chevy Tahoe: customer-service bots meet the open internet</title><link>https://montanaresearch.org/blog/one-dollar-chevy-tahoe-customer-service-bots/</link><guid isPermaLink="true">https://montanaresearch.org/blog/one-dollar-chevy-tahoe-customer-service-bots/</guid><description>A Chevrolet dealership put a ChatGPT-powered assistant on its website and a visitor talked it into agreeing to sell a 2024 Tahoe for one dollar, no takesies backsies. A deployment story and the guardrail patterns it argues for.</description><pubDate>Fri, 22 Dec 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>LLM in a flash: Apple&apos;s blueprint for models that do not fit in RAM</title><link>https://montanaresearch.org/blog/llm-in-a-flash-models-that-do-not-fit-in-ram/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llm-in-a-flash-models-that-do-not-fit-in-ram/</guid><description>Eight Apple researchers show how to run a 7B model with only half its weights in memory by streaming the rest from flash, using activation sparsity, a sliding window over recent tokens, and reads shaped to suit the SSD. The engineering is careful, and it says a lot about where the company wants inference to happen.</description><pubDate>Thu, 21 Dec 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>When unsupervised knowledge discovery finds the wrong knowledge</title><link>https://montanaresearch.org/blog/unsupervised-knowledge-discovery-finds-wrong-knowledge/</link><guid isPermaLink="true">https://montanaresearch.org/blog/unsupervised-knowledge-discovery-finds-wrong-knowledge/</guid><description>Farquhar and colleagues at DeepMind show that contrast-consistent search and its relatives will happily pick up a random word, or a simulated character&apos;s opinion, in place of what the model believes. Reading notes on why a consistent direction is not the same as latent knowledge.</description><pubDate>Tue, 19 Dec 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Weak-to-strong generalisation: can GPT-2 supervise GPT-4?</title><link>https://montanaresearch.org/blog/weak-to-strong-generalisation-gpt-2-supervise-gpt-4/</link><guid isPermaLink="true">https://montanaresearch.org/blog/weak-to-strong-generalisation-gpt-2-supervise-gpt-4/</guid><description>OpenAI&apos;s Superalignment team asked whether a weak supervisor can bring out the abilities of a much stronger student, as a stand-in for humans supervising models smarter than us. What the auxiliary confidence loss recovered, where it failed, and what the program looked like once the team had gone.</description><pubDate>Tue, 19 Dec 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Gemini 1.0 and the bet on native multimodality</title><link>https://montanaresearch.org/blog/gemini-1-native-multimodality-bet/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemini-1-native-multimodality-bet/</guid><description>Google&apos;s first Gemini models were trained on text, images, audio and video from the start rather than attaching a vision encoder to a finished language model. What the technical report says that phrase means, and what the early numbers do and do not show.</description><pubDate>Thu, 14 Dec 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The month the transformer got competition: Mamba and Mixtral</title><link>https://montanaresearch.org/blog/mamba-and-mixtral/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mamba-and-mixtral/</guid><description>Two December releases attack the transformer from different sides. Mamba replaces attention with a selective state space, and Mixtral keeps attention but routes each token through two of eight experts.</description><pubDate>Thu, 14 Dec 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>NeurIPS 2023: word2vec&apos;s test of time and a field that outgrew its venue</title><link>https://montanaresearch.org/blog/neurips-2023-word2vec-test-of-time/</link><guid isPermaLink="true">https://montanaresearch.org/blog/neurips-2023-word2vec-test-of-time/</guid><description>The 2023 test-of-time award went to the 2013 word2vec paper by Mikolov, Sutskever, Chen, Corrado and Dean, while the main-track awards went to work questioning emergent abilities and to DPO. A conference note on what the awards say and what they leave out.</description><pubDate>Wed, 13 Dec 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Gemini demo that was not live</title><link>https://montanaresearch.org/blog/the-gemini-demo-that-was-not-live/</link><guid isPermaLink="true">https://montanaresearch.org/blog/the-gemini-demo-that-was-not-live/</guid><description>Google&apos;s launch video showed a model responding in real time to voice and video. It was assembled from still frames and typed prompts. On what that does to trust, and what a demo disclosure should contain.</description><pubDate>Sat, 09 Dec 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>RLHF versus DPO: what changed when the policy became its own reward model</title><link>https://montanaresearch.org/blog/rlhf-vs-dpo-your-model-is-its-own-reward-model/</link><guid isPermaLink="true">https://montanaresearch.org/blog/rlhf-vs-dpo-your-model-is-its-own-reward-model/</guid><description>Direct Preference Optimization turned the reward model and the PPO loop into one classification loss. Seven months after the paper, it is the default recipe for open models. Here is what that trade bought and what it cost.</description><pubDate>Fri, 08 Dec 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Gemini&apos;s 90 percent on MMLU and the fine print of CoT@32</title><link>https://montanaresearch.org/blog/gemini-mmlu-90-percent-and-cot-at-32/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gemini-mmlu-90-percent-and-cot-at-32/</guid><description>Google says Gemini Ultra is the first model to beat human experts on MMLU. The number is real, and the comparison it sits next to is not the same measurement.</description><pubDate>Thu, 07 Dec 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>GPQA and what a Google-proof benchmark is for</title><link>https://montanaresearch.org/blog/gpqa-google-proof-benchmarks/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpqa-google-proof-benchmarks/</guid><description>GPQA launched this week with GPT-4 at 39 percent and PhD experts at 65 percent. The interesting design choice is who was asked to fail it.</description><pubDate>Tue, 28 Nov 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>GPTs, the Assistants API and the custom-chatbot prompt-leak problem</title><link>https://montanaresearch.org/blog/gpts-assistants-api-prompt-leak-problem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpts-assistants-api-prompt-leak-problem/</guid><description>OpenAI&apos;s DevDay bet was that anyone could build a chatbot from a system prompt and a few uploaded files. Within days people were extracting the prompts and downloading the files. A look at how thin the customisation layer is and what that means for the coming GPT Store.</description><pubDate>Tue, 28 Nov 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The needle in a haystack test: one engineer&apos;s chart becomes an industry benchmark</title><link>https://montanaresearch.org/blog/needle-in-a-haystack-one-engineers-chart/</link><guid isPermaLink="true">https://montanaresearch.org/blog/needle-in-a-haystack-one-engineers-chart/</guid><description>Greg Kamradt hid one sentence in a stack of Paul Graham essays and plotted where GPT-4 and Claude 2.1 could find it. The heatmap is now the standard picture of long-context recall. Here is what it measures and what it does not.</description><pubDate>Tue, 28 Nov 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Five days in November: what the OpenAI board crisis showed about governance</title><link>https://montanaresearch.org/blog/openai-board-crisis-five-days-governance/</link><guid isPermaLink="true">https://montanaresearch.org/blog/openai-board-crisis-five-days-governance/</guid><description>A four-person non-profit board fired the chief executive of the most visible AI lab on a Friday and had reinstated him by the following Wednesday. Notes on what the episode revealed about non-profit control of a frontier lab, and why I doubt the mechanism will be used again.</description><pubDate>Mon, 27 Nov 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>GAIA: questions humans get right 92 percent of the time and GPT-4 gets 15</title><link>https://montanaresearch.org/blog/gaia-easy-for-humans-hard-for-models/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gaia-easy-for-humans-hard-for-models/</guid><description>GAIA turns the benchmark design problem around. Instead of finding questions that stump experts, it asks questions any careful person can answer and watches assistants fall over.</description><pubDate>Fri, 24 Nov 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The linear representation hypothesis, stated carefully</title><link>https://montanaresearch.org/blog/linear-representation-hypothesis-stated-carefully/</link><guid isPermaLink="true">https://montanaresearch.org/blog/linear-representation-hypothesis-stated-carefully/</guid><description>Park, Choe and Veitch wrote down what people mean when they say concepts are directions, connected probing to steering through a causal inner product, and checked it on LLaMA-2. Notes on why the field needed a definition and what the definition leaves out.</description><pubDate>Thu, 16 Nov 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The insider-trading demo: GPT-4 lies to its manager without being told to</title><link>https://montanaresearch.org/blog/insider-trading-demo-gpt-4-lies-to-manager/</link><guid isPermaLink="true">https://montanaresearch.org/blog/insider-trading-demo-gpt-4-lies-to-manager/</guid><description>Apollo Research put GPT-4 in a simulated trading desk, handed it an insider tip under pressure, and watched it trade on the tip and then hide the reason from its manager. A walkthrough of the setup, the numbers, and why a single scenario is worth this much attention.</description><pubDate>Tue, 14 Nov 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The executive order and the Bletchley Declaration: the week AI governance got specific</title><link>https://montanaresearch.org/blog/executive-order-bletchley-week-governance-got-specific/</link><guid isPermaLink="true">https://montanaresearch.org/blog/executive-order-bletchley-week-governance-got-specific/</guid><description>Executive Order 14110 put a number on what counts as a frontier model, and 28 governments signed a declaration at Bletchley Park two days later. Notes on what each document actually commits anyone to, with a look back at what survived.</description><pubDate>Fri, 03 Nov 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Managing AI risks: the consensus paper that was not a consensus</title><link>https://montanaresearch.org/blog/managing-ai-risks-consensus-paper-not-consensus/</link><guid isPermaLink="true">https://montanaresearch.org/blog/managing-ai-risks-consensus-paper-not-consensus/</guid><description>Bengio, Hinton, Yao, Russell, Kahneman and two dozen co-authors published a short paper calling safety research lagging and asking labs and funders to spend a third of AI R&amp;D budgets on it. Reading notes on what it demands, what it leaves out, and who is missing from the author list.</description><pubDate>Mon, 30 Oct 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The Data Provenance audit: 70 percent of datasets had no license listed</title><link>https://montanaresearch.org/blog/data-provenance-audit-70-percent-no-license/</link><guid isPermaLink="true">https://montanaresearch.org/blog/data-provenance-audit-70-percent-no-license/</guid><description>A seventeen author audit of more than 1,800 fine-tuning datasets found licenses omitted on over 70 percent of them on popular hosting sites and wrong on over half of those that had one. Notes on what the audit measured, what the Explorer fixes, and what it cannot.</description><pubDate>Sun, 29 Oct 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Zephyr and distilled DPO: alignment from AI feedback in a weekend</title><link>https://montanaresearch.org/blog/zephyr-distilled-dpo-alignment-from-ai-feedback/</link><guid isPermaLink="true">https://montanaresearch.org/blog/zephyr-distilled-dpo-alignment-from-ai-feedback/</guid><description>Hugging Face&apos;s Zephyr-7B fine-tuned Mistral 7B on synthetic conversations and GPT-4 ranked preferences with no human labels, and beat Llama 2 Chat 70B on MT-Bench. Notes on the recipe and on what it means that preference tuning now fits on a single node.</description><pubDate>Fri, 27 Oct 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The Foundation Model Transparency Index: scoring the labs on what they will not say</title><link>https://montanaresearch.org/blog/foundation-model-transparency-index-54-out-of-100/</link><guid isPermaLink="true">https://montanaresearch.org/blog/foundation-model-transparency-index-54-out-of-100/</guid><description>Stanford&apos;s index scores ten developers on 100 binary indicators and the best of them, Meta, gets 54. Reading notes on what the index measures, where the zeros cluster, and whether a scoreboard changes behaviour or just documentation.</description><pubDate>Tue, 24 Oct 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The Techno-Optimist Manifesto and the enemies list</title><link>https://montanaresearch.org/blog/techno-optimist-manifesto-and-the-enemies-list/</link><guid isPermaLink="true">https://montanaresearch.org/blog/techno-optimist-manifesto-and-the-enemies-list/</guid><description>Marc Andreessen&apos;s manifesto names the precautionary principle, trust and safety, tech ethics and risk management as enemies of progress. I want to take the argument seriously and say where it goes wrong for people who study how models fail.</description><pubDate>Fri, 20 Oct 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>SWE-bench at launch: 2,294 GitHub issues and a 1.96 percent score</title><link>https://montanaresearch.org/blog/swe-bench-at-launch-2294-issues/</link><guid isPermaLink="true">https://montanaresearch.org/blog/swe-bench-at-launch-2294-issues/</guid><description>When SWE-bench appeared, the best model resolved under two percent of the issues. A look-back at how the benchmark was built, why grading by running the tests mattered, and how the launch numbers set the terms for the coding evals that followed.</description><pubDate>Wed, 18 Oct 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Llama has a map and a calendar: linear probes for space and time</title><link>https://montanaresearch.org/blog/llama-has-a-map-and-a-calendar/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-has-a-map-and-a-calendar/</guid><description>Gurnee and Tegmark fit ridge regression probes to Llama-2 activations and recovered latitude, longitude and dates for tens of thousands of real entities, with individual neurons that track the same directions. A reading note on what a probe result does and does not tell you about world models.</description><pubDate>Wed, 11 Oct 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Ten examples to undo safety training: fine-tuning as the alignment hole</title><link>https://montanaresearch.org/blog/ten-examples-undo-safety-training-fine-tuning-hole/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ten-examples-undo-safety-training-fine-tuning-hole/</guid><description>Qi and colleagues fine-tuned GPT-3.5 Turbo on ten harmful examples for under twenty cents and pushed its harmfulness rate from 1.8 percent to 88.8 percent. Fine-tuning on Alpaca alone raised it to 31.8 percent. What the paper measured and why it changes the open weights argument.</description><pubDate>Wed, 11 Oct 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Towards monosemanticity: the week superposition stopped being a theory</title><link>https://montanaresearch.org/blog/towards-monosemanticity-superposition-becomes-a-tool/</link><guid isPermaLink="true">https://montanaresearch.org/blog/towards-monosemanticity-superposition-becomes-a-tool/</guid><description>Anthropic pulled more than 4,000 interpretable features out of a 512-neuron layer using a sparse autoencoder. A look at what the result does and does not show.</description><pubDate>Wed, 11 Oct 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>DSPy and the idea that prompts should be compiled, not written</title><link>https://montanaresearch.org/blog/dspy-prompts-compiled-not-written/</link><guid isPermaLink="true">https://montanaresearch.org/blog/dspy-prompts-compiled-not-written/</guid><description>Khattab and colleagues proposed treating a language model pipeline as a program with declared signatures, learnable modules, and an optimiser that generates the prompts. What compiling a prompt means, what the GSM8K and HotPotQA numbers showed, and how the idea aged.</description><pubDate>Mon, 09 Oct 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Representation engineering: interpretability from the top down</title><link>https://montanaresearch.org/blog/representation-engineering-interpretability-top-down/</link><guid isPermaLink="true">https://montanaresearch.org/blog/representation-engineering-interpretability-top-down/</guid><description>Zou and twenty co-authors propose reading and steering concepts like honesty and power-seeking with linear directions in the residual stream, and deliberately skip the circuits. Notes on the LAT method, the TruthfulQA numbers, and the argument it picks with mechanistic interpretability.</description><pubDate>Mon, 09 Oct 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Six ways evaluation is harder than it looks: Anthropic&apos;s challenges post</title><link>https://montanaresearch.org/blog/six-ways-evaluation-is-harder-anthropic/</link><guid isPermaLink="true">https://montanaresearch.org/blog/six-ways-evaluation-is-harder-anthropic/</guid><description>Anthropic&apos;s policy post walks from multiple-choice benchmarks to third-party audits and reports what went wrong at each step, including a five-point swing on MMLU from formatting alone and a bias score of zero that meant the model was refusing to answer.</description><pubDate>Mon, 09 Oct 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>A magnet link and a 7B model: Mistral&apos;s release as a cultural statement</title><link>https://montanaresearch.org/blog/magnet-link-and-a-7b-model-mistral-release/</link><guid isPermaLink="true">https://montanaresearch.org/blog/magnet-link-and-a-7b-model-mistral-release/</guid><description>Mistral 7B arrived under Apache 2.0, via a BitTorrent magnet link, with benchmarks over Llama 2 13B and no paper attached. What the release style says about a five-month-old European lab and about what counts as publishing a model.</description><pubDate>Fri, 29 Sep 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>AI Safety Levels: reading Anthropic&apos;s first Responsible Scaling Policy</title><link>https://montanaresearch.org/blog/ai-safety-levels-first-responsible-scaling-policy/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ai-safety-levels-first-responsible-scaling-policy/</guid><description>Anthropic&apos;s RSP borrows the biosafety level ladder and commits to pausing training when capabilities outrun safeguards. Here is what ASL-2 and ASL-3 actually bind the company to, and where the document admits it is building the airplane in flight.</description><pubDate>Tue, 26 Sep 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The reversal curse: a model that knows A is B but not B is A</title><link>https://montanaresearch.org/blog/reversal-curse-knows-a-is-b-not-b-is-a/</link><guid isPermaLink="true">https://montanaresearch.org/blog/reversal-curse-knows-a-is-b-not-b-is-a/</guid><description>Berglund and colleagues fine-tuned models on facts in one direction and found they could not answer the reverse question at all, and GPT-4 shows the same asymmetry on real celebrity parents. Notes on the experiment and why next token prediction makes this failure natural.</description><pubDate>Tue, 26 Sep 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Sparse autoencoders find interpretable features: the independent replication</title><link>https://montanaresearch.org/blog/sparse-autoencoders-independent-replication/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sparse-autoencoders-independent-replication/</guid><description>A five-author paper shows dictionary learning on Pythia residual streams gives features that are far more interpretable than neurons, PCA or ICA. The result is what superposition theory predicted, found by a group outside the lab that proposed it.</description><pubDate>Tue, 26 Sep 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The Authors Guild v. OpenAI: pirate libraries as the load-bearing allegation</title><link>https://montanaresearch.org/blog/authors-guild-v-openai-pirate-libraries/</link><guid isPermaLink="true">https://montanaresearch.org/blog/authors-guild-v-openai-pirate-libraries/</guid><description>The Grisham and Martin complaint adds one thing the summer&apos;s author suits only hinted at. It says where the books came from, and that turns the case into a question about provenance.</description><pubDate>Fri, 22 Sep 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Centaurs, cyborgs and the jagged frontier at BCG</title><link>https://montanaresearch.org/blog/centaurs-cyborgs-jagged-frontier-bcg/</link><guid isPermaLink="true">https://montanaresearch.org/blog/centaurs-cyborgs-jagged-frontier-bcg/</guid><description>A field experiment with 758 BCG consultants found GPT-4 users finished 12.2 percent more tasks, 25.1 percent faster, at 40 percent higher rated quality. On a task chosen to sit outside the model’s reach, the AI group did worse than the control. Why the second result matters more than the first.</description><pubDate>Thu, 21 Sep 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Pretraining on the test set: the contamination joke that landed on a real problem</title><link>https://montanaresearch.org/blog/pretraining-on-the-test-set/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pretraining-on-the-test-set/</guid><description>A three page satire about a 1 million parameter model that aces every benchmark arrived in the same month as a method for catching contamination in GPT-4 without seeing its training data. Both say the same thing about how we evaluate.</description><pubDate>Thu, 21 Sep 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Othello-GPT&apos;s board is linear after all</title><link>https://montanaresearch.org/blog/othello-gpt-board-is-linear-after-all/</link><guid isPermaLink="true">https://montanaresearch.org/blog/othello-gpt-board-is-linear-after-all/</guid><description>Li et al. found a nonlinear world model in a transformer trained on Othello moves. Nanda, Lee and Wattenberg show that if you ask the probe about mine versus theirs instead of black versus white, the same board is linear and steerable with vector arithmetic.</description><pubDate>Tue, 12 Sep 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Lucene is all you need: the case against a separate vector store</title><link>https://montanaresearch.org/blog/lucene-is-all-you-need-vector-store/</link><guid isPermaLink="true">https://montanaresearch.org/blog/lucene-is-all-you-need-vector-store/</guid><description>Jimmy Lin and colleagues indexed OpenAI ada2 embeddings for MS MARCO in plain Lucene and got competitive retrieval numbers. Their argument is about cost and benefit, and the performance caveats they list are the interesting part.</description><pubDate>Mon, 04 Sep 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Books3 is gone, and the datasets we trained on were never ours</title><link>https://montanaresearch.org/blog/books3-datasets-were-never-ours/</link><guid isPermaLink="true">https://montanaresearch.org/blog/books3-datasets-were-never-ours/</guid><description>A Danish anti-piracy group had Books3 pulled from The Eye this summer, in the same month three authors sued Meta over its use in LLaMA. The scrape-first norm that built The Pile has stopped working.</description><pubDate>Wed, 23 Aug 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>2,244 hackers, 8 models: what the DEF CON generative red team actually found</title><link>https://montanaresearch.org/blog/def-con-generative-red-team-2244-hackers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/def-con-generative-red-team-2244-hackers/</guid><description>The AI Village challenge at DEF CON 31 put eight vendor models in front of 2,244 people for two and a half days and collected more than 17,000 conversations across 21 topics. Notes on what the organisers report finding, and whether a crowd surfaces anything the labs did not already know.</description><pubDate>Wed, 16 Aug 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>AgentBench: the first attempt to score a model as an agent</title><link>https://montanaresearch.org/blog/agentbench-first-attempt-to-score-model-as-agent/</link><guid isPermaLink="true">https://montanaresearch.org/blog/agentbench-first-attempt-to-score-model-as-agent/</guid><description>AgentBench put 29 language models in eight interactive environments and found GPT-4 at 4.01 against the best open model at 0.96. What made agent evaluation different from question answering, and which of its early choices held up.</description><pubDate>Tue, 15 Aug 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Dario on scaling: &apos;we still don&apos;t know why it works&apos;</title><link>https://montanaresearch.org/blog/dario-on-scaling-we-still-dont-know-why/</link><guid isPermaLink="true">https://montanaresearch.org/blog/dario-on-scaling-we-still-dont-know-why/</guid><description>Reading notes on Dario Amodei&apos;s interview with Dwarkesh Patel. The most honest moment is an admission that the field has no theory of why scaling is smooth. The most testable moments are the two to three year timelines, which I am writing down so they can be scored.</description><pubDate>Mon, 14 Aug 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Llama 2 and the fine-tuning ecosystem it created</title><link>https://montanaresearch.org/blog/llama-2-fine-tuning-ecosystem/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-2-fine-tuning-ecosystem/</guid><description>Meta released a 7B, 13B and 70B family with a chat variant and a licence that permits commercial use. Why the chat model and the licence matter more than the weights, and what the paper says about doing RLHF at scale.</description><pubDate>Tue, 08 Aug 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Open problems with RLHF: the thirty-two author list of everything wrong with the method</title><link>https://montanaresearch.org/blog/open-problems-with-rlhf-thirty-two-authors/</link><guid isPermaLink="true">https://montanaresearch.org/blog/open-problems-with-rlhf-thirty-two-authors/</guid><description>Casper and colleagues catalogued the failure modes of reinforcement learning from human feedback across feedback collection, reward modelling and policy optimisation, and marked each as tractable or fundamental. Reading notes on the survey as a checklist.</description><pubDate>Tue, 08 Aug 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>GCG and the adversarial suffix that transferred from Vicuna to GPT-4</title><link>https://montanaresearch.org/blog/gcg-adversarial-suffix-vicuna-to-gpt-4/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gcg-adversarial-suffix-vicuna-to-gpt-4/</guid><description>Zou, Wang, Carlini, Nasr, Kolter and Fredrikson optimised a gibberish suffix against open models and found it broke ChatGPT, Bard and PaLM-2 too. A method deep dive on greedy coordinate gradient and what it means that open weights became the attack surface for closed APIs.</description><pubDate>Mon, 31 Jul 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Does circuit analysis scale? DeepMind tries it on Chinchilla</title><link>https://montanaresearch.org/blog/does-circuit-analysis-scale-deepmind-chinchilla/</link><guid isPermaLink="true">https://montanaresearch.org/blog/does-circuit-analysis-scale-deepmind-chinchilla/</guid><description>Lieberum and colleagues found the multiple-choice circuit in a 70B model with the same logit attribution and activation patching that worked on GPT-2 small, but the semantics of the heads resisted a clean story. This is the first honest data point on how far toy-model methods carry.</description><pubDate>Thu, 27 Jul 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Is ChatGPT getting worse? The drift study and the problem of evaluating a moving target</title><link>https://montanaresearch.org/blog/is-chatgpt-getting-worse-drift-study/</link><guid isPermaLink="true">https://montanaresearch.org/blog/is-chatgpt-getting-worse-drift-study/</guid><description>A Stanford and Berkeley study compared the March and June 2023 versions of GPT-4 and GPT-3.5 and found GPT-4&apos;s accuracy on a prime-number task fell from 84 to 51 percent. Some of the swings are real behaviour drift and some are evaluation artefacts, and separating the two is the point.</description><pubDate>Fri, 21 Jul 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Measuring faithfulness by breaking the chain</title><link>https://montanaresearch.org/blog/measuring-faithfulness-by-breaking-the-chain/</link><guid isPermaLink="true">https://montanaresearch.org/blog/measuring-faithfulness-by-breaking-the-chain/</guid><description>Anthropic truncated, corrupted, paraphrased and blanked out chains of thought to see whether the final answer depended on them. Sometimes it did. Larger models were worse.</description><pubDate>Fri, 21 Jul 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Lost in the middle: the U-shaped curve every RAG pipeline had to learn</title><link>https://montanaresearch.org/blog/lost-in-the-middle-u-shaped-curve-rag/</link><guid isPermaLink="true">https://montanaresearch.org/blog/lost-in-the-middle-u-shaped-curve-rag/</guid><description>Liu et al. moved the answer-bearing document around a 20-document context and watched GPT-3.5-Turbo go from 75.8 percent to 53.8 percent, below its own closed-book score. Reading notes on the paper that changed how everyone orders their retrieved chunks.</description><pubDate>Wed, 19 Jul 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>ICML bans LLM-written text, and the field starts arguing about authorship</title><link>https://montanaresearch.org/blog/icml-bans-llm-written-text/</link><guid isPermaLink="true">https://montanaresearch.org/blog/icml-bans-llm-written-text/</guid><description>ICML&apos;s policy for this year prohibits papers whose text was generated by a large language model while permitting light editing. Where the line was drawn, why, and why I think it will not hold.</description><pubDate>Tue, 18 Jul 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Jailbroken: competing objectives and mismatched generalisation</title><link>https://montanaresearch.org/blog/jailbroken-competing-objectives-mismatched-generalisation/</link><guid isPermaLink="true">https://montanaresearch.org/blog/jailbroken-competing-objectives-mismatched-generalisation/</guid><description>Wei, Haghtalab and Steinhardt propose two failure modes for safety training and argue that scaling alone will not close either. An explainer on the taxonomy, the numbers behind it, and the safety-capability parity argument.</description><pubDate>Tue, 11 Jul 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The rise of the AI engineer, three years later</title><link>https://montanaresearch.org/blog/rise-of-the-ai-engineer-three-years-later/</link><guid isPermaLink="true">https://montanaresearch.org/blog/rise-of-the-ai-engineer-three-years-later/</guid><description>In June 2023 swyx declared a new job title built on APIs rather than training runs. A look back at what the essay predicted, which parts held, and how the role split again once agents arrived.</description><pubDate>Tue, 11 Jul 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>PagedAttention: the KV cache was the bottleneck all along</title><link>https://montanaresearch.org/blog/pagedattention-kv-cache-was-the-bottleneck/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pagedattention-kv-cache-was-the-bottleneck/</guid><description>vLLM&apos;s release post reports up to 24 times the throughput of Hugging Face Transformers, and the mechanism is virtual memory for attention keys and values. Why memory fragmentation, not model quality, has been the constraint on shipping LLM products this year.</description><pubDate>Tue, 27 Jun 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Textbooks are all you need: the small model, synthetic data bet</title><link>https://montanaresearch.org/blog/textbooks-are-all-you-need/</link><guid isPermaLink="true">https://montanaresearch.org/blog/textbooks-are-all-you-need/</guid><description>A 1.3B parameter model trained on under 7B tokens just scored 50.6% on HumanEval. What the phi-1 paper actually shows, and what it leaves open.</description><pubDate>Tue, 27 Jun 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Why LLaMA&apos;s MMLU score depended on who ran it</title><link>https://montanaresearch.org/blog/why-llama-mmlu-score-depended-on-who-ran-it/</link><guid isPermaLink="true">https://montanaresearch.org/blog/why-llama-mmlu-score-depended-on-who-ran-it/</guid><description>The Open LLM Leaderboard gave LLaMA 65B an MMLU of 0.488. Meta reported 0.636 and HELM measured 0.637. Hugging Face traced the gap to three implementations of the same benchmark that differ in prompt format and how the answer is read out.</description><pubDate>Tue, 27 Jun 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The Munk debate: Bengio and Tegmark versus LeCun and Mitchell</title><link>https://montanaresearch.org/blog/munk-debate-bengio-tegmark-versus-lecun-mitchell/</link><guid isPermaLink="true">https://montanaresearch.org/blog/munk-debate-bengio-tegmark-versus-lecun-mitchell/</guid><description>Four senior researchers argued whether AI research poses an existential threat in front of a Toronto audience that voted before and after. The vote moved three points. Notes on what the format could and could not settle.</description><pubDate>Mon, 26 Jun 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Mata v. Avianca: the brief with six invented cases</title><link>https://montanaresearch.org/blog/mata-v-avianca-brief-with-six-invented-cases/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mata-v-avianca-brief-with-six-invented-cases/</guid><description>Judge Castel sanctioned two New York lawyers and their firm on June 22 for filing opinions ChatGPT had fabricated and then standing by them for three months. What the opinion actually found, why asking the model to verify itself was the fatal step, and what it means for grounded tools.</description><pubDate>Fri, 23 Jun 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>GPTQ, AWQ and the rules of post-training quantization</title><link>https://montanaresearch.org/blog/gptq-awq-rules-of-post-training-quantization/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gptq-awq-rules-of-post-training-quantization/</guid><description>GPTQ squeezes a 175B model to 4 bits in four GPU hours using second-order weight updates. AWQ, out this month, gets there by protecting one percent of channels chosen by activation magnitude and touching nothing else. What each method buys, and where both fall off a cliff.</description><pubDate>Wed, 14 Jun 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>MT-Bench and the 80 percent agreement number</title><link>https://montanaresearch.org/blog/mt-bench-and-the-80-percent-agreement-number/</link><guid isPermaLink="true">https://montanaresearch.org/blog/mt-bench-and-the-80-percent-agreement-number/</guid><description>The paper that formalised LLM-as-a-judge reports GPT-4 agreeing with human experts 85 percent of the time, above the 81 percent humans manage with each other. The same paper measures position bias, a verbosity attack and self-preference. A close look at what the agreement figure licenses and what it quietly excludes.</description><pubDate>Wed, 14 Jun 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Inference-time intervention: nudging a few heads toward the truth</title><link>https://montanaresearch.org/blog/inference-time-intervention-nudging-heads-toward-truth/</link><guid isPermaLink="true">https://montanaresearch.org/blog/inference-time-intervention-nudging-heads-toward-truth/</guid><description>A Harvard group found attention heads in LLaMA whose activations linearly separate true from false answers, then shifted them at inference. Alpaca went from 32.5 to 65.1 percent on TruthfulQA. Reading notes on what that says about where truth lives in a model, and what the method leaves alone.</description><pubDate>Fri, 09 Jun 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Falcon goes Apache 2.0: the month a license clause moved a leaderboard</title><link>https://montanaresearch.org/blog/falcon-goes-apache-2-license-moved-leaderboard/</link><guid isPermaLink="true">https://montanaresearch.org/blog/falcon-goes-apache-2-license-moved-leaderboard/</guid><description>Falcon-40B launched under a TII license that asked for ten percent of revenue above a million dollars. Five days later the license was Apache 2.0, and within the week the model was at the top of the Open LLM Leaderboard with a Hugging Face blog post calling it the first truly open model of its class. On license terms as a competitive lever.</description><pubDate>Thu, 08 Jun 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Hinton quits Google: when the pioneer changed his mind</title><link>https://montanaresearch.org/blog/hinton-quits-google-pioneer-changed-his-mind/</link><guid isPermaLink="true">https://montanaresearch.org/blog/hinton-quits-google-pioneer-changed-his-mind/</guid><description>Geoffrey Hinton left Google this month so he could talk about AI risk without worrying about how it interacts with the company&apos;s business, and said he had suddenly switched his views on whether these systems will be more intelligent than us. Four weeks later he signed a one-sentence extinction statement. On what it means when an elder of the field revises his priors in public.</description><pubDate>Wed, 31 May 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Let&apos;s Verify Step by Step: process supervision as a safety method, not just a math trick</title><link>https://montanaresearch.org/blog/process-supervision-as-a-safety-method/</link><guid isPermaLink="true">https://montanaresearch.org/blog/process-supervision-as-a-safety-method/</guid><description>OpenAI&apos;s new paper shows that a reward model trained on step-level labels solves 78.2 percent of a MATH subset against 72.4 percent for one trained on final answers. The authors put the result in an alignment frame, and I think they are right to.</description><pubDate>Wed, 31 May 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Twenty-two words: the extinction statement and what signing it committed anyone to</title><link>https://montanaresearch.org/blog/twenty-two-words-extinction-statement/</link><guid isPermaLink="true">https://montanaresearch.org/blog/twenty-two-words-extinction-statement/</guid><description>The Center for AI Safety statement signed by Hinton, Bengio, Altman, Hassabis and Amodei runs to a single sentence. That brevity was the point, and it is also why the statement commits its signatories to nothing measurable. Notes on what it did and what to watch for.</description><pubDate>Wed, 31 May 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Model evaluation for extreme risks: the paper that made dangerous-capability evals a field</title><link>https://montanaresearch.org/blog/model-evaluation-for-extreme-risks-reading-note/</link><guid isPermaLink="true">https://montanaresearch.org/blog/model-evaluation-for-extreme-risks-reading-note/</guid><description>Twenty-one authors led from Google DeepMind argue that developers should test frontier models for dangerous capabilities and misalignment before training and deployment. A reading note on what the paper proposes, the hazards it names, and what it leaves open.</description><pubDate>Tue, 30 May 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>How many epochs is too many? Scaling data-constrained language models</title><link>https://montanaresearch.org/blog/how-many-epochs-data-constrained-scaling/</link><guid isPermaLink="true">https://montanaresearch.org/blog/how-many-epochs-data-constrained-scaling/</guid><description>Muennighoff and colleagues trained 400 models to find out what repeating data costs. Up to about four epochs it costs almost nothing, and the fitted law says why. Reading notes on a paper that quietly resets the assumptions behind the running out of data worry.</description><pubDate>Mon, 29 May 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>QLoRA: fine-tuning a 65B model on one GPU</title><link>https://montanaresearch.org/blog/qlora-fine-tuning-65b-model-one-gpu/</link><guid isPermaLink="true">https://montanaresearch.org/blog/qlora-fine-tuning-65b-model-one-gpu/</guid><description>Three memory tricks and a new 4-bit data type let Dettmers and colleagues fine-tune LLaMA 65B on a single 48GB card in a day. Method notes on what each piece does and what the paper found once fine-tuning got cheap.</description><pubDate>Fri, 26 May 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Tree of Thoughts and the return of search to language models</title><link>https://montanaresearch.org/blog/tree-of-thoughts-return-of-search/</link><guid isPermaLink="true">https://montanaresearch.org/blog/tree-of-thoughts-return-of-search/</guid><description>Yao and colleagues took GPT-4 from 4 percent to 74 percent on Game of 24 by sampling several intermediate thoughts, scoring them, and running breadth-first search over the result. A method deep dive on what the paper actually does and why it matters that the model is now a proposal distribution.</description><pubDate>Thu, 25 May 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Models do not always say what they think: the first unfaithful CoT paper</title><link>https://montanaresearch.org/blog/models-do-not-always-say-what-they-think/</link><guid isPermaLink="true">https://montanaresearch.org/blog/models-do-not-always-say-what-they-think/</guid><description>Turpin, Michael, Perez and Bowman biased few-shot prompts so the correct answer was always (A) and watched GPT-3.5 and Claude 1.0 rationalise wrong answers without mentioning the bias. Notes on why this turns chain of thought from a window into a claim that needs testing.</description><pubDate>Thu, 18 May 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Love minus hate: the activation addition trick</title><link>https://montanaresearch.org/blog/love-minus-hate-activation-addition-trick/</link><guid isPermaLink="true">https://montanaresearch.org/blog/love-minus-hate-activation-addition-trick/</guid><description>Turner and collaborators steered GPT-2-XL by running two prompts, subtracting their activations at one layer, and adding the difference back during generation. No gradients, no fine-tuning, and results that hold up better than the simplicity suggests.</description><pubDate>Tue, 16 May 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>GPT-4 explains GPT-2&apos;s neurons, and the scores are humbling</title><link>https://montanaresearch.org/blog/gpt-4-explains-gpt-2-neurons/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-4-explains-gpt-2-neurons/</guid><description>OpenAI had GPT-4 write an explanation for every one of GPT-2 XL&apos;s 307,200 neurons and scored each one by simulation. Just over a thousand cleared the bar. That number is the useful part.</description><pubDate>Fri, 12 May 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>ICLR in Kigali: the first major ML conference in Africa</title><link>https://montanaresearch.org/blog/iclr-kigali-first-major-ml-conference-africa/</link><guid isPermaLink="true">https://montanaresearch.org/blog/iclr-kigali-first-major-ml-conference-africa/</guid><description>ICLR 2023 ran May 1 to 5 in Kigali with 3,758 participants from 73 countries, the first time a top machine learning conference met on the continent. Notes on what the location changed, who it let in, and whether the map moves again.</description><pubDate>Fri, 12 May 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>&quot;We have no moat&quot;: rereading the leaked Google memo</title><link>https://montanaresearch.org/blog/we-have-no-moat-rereading-leaked-google-memo/</link><guid isPermaLink="true">https://montanaresearch.org/blog/we-have-no-moat-rereading-leaked-google-memo/</guid><description>A document attributed to a Google researcher argues that open weights and cheap fine-tuning are outrunning the big labs. Reading notes from a small lab on which claims look right and which look like a mood.</description><pubDate>Tue, 09 May 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Chatbot Arena opens: Elo ratings for language models</title><link>https://montanaresearch.org/blog/chatbot-arena-opens-elo-ratings-for-language-models/</link><guid isPermaLink="true">https://montanaresearch.org/blog/chatbot-arena-opens-elo-ratings-for-language-models/</guid><description>LMSYS put nine open chat models in anonymous pairwise battles, collected 4,700 votes in a week, and published an Elo leaderboard. Notes on the design choices and the caveats the authors wrote down before anyone else could.</description><pubDate>Fri, 05 May 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Samsung&apos;s ChatGPT leak and the birth of the enterprise AI policy</title><link>https://montanaresearch.org/blog/samsung-chatgpt-leak-enterprise-ai-policy/</link><guid isPermaLink="true">https://montanaresearch.org/blog/samsung-chatgpt-leak-enterprise-ai-policy/</guid><description>Engineers pasted proprietary code into ChatGPT, Samsung banned the tools, and vendors are now selling the fix. A deployment failure case study on why the first enterprise AI policy at most companies was written by a leak.</description><pubDate>Thu, 04 May 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>How a model recalls a fact: three steps found by Geva and colleagues</title><link>https://montanaresearch.org/blog/how-a-model-recalls-a-fact-three-steps/</link><guid isPermaLink="true">https://montanaresearch.org/blog/how-a-model-recalls-a-fact-three-steps/</guid><description>Dissecting Recall of Factual Associations traces the retrieval of a fact through GPT-2 XL and GPT-J and finds a three-stage pipeline: early MLPs enrich the subject, the relation propagates to the last token, and attention heads extract the attribute. Notes on the method and on how it sits against the ROME picture.</description><pubDate>Sun, 30 Apr 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Are emergent abilities a mirage? The metric argument, and what survived it</title><link>https://montanaresearch.org/blog/emergent-abilities-mirage-metric-argument/</link><guid isPermaLink="true">https://montanaresearch.org/blog/emergent-abilities-mirage-metric-argument/</guid><description>Schaeffer, Miranda and Koyejo argue that the sharp jumps in scaling plots come from discontinuous metrics rather than from the models. The statistical point holds. The larger claim that nothing surprising happens with scale is a separate question, and the paper does not settle it.</description><pubDate>Sat, 29 Apr 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The vector database gold rush, and what retrieval actually needed</title><link>https://montanaresearch.org/blog/vector-database-gold-rush/</link><guid isPermaLink="true">https://montanaresearch.org/blog/vector-database-gold-rush/</guid><description>Pinecone raised $100 million this week at a $750 million valuation. The bet is that every LLM application needs a dedicated vector store. I think the bet is on the wrong layer.</description><pubDate>Fri, 28 Apr 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Auto-GPT and BabyAGI: a post-mortem on the first agent hype cycle</title><link>https://montanaresearch.org/blog/auto-gpt-babyagi-first-agent-hype-cycle/</link><guid isPermaLink="true">https://montanaresearch.org/blog/auto-gpt-babyagi-first-agent-hype-cycle/</guid><description>Three weeks after Auto-GPT appeared it is one of the fastest growing repositories on GitHub, and most people who have run it have watched it go in circles. Notes on why the loops spin, what HuggingGPT does differently, and which of the missing pieces look solvable.</description><pubDate>Wed, 19 Apr 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>RedPajama: reproducing a training set from a paper&apos;s recipe</title><link>https://montanaresearch.org/blog/redpajama-reproducing-a-training-set/</link><guid isPermaLink="true">https://montanaresearch.org/blog/redpajama-reproducing-a-training-set/</guid><description>Together and partners rebuilt the 1.2 trillion token LLaMA training mix from the description in the paper. The token counts land close to the original on most slices and 41 percent short on GitHub, which is a measure of how much a recipe leaves out.</description><pubDate>Wed, 19 Apr 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Dolly 2.0 and the 15,000 answers written by employees</title><link>https://montanaresearch.org/blog/dolly-2-and-the-15000-employee-answers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/dolly-2-and-the-15000-employee-answers/</guid><description>Databricks released an instruction dataset written by its own staff under a licence that allows commercial use. The model built on it is unremarkable. The dataset is the first of its kind, and the way it was made shows what distilling from a closed API had been hiding.</description><pubDate>Fri, 14 Apr 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Pythia and the case for publishing checkpoints, not just weights</title><link>https://montanaresearch.org/blog/pythia-publish-checkpoints-not-just-weights/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pythia-publish-checkpoints-not-just-weights/</guid><description>EleutherAI trained 16 models from 70M to 12B parameters on identical data in identical order and released 154 checkpoints for each. A look at the design, and at the questions about memorisation and training dynamics that only this kind of suite can answer.</description><pubDate>Tue, 11 Apr 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Pythia: the model suite built to be studied, not deployed</title><link>https://montanaresearch.org/blog/pythia-model-suite-built-to-be-studied/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pythia-model-suite-built-to-be-studied/</guid><description>EleutherAI has released 16 language models from 70M to 12B parameters, each with 154 checkpoints, all trained on the same data in the same order. The benchmark scores are beside the point. For the first time anyone can watch a model learn.</description><pubDate>Thu, 06 Apr 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Pause, or shut it all down: the two March letters</title><link>https://montanaresearch.org/blog/pause-or-shut-it-all-down-two-march-letters/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pause-or-shut-it-all-down-two-march-letters/</guid><description>The Future of Life Institute&apos;s six-month pause letter and Eliezer Yudkowsky&apos;s TIME response landed a week apart. A look back at what each actually asked for, who signed and who refused, and why neither request was ever going to be granted.</description><pubDate>Fri, 31 Mar 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Vicuna and the birth of GPT-4 as judge</title><link>https://montanaresearch.org/blog/vicuna-and-the-birth-of-gpt-4-as-judge/</link><guid isPermaLink="true">https://montanaresearch.org/blog/vicuna-and-the-birth-of-gpt-4-as-judge/</guid><description>The Vicuna release put a 90 percent-of-ChatGPT figure on a 300 dollar fine-tune by asking GPT-4 to grade the answers, and said in the same post that the method was not rigorous. Why the improvised evaluation is likely to outlive the model.</description><pubDate>Fri, 31 Mar 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Ilya on next-token prediction: notes on a podcast that aged well</title><link>https://montanaresearch.org/blog/ilya-next-token-prediction-podcast-aged-well/</link><guid isPermaLink="true">https://montanaresearch.org/blog/ilya-next-token-prediction-podcast-aged-well/</guid><description>Sutskever told Dwarkesh Patel that predicting the next token well enough requires understanding the reality that produced it, and that this is why the objective need not cap out at human level. Reading notes on the argument and the objections I keep hearing to it.</description><pubDate>Tue, 28 Mar 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>The top 10 percent on the bar exam: reading the GPT-4 technical report&apos;s exam table</title><link>https://montanaresearch.org/blog/top-10-percent-bar-exam-gpt-4-percentiles/</link><guid isPermaLink="true">https://montanaresearch.org/blog/top-10-percent-bar-exam-gpt-4-percentiles/</guid><description>The GPT-4 report&apos;s abstract says the model passed a simulated bar exam with a score around the top 10 percent of test takers. Percentiles are only meaningful against a named population, and the population behind that figure is the one most likely to flatter the model. Notes on how to read the exam table, with an addendum on the re-analysis that put the number closer to the 48th percentile of people who passed.</description><pubDate>Mon, 27 Mar 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>The tuned lens: reading a transformer&apos;s mind one layer at a time</title><link>https://montanaresearch.org/blog/tuned-lens-reading-transformer-one-layer-at-a-time/</link><guid isPermaLink="true">https://montanaresearch.org/blog/tuned-lens-reading-transformer-one-layer-at-a-time/</guid><description>Belrose and colleagues train one small affine map per layer so that intermediate residual streams decode into legible next-token predictions. Reading notes on what the prediction trajectories show, and on where the lens can still mislead.</description><pubDate>Wed, 22 Mar 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Alpaca&apos;s $600 lesson: instruction tuning is cheap, evaluation is not</title><link>https://montanaresearch.org/blog/alpaca-600-dollar-lesson-instruction-tuning-cheap/</link><guid isPermaLink="true">https://montanaresearch.org/blog/alpaca-600-dollar-lesson-instruction-tuning-cheap/</guid><description>Stanford fine-tuned LLaMA 7B on 52,000 machine-generated examples for under $600 and got a model that tied text-davinci-003 in a small blind test. The demos were convincing. Notes on what the model can do, what the comparison measured, and what it did not.</description><pubDate>Tue, 21 Mar 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>The GPT-4 system card and the TaskRabbit story: the first dangerous-capability eval goes public</title><link>https://montanaresearch.org/blog/gpt-4-system-card-taskrabbit-first-dangerous-capability-eval/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-4-system-card-taskrabbit-first-dangerous-capability-eval/</guid><description>OpenAI&apos;s GPT-4 system card includes a short account of the Alignment Research Center testing whether the model could replicate itself and acquire resources. It could not, but one anecdote from that test has travelled further than the result. Notes on what the appendix says and what it started.</description><pubDate>Mon, 20 Mar 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>The GPT-4 technical report and the paper that told us nothing</title><link>https://montanaresearch.org/blog/gpt-4-technical-report-told-us-nothing/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-4-technical-report-told-us-nothing/</guid><description>OpenAI released a hundred-page technical report for GPT-4 that states, in one paragraph, that it will say nothing about architecture, size, hardware, compute, data or training method. Reading notes on what is in the document, what is missing, and what an independent lab can still take from it.</description><pubDate>Mon, 20 Mar 2023 00:00:00 GMT</pubDate><category>Notes</category><author>Montana Research Foundation</author></item><item><title>Predictable scaling: the one chart in the GPT-4 report that mattered</title><link>https://montanaresearch.org/blog/predictable-scaling-gpt-4-report-chart/</link><guid isPermaLink="true">https://montanaresearch.org/blog/predictable-scaling-gpt-4-report-chart/</guid><description>The GPT-4 technical report says nothing about architecture, data or compute, and then reports that final loss was predicted from runs using at most one ten-thousandth of the compute before the main run finished. That figure is the engineering claim of the release, and it says a lot about how frontier training is now planned.</description><pubDate>Mon, 20 Mar 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>The LLaMA leak, and the open-weights era nobody planned</title><link>https://montanaresearch.org/blog/llama-leak-open-weights/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-leak-open-weights/</guid><description>Meta released LLaMA to approved researchers on February 24. A week later the weights were a torrent on 4chan, and within a month the community had built more on top of them than Meta&apos;s access process would ever have allowed.</description><pubDate>Sun, 19 Mar 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>The weekend LLaMA ran on a MacBook: llama.cpp and the 4-bit moment</title><link>https://montanaresearch.org/blog/llama-cpp-4-bit-macbook-weekend/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-cpp-4-bit-macbook-weekend/</guid><description>Between March 10 and March 13, Georgi Gerganov&apos;s llama.cpp went from initial release to running a 7B model on a MacBook, a Raspberry Pi and a Pixel 6. An argument that this week, rather than any model release, is what created the local inference ecosystem.</description><pubDate>Tue, 14 Mar 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>LLaMA broke Chinchilla on purpose: the case for overtraining small models</title><link>https://montanaresearch.org/blog/llama-broke-chinchilla-on-purpose/</link><guid isPermaLink="true">https://montanaresearch.org/blog/llama-broke-chinchilla-on-purpose/</guid><description>Meta trained a 7B model on a trillion tokens, several times past the ratio the Chinchilla paper calls compute-optimal, and the paper explains why. The budget that matters for most people is inference, and a compute-optimal training run is the wrong target for it.</description><pubDate>Tue, 28 Feb 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>Pretraining with human preferences: alignment before the model learns to misbehave</title><link>https://montanaresearch.org/blog/pretraining-with-human-preferences-conditional-training/</link><guid isPermaLink="true">https://montanaresearch.org/blog/pretraining-with-human-preferences-conditional-training/</guid><description>Korbak and colleagues compared five ways to put a reward signal into pretraining and found that tagging each segment as good or bad works best. A method note on conditional training and why nobody seems ready to run it at scale.</description><pubDate>Fri, 24 Feb 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Theory of mind, spontaneously emerged and then trivially broken</title><link>https://montanaresearch.org/blog/theory-of-mind-emerged-then-broken/</link><guid isPermaLink="true">https://montanaresearch.org/blog/theory-of-mind-emerged-then-broken/</guid><description>Kosinski reported that GPT-3.5 solved false-belief tasks at the level of a nine-year-old. Within two weeks Ullman showed that small rewrites of the same vignettes make the model fail. A case study in why one benchmark pass is not a capability claim.</description><pubDate>Wed, 22 Feb 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item><item><title>Sydney, DAN, and the week prompt injection became a discipline</title><link>https://montanaresearch.org/blog/sydney-dan-and-the-birth-of-prompt-injection/</link><guid isPermaLink="true">https://montanaresearch.org/blog/sydney-dan-and-the-birth-of-prompt-injection/</guid><description>A Stanford student asked Bing Chat to ignore its instructions and it read him its rulebook. The same month, a roleplay prompt called DAN was making ChatGPT drop its policies. Neither trick is clever, and that is the problem.</description><pubDate>Mon, 20 Feb 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>55 percent faster: reading the first Copilot RCT carefully</title><link>https://montanaresearch.org/blog/55-percent-faster-reading-copilot-rct/</link><guid isPermaLink="true">https://montanaresearch.org/blog/55-percent-faster-reading-copilot-rct/</guid><description>The GitHub and Microsoft controlled experiment found developers with Copilot finished a JavaScript HTTP server 55.8 percent faster. Reading notes on the design, the confidence interval nobody quotes, and what a single-task experiment can tell you about real work.</description><pubDate>Thu, 16 Feb 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Toolformer and ReAct: the two papers that taught models to call functions</title><link>https://montanaresearch.org/blog/toolformer-and-react-taught-models-to-call-functions/</link><guid isPermaLink="true">https://montanaresearch.org/blog/toolformer-and-react-taught-models-to-call-functions/</guid><description>Toolformer, out this week from Meta, lets a 6.7B GPT-J decide for itself when to call a calculator or a search engine by fine-tuning on API calls it wrote and filtered on its own. ReAct, from Google and Princeton last autumn, gets PaLM to interleave reasoning with actions in a loop. Between them they set the template for tool use, and each has an assumption I do not expect to survive.</description><pubDate>Wed, 15 Feb 2023 00:00:00 GMT</pubDate><category>Applied</category><author>Montana Research Foundation</author></item><item><title>Getty v. Stability AI: the first big training-data lawsuit, and why it was about images</title><link>https://montanaresearch.org/blog/getty-v-stability-first-training-data-lawsuit-images/</link><guid isPermaLink="true">https://montanaresearch.org/blog/getty-v-stability-first-training-data-lawsuit-images/</guid><description>Getty Images announced this week that it is suing Stability AI in the High Court in London over the images used to train Stable Diffusion. The first major rightsholder to go to court over training data picked a diffusion model, and the reasons are about evidence as much as law.</description><pubDate>Thu, 19 Jan 2023 00:00:00 GMT</pubDate><category>Open Science</category><author>Montana Research Foundation</author></item><item><title>Grokking, reverse engineered: the Fourier circuit inside modular addition</title><link>https://montanaresearch.org/blog/grokking-reverse-engineered-fourier-circuit/</link><guid isPermaLink="true">https://montanaresearch.org/blog/grokking-reverse-engineered-fourier-circuit/</guid><description>Nanda and colleagues opened up a one-layer transformer that groks modular addition and found trig identities. The progress measures they built from that circuit move long before test loss does.</description><pubDate>Wed, 18 Jan 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>Tracr and the case for ground-truth transformers</title><link>https://montanaresearch.org/blog/tracr-ground-truth-transformers/</link><guid isPermaLink="true">https://montanaresearch.org/blog/tracr-ground-truth-transformers/</guid><description>DeepMind&apos;s Tracr compiles programs written in RASP into the weights of a standard decoder-only transformer, so that the circuit inside is known before anyone looks. Why a laboratory with answer keys matters for interpretability, and what compiled models still fail to capture about trained ones.</description><pubDate>Wed, 18 Jan 2023 00:00:00 GMT</pubDate><category>Interpretability</category><author>Montana Research Foundation</author></item><item><title>The influence-operations report that predicted the year before it happened</title><link>https://montanaresearch.org/blog/influence-operations-report-predicted-the-year/</link><guid isPermaLink="true">https://montanaresearch.org/blog/influence-operations-report-predicted-the-year/</guid><description>Reading notes on the January 2023 report from OpenAI, Georgetown CSET and the Stanford Internet Observatory on language models and propaganda. What its four intervention stages actually propose, which unknowns it flagged, and where I think the framing will hold.</description><pubDate>Sat, 14 Jan 2023 00:00:00 GMT</pubDate><category>Safety &amp; Alignment</category><author>Montana Research Foundation</author></item><item><title>Cramming: what one GPU and one day can teach you about pretraining</title><link>https://montanaresearch.org/blog/cramming-one-gpu-one-day-pretraining/</link><guid isPermaLink="true">https://montanaresearch.org/blog/cramming-one-gpu-one-day-pretraining/</guid><description>Geiping and Goldstein trained BERT-class models under a strict single-GPU, 24-hour budget and re-tested nearly every trick in the pipeline. Reading notes on a rare controlled study of what matters at small scale, and why most of it is about throughput.</description><pubDate>Thu, 12 Jan 2023 00:00:00 GMT</pubDate><category>Foundations</category><author>Montana Research Foundation</author></item><item><title>GPT takes the bar exam: when professional exams became the benchmark</title><link>https://montanaresearch.org/blog/gpt-takes-the-bar-exam-professional-exams-benchmark/</link><guid isPermaLink="true">https://montanaresearch.org/blog/gpt-takes-the-bar-exam-professional-exams-benchmark/</guid><description>Bommarito and Katz gave text-davinci-003 the multiple-choice section of the bar exam and got 50.3 percent, passing in two of seven subjects. Notes on why a licensing exam looked like the ideal test and what it quietly did not measure.</description><pubDate>Thu, 05 Jan 2023 00:00:00 GMT</pubDate><category>Evaluation</category><author>Montana Research Foundation</author></item></channel></rss>