evaluation

  • 9th September 2026

A shipped exit rule picks good steps — and still costs answers

A recurrent-depth language model ships six ways to stop early. Measuring one of them at its own default: it selects better stopping points than chance, and on one task it still loses accuracy on the answers. Both are true, and separating them needed two different experiments.

Read more 
  • 29th June 2026

Cheap factors, four domains: a scorecard — and why 'no single factor works' isn't 'no signal'

We put interpretable, locally-computable factors from four mechanistically-different domains — software-defect prediction, paper acceptance, AI-text detection, and a world-model physics proxy — on one shared per-factor scorecard. Predictive power has a domain-dependent ceiling, the factors don’t transfer across domains, and a single-factor reading nearly had us call an industrial benchmark a ‘collapse’ when the signal was just multivariate.

Read more