evaluation
A shipped exit rule picks good steps — and still costs answers
A recurrent-depth language model ships six ways to stop early. Measuring one of them at its own default: it selects better stopping points than chance, and on one task it still loses accuracy on the answers. Both are true, and separating them needed two different experiments.
Cheap factors, four domains: a scorecard — and why 'no single factor works' isn't 'no signal'
We put interpretable, locally-computable factors from four mechanistically-different domains — software-defect prediction, paper acceptance, AI-text detection, and a world-model physics proxy — on one shared per-factor scorecard. Predictive power has a domain-dependent ceiling, the factors don’t transfer across domains, and a single-factor reading nearly had us call an industrial benchmark a ‘collapse’ when the signal was just multivariate.