deployability-eval
Cheap factors, four domains: a scorecard — and why 'no single factor works' isn't 'no signal'
We put interpretable, locally-computable factors from four mechanistically-different domains — software-defect prediction, paper acceptance, AI-text detection, and a world-model physics proxy — on one shared per-factor scorecard. Predictive power has a domain-dependent ceiling, the factors don’t transfer across domains, and a single-factor reading nearly had us call an industrial benchmark a ‘collapse’ when the signal was just multivariate.
Does a multi-agent panel beat a single LLM at evolutionary code search? (We tested it. No.)
We replaced the single-LLM mutation step of an AlphaEvolve-style evolve loop (OpenEvolve) with a proposer→critic→aggregator panel and tested it rigorously across three problems. It never won — equal on easy optimization, significantly worse on hard optimization (p=0.01), and on correctness-gated ARC it cracked nothing a single call couldn’t. A clean negative result, with the baseline/ablation harness that makes it trustworthy.