llm
Does a multi-agent panel beat a single LLM at evolutionary code search? (We tested it. No.)
We replaced the single-LLM mutation step of an AlphaEvolve-style evolve loop (OpenEvolve) with a proposer→critic→aggregator panel and tested it rigorously across three problems. It never won — equal on easy optimization, significantly worse on hard optimization (p=0.01), and on correctness-gated ARC it cracked nothing a single call couldn’t. A clean negative result, with the baseline/ablation harness that makes it trustworthy.
The moat in LLM chip design isn't the model — it's the environment
A field note on 2025–2026 agentic RTL generation: every result that moved the needle came from the verification loop wrapped around the LLM, not the LLM itself. For someone who builds compilers and ML infra, that loop is the product.
When the Government Pulls a Model: The Fable 5 / Mythos 5 Export-Control Suspension
On June 12, 2026, a US government directive forced Anthropic to abruptly disable Claude Fable 5 and Mythos 5 for all customers. A careful, source-verified account of what actually happened — separating the confirmed facts from Anthropic’s characterization, and from the jailbreak-research claims that are not yet independently established.