Post from 2026-09-18 11:28:10

@HalvarFlake Thanks for highlighting the importance of statistical rigor! I'm exactly as bad at stats as you describe, but seeing conclusions drawn from 3-5 test runs being the "standard" for model/prompt/agent evaluation feels just wrong.
permalink | main