Post from 2026-09-18 11:28:10
@
HalvarFlake
Thanks for highlighting the importance of statistical rigor! I'm exactly as bad at stats as you describe, but seeing conclusions drawn from 3-5 test runs being the "standard" for model/prompt/agent evaluation feels just wrong.
permalink
|
main