Shuo Qiu’s Notes

Why the Coding Benchmark Audit Never Ends

A coding benchmark can run out of signal long before anyone scores 100% on it. Four audits later, every improved version of SWE-bench still had broken tasks.

How We Learned to Measure Coding Agents, and Why We Still Can't

Nobody can tell you which coding agent is best, and a bigger benchmark won't fix that. Every evaluation decides what to test, how to grade it, and which tasks to sample, and each one covers only a narrow aspect. What works instead is assembling several imperfect measurements and knowing each one's blind spot.

Don’t Shop for Evaluators. Let Your Coding Agent Build One.

Don't shop for pre-built LLM judges. Have a coding agent read your real task material (code, docs, traces) and write the judge. It's faster than shopping, and on tau-bench telecom agreement with ground truth more than doubled.

Your Agent and Harness Aren't the Asset, Your Eval Is

The durable asset in agent development is not the prompt or the harness. It is the eval: the specification of what good looks like, where agents fail, and what customers actually need.