Grading our own homework
Khaliq Gant6 min readWe benchmarked RelayFlow against a raw coding agent on Terminal-Bench 2.0. The first result was a tie, and the reason was two bugs, both of which were quietly working against us. Here is what we found, what we fixed, and the honest number that came out the other side.

