12. Racing Three Coding Agents in Isolated Git Worktrees

I build DataNexus with several coding agents running at once. Claude Code fixes router logic while Codex fills in test coverage. The workspace was the problem. With one checkout and multiple terminals, agents step on each other’s files. One reinstalls dependencies while another’s build breaks mid-run. I juggled stashes for a while, then gave up and ran them one at a time. Agents I added for parallelism were running serially. ...

July 11, 2026 · 4 min · Junho Lee

10. Spider 72%: The Dataset, Not the Model, Shapes the Error Profile

Three multi-candidate experiments from post 9 produced no meaningful improvement. Before moving on to the Schema Binding Plan, I wanted to test one more hypothesis: model tier as the bottleneck. I expected Gemini Pro to improve the current 56% accuracy baseline over flash-lite. I also ran the error-type classification and Spider experiment as part of that test. Switching to Pro Made Things Worse Prompts, context, and pipeline stayed the same. Only the model changed, and the same 50 BIRD questions ran against it. The metric is EX (Execution Accuracy): whether the generated SQL produces the correct result when executed against the database. ...

April 24, 2026 · 7 min · Junho Lee

9. BIRD 56%: Nine Experiments and What Got Ruled Out

I hit 80% on my own 30-question benchmark, but only 56% on BIRD Mini-Dev’s 50 public questions. Nine experiments later, I had ruled out the multi-candidate hypothesis from three different angles. What’s left is schema understanding and methodology.

April 19, 2026 · 6 min · Junho Lee