10. Spider 72%: The Dataset, Not the Model, Shapes the Error Profile

Three multi-candidate experiments from post 9 produced no meaningful improvement. Before moving on to the Schema Binding Plan, I wanted to test one more hypothesis: model tier as the bottleneck. I expected Gemini Pro to improve the current 56% accuracy baseline over flash-lite. I also ran the error-type classification and Spider experiment as part of that test. Switching to Pro Made Things Worse Prompts, context, and pipeline stayed the same. Only the model changed, and the same 50 BIRD questions ran against it. The metric is EX (Execution Accuracy): whether the generated SQL produces the correct result when executed against the database. ...

April 24, 2026 · 7 min · Junho Lee