Quail: Speeding up AI-SQL by jointly optimizing query planner and inference engine
Blog post from Modal
Quail is a query-aware inference engine developed by Modal and Carnegie Mellon University’s Full Stack Data Lab to accelerate AI-SQL workloads, where SQL queries generate prompts from database records to perform tasks such as filtering, classification, and fuzzy joins. Unlike chatbot and coding-agent inference, these workloads can involve millions of mostly independent, relatively low-intelligence requests and often need only a single Boolean token rather than full text generation. Quail uses knowledge of an entire SQL query plan to optimize request ordering, reuse and precisely evict transformer KV caches, avoid decode-phase overhead, reduce CPU bottlenecks, and tailor GPU kernels for prefill-heavy processing. On a complex multi-join query, the authors report throughput above one billion tokens per minute per H100 GPU, more than ten times their vLLM baseline, while their broader AI-SQL benchmark shows a geometric-average speedup of 1.84 times. The system combines conventional database planning techniques with KV-aware join ordering and hardware-based cost estimation, and it relies on open-source components such as sqlglot, PyArrow, FlashAttention, DeepGEMM, and Triton. The authors describe Quail as an early step toward specialized but reusable inference infrastructure, identifying future opportunities in tiered and cross-query caching, larger datasets, prefix-sharing indexes, improved kernel overlap, and potentially adaptive model optimization during query execution.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 6 | 747 | 162 | 79 | -85% |
| Jev | 4 | No monthly metrics for this publish month. | |||
| AI Model Fine-tuning | 1 | 139 | 28 | 14 | -75% |
| Reinforcement learning | 1 | 17 | 7 | 5 | -82% |
| Vector Search | 1 | 265 | 57 | 33 | -89% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.