Modal and Carnegie Mellon University’s Full Stack Data Lab introduced Quail, the Query-Aware Inference Layer, for AI-SQL workloads that apply language models to database rows and produce tabular results. Unlike NL2SQL, AI-SQL extends SQL with functions such as AI_FILTER, AI_JOIN, and AI.IF; these operations can create millions of short model requests, a workload that differs from chatbot and coding-agent inference. Quail jointly optimizes the SQL query plan and model execution, using request ordering, KV-cache awareness, predicate pushdown, selectivity-based filter ordering, and join-order search that keeps non-dominated plans with different token, attention, and cache costs. Its execution engine is derived from vLLM and adds Triton kernel fusion, specialized attention handling for shared join prefixes, and a reduced output vocabulary of eight truth-value options. The system focuses on prefill-only inference for Boolean classification, so it does not use sampling, decode separation, CUDA Graph capture, or speculative decoding. In the reported results, Quail exceeded one billion tokens per minute per H100 GPU on a multi-join query and cost under six cents per billion tokens on Modal; across the authors’ AI-SQL benchmark, it was 1.84 times faster than vLLM on a geometric mean, although two queries were included specifically to expose areas needing improvement. The benchmark also contained an agent-trace case where Quail lagged vLLM because it could not yet exploit cross-document prefix sharing. The release includes a runnable example using an H100, the Qwen3-4B-FP8 model, an IMDb dataset, and the `quail-engine` package. The authors identify multi-tier and cross-query KV caching, larger-than-memory datasets, radix indexes, better kernel overlap, and possible on-the-fly model specialization as future work. They also argue that specialized engines may increasingly share reusable open-source cores rather than being built independently for every workload.
