How Quail Estimates the Cost of AI-Powered Filters
Summary
The article explains how Quail, an AI-SQL execution engine, estimates the latency of filters that ask an LLM to classify each document as true or false. Instead of profiling every model, GPU, and workload combination, Quail uses a roofline model that converts floating-point operations and HBM traffic into a speed-of-light latency lower bound. The cost model separately accounts for transformer projections, attention, and the MLP, using FP8 throughput for projections and MLPs and BF16 throughput for attention on an NVIDIA H100. The example uses Qwen3-4B FP8, whose grouped-query attention reduces KV-cache storage, and estimates costs from request length, token count, attention comparisons, and the number of packed forward passes. For 5,000 IMDB reviews, one filter is estimated at 6.64 seconds, with the MLP and projections compute-bound. For multiple filters, Quail retains document-prefix KV vectors in HBM, limits batches to roughly 1,200 documents in the example, and estimates later-filter work from selectivity and instruction length. It orders filters by ask cost divided by rejection rate, tests each possible first filter, and selects F2 → F1 → F3 for the example. The ordering score is 6.8665 seconds, while a query-wide roofline calculation gives a 6.86-second lower bound after aggregating work and allowing arithmetic-memory overlap across filters. The authors caution that actual runtime may be higher because peak hardware rates and perfect overlap are unlikely. They identify extensions for AI joins and classification, hybrid models, and newer hardware such as Blackwell GPUs.