Measuring When Off-the-Shelf Small Language Models Are Enough for Agent Microtasks
Summary
This study asks whether off-the-shelf small language models (SLMs) can reliably handle the microtasks surrounding a frontier LLM planner, including shell-command approval, memory writing, tool selection, and ranking prior turns. The authors build a four-task benchmark with fixed prompts, automatic metrics, and task-specific thresholds anchored to a cheap non-LLM baseline. A configuration is eligible only when its confidence bound clears the threshold. They test Qwen3 models with 0.6B, 1.7B, 4B, and 8B parameters using FP16, greedy decoding, one frozen prompt, and no tuning. None of the 16 model-task configurations passes. A log-probability decision-threshold diagnostic separates failures into four regimes, distinguishing capability deficits from cases that might be improved by changing the decoding threshold. Four-bit RTN, GPTQ, and AWQ quantization changes performance in a size-dependent way but moves no configuration into eligibility; the certification is strongest on the reconstructable hard-label tasks, with diagnostic or windowed robustness checks for the others. The result is replicated on Llama-3.x, where 12 of 12 configurations are ineligible. A threshold sweep and three neutral prompt paraphrases also preserve the finding: across the original and paraphrased cells, 0 of 112 configurations qualify. The authors therefore recommend putting SLMs behind a baseline that already meets the confidence-backed threshold, and using an SLM only when that baseline fails. In one separate example, a 4B reranker improves a BM25 shortlist by 0.047, with a confidence interval of [0.020, 0.073], but the improvement itself does not certify eligibility.