Back to News
User submissiongithub.com

Jeff Releases Small Models for Fast Zero-Shot Classification

Summary

Jeff is an independent project that fine-tunes 0.8B to 2B-parameter Qwen3.5 and Gemma 4 models for zero-shot classification. Instead of generating text, the models return a calibrated probability for each described option from a single forward pass, supporting choice, yes/no, and score questions. The project reports median decision times of 22 ms on an RTX PRO 6000 and 28 ms for the 0.8B Qwen model on an Apple M4 Max, while CPU inference is slower. Across 4,599 questions from five public benchmarks, the fine-tuned models scored 79.1, 83.1, and 81.6 overall for Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, and Jeff-Gemma4-E2B, compared with 83.0 for published Jev and 84.9 for AutoJev-27B. The results were stronger on classification and grounding tasks and weaker on reasoning-heavy benchmarks. In game tests, the 0.8B model generally performed better than the 2B version, showing that benchmark scores did not predict game play. A voice-navigation fine-tune reportedly increased held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU. Jeff was trained and tested with local hardware and synthetic data generated by an open model; the project says it does not release the training data, but releases MIT-licensed code and Apache 2.0 model weights. The authors emphasize that Jeff is a fast classifier rather than a planner, and that wording, option descriptions, and task-specific fine-tuning strongly affect results.