QVAC Genesis III is an open synthetic corpus designed to address the shortage of high-quality STEM pre-training data for smaller language models aimed at edge AI and on-device deployment. The corpus contains 191.43 billion tokens spanning 19 STEM domains, multiple difficulty levels, and different educational styles. Its data-generation process uses targeted teacher distillation with a weak edge-scale student model: student failures are turned into corrective explanations, while successful answers are expanded into contrastive reasoning over all answer choices. The authors also introduce an LLM-as-a-parser protocol that extracts final answers from free-form outputs and measures both accuracy and answer validity. In controlled from-scratch ablations, 1.7B-parameter models trained on QVAC Genesis III consistently outperform models trained on the open synthetic corpus Cosmopedia-v2 and the public Cosmo-1B model across ARC, GPQA Diamond, and MMLU STEM. The reported improvements reach 28.57% on ARC-E and 21.35% on ARC-C, while the Valid Answer Rate reaches 99.45%. The results are presented as evidence that carefully generated, STEM-focused data can raise the learning value of each token for smaller models.
AI News
The latest AI releases, research, products, and industry updates.
Loading...