Back to News
RSS feedarxiv.org

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

Summary

A new study asks whether large language models need human-readable text to fine-tune effectively. It introduces Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to optimize continuous synthetic input embeddings. The embeddings are used directly for downstream fine-tuning, while discrete token projections are reserved for qualitative inspection rather than training. In experiments covering six Llama and Qwen models ranging from 1B to 32B parameters, the method was evaluated on six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA settings, DASA performed comparably to source natural-language data and exceeded it in multiple configurations. It also outperformed GRADMM in most comparisons, with tests covering both general-domain and task-specialized source data. Under the reported synthesis settings, DASA was 3.6 to 4.9 times faster than GRADMM while using comparable peak GPU memory. The results support the study’s claim that useful fine-tuning updates can be targeted without first producing human-readable source text, although the findings are limited to the evaluated models, tasks, and synthesis settings.