Back to News
RSS feedarxiv.org

Sigma-Hunter Adapts Language Models for Threat Detection Rules

Summary

Detection engineers often turn threat reports, forensic observations, and hunt hypotheses into Sigma rules, but general-purpose language models can produce invalid YAML, use the wrong log sources, reference unsupported fields, or create overly broad logic. The paper introduces Sigma-Hunter, a domain-adapted language model intended to assist with Sigma rule generation and threat hunting. The authors derive 7,663 question-answer and analyst-reasoning examples from 3,635 validated open-source Sigma rules, assigning each source rule to a train, validation, or test split before expansion to prevent rule leakage. They fine-tune a 7B Mistral model and a Phi-4 model with LoRA, then evaluate held-out generations for syntax, approximate field consistency, and the semantic quality of detection logic, including completeness, selectivity, and log-source alignment. Sigma-Hunter-Mistral achieves an overall score of 8.17, compared with 7.88 for the strongest general-purpose baseline and 4.61 for untuned Mistral. The results indicate that domain adaptation can make a compact 7B model competitive with larger general-purpose models on this structured task. They also show that syntactic validity alone is a weak indicator of useful detection logic, because some baselines produce valid YAML with poor semantics. The adapted models run locally, which supports use in disconnected environments where analysts cannot access hosted model services.