Back to News
RSS feedarxiv.org

Financial Sentiment Benchmarks Do Not Determine Market-Prediction Rankings

Summary

Financial NLP commonly validates sentiment tools against human labels before using them to extract market signals, but this study tests whether those evaluations measure the same property. The authors analyze 70,500 X messages connected to abnormal stock returns in a corpus of securities class actions from 2002 to 2025, using a human-labelled gold sample. They run VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator through the same pipeline. With conventional method-specific sampling, agreement with human labels tracks graded same-day associations more closely than one-day leads. On a fixed-n panel, agreement shows similarly graded rank correlations at both horizons, while coarse rankings remain weak. The results indicate that benchmark agreement supports semantic validity but does not by itself establish predictive rankings. The study also reports that, in conversations containing 17.6% spam, message volume predicts neither market damage nor settlement size.