LatticeFlow Introduces Independent Framework to Measure Political Bias in LLMs
Summary
LatticeFlow AI has introduced what it describes as an independent framework for measuring political bias in large language models. The framework compares model responses across Chinese, US, and European political spectra without using human-written neutrality rubrics or another AI model as a judge. It decomposes answers into claims, derives positions from agreement and disagreement among reference models, and ties each result to a SHA-256 hash of the tested weights. In its first evaluation, 17 models were assessed across 11 spectra using roughly 554 samples per axis, including direct and indirect probes and control datasets. LatticeFlow reports that Qwen 3.7 Max was farther toward the Chinese pole than Qwen3 32B across all six China-politics categories, and was the most Chinese-aligned tested model on freedom of religion and ethnic issues. GLM 5.2, Kimi K2.6, Qwen 3.7 Max, MiniMax M2.7 FP4, and DeepSeek V4 Pro were positioned toward the Chinese end across every tested category, generally through reframing rather than refusal. The framework also separated Western models: Grok 4.3 and Grok 3 appeared at one end of the US-politics axis, while GPT-5.5 and GPT-5.4 appeared at the other. LatticeFlow emphasizes that the results are dated records rather than a continuously updated leaderboard, and that all models were self-hosted with provider-side guardrails disabled. The framework can be applied to self-hostable models and, when providers participate, proprietary models through APIs, with custom categories available for specific markets or regulatory contexts.