Thomson Reuters has announced early benchmarking results for Thomson, an AI model built for legal and other professional work. The company says Thomson performed competitively with leading frontier models, including Claude Opus 4.8, and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro on the evaluated categories. The model starts from an open-source foundation and was further trained with mid-training and post-training techniques on content from Westlaw, Practical Law, Checkpoint, and Reuters, with hundreds of subject-matter experts evaluating outputs and failure modes. The company says customer data is not used for training, and that Thomson is smaller and less costly to train and operate than many general-purpose models. Evaluations covered legal work, coding, tax, accounting, multilingualism, journalism, agentic tasks, safety, long context, reasoning, and instruction following, using public and internal benchmarks. Thomson Reuters says less than 10% of its content has been used so far, leaving room for further training and validation. In a separate evaluation of 53 legal research queries, the model’s access to proprietary sources was assessed for completeness and factuality against frontier models with web access. Thomson is scheduled to become the default model for Tabular Analysis in CoCounsel Legal in August, followed by planned integration into legal and tax products over the next year.
