• 4 min read
Thomson Reuters says its model beats GPT-5.5
Thomson Reuters says its new Thomson model beats GPT-5.5 on selected benchmarks and will debut in CoCounsel Legal in August.

Image: Hacker News
Thomson Reuters says its new Thomson model is competitive with leading frontier systems, including Claude Opus 4.8, while outperforming GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro on selected evaluations. The company plans to launch it later this summer, with its first production deployment scheduled for August.
How Thomson was built
Thomson Reuters acquired AI research company Safe Sign Technologies in 2024. Rather than relying solely on increasingly capable general-purpose models, the company says it developed a model specifically for professional domains, standards, and accountability requirements.
Thomson starts with an open-source foundation for general-purpose capabilities. Thomson Reuters then applied mid-training and post-training techniques using decades of content from Westlaw, Practical Law, Checkpoint, and Reuters. Hundreds of subject-matter experts evaluated outputs, documented failure modes, and validated whether the model’s reasoning reflected how legal professionals work.

Recommended reading
Qualcomm robot collapses during demo after connection loss
The company says customer data is never used to train the model. It also claims Thomson delivers performance comparable to much larger models while requiring a fraction of their size and training and operating costs. Less than 10% of Thomson Reuters' content has been used in training so far, according to the company.
Benchmark results across legal and general tasks
Thomson Reuters evaluated Thomson-1-Large against Google DeepMind’s Gemini 3.1 Pro, Anthropic’s Opus 4.8, and OpenAI’s GPT-5.5 across legal and general-purpose categories, including coding, tax, accounting, multilingualism, journalism, agentic tasks, safety, long context, reasoning, and instruction following.
The published scores include:
- Stanford LegalBench: Thomson 0.823; Gemini 3.1 Pro 0.843; Opus 4.8 0.818; GPT-5.5 0.832.
- PrBench Legal Hard: Thomson 0.352, the highest score; Gemini 3.1 Pro 0.293; Opus 4.8 0.315; GPT-5.5 0.333.
- Harvey Legal Agent Benchmark: Thomson 0.857; Gemini 3.1 Pro 0.555; Opus 4.8 0.869; GPT-5.5 0.781.
- Instruction following: Thomson 0.914, the highest score; Gemini 3.1 Pro 0.848; Opus 4.8 0.861; GPT-5.5 0.885.
- Reasoning: Thomson 0.684; Gemini 3.1 Pro 0.748; Opus 4.8 0.737; GPT-5.5 0.589.
- Coding: Thomson 0.399; Gemini 3.1 Pro 0.500; Opus 4.8 0.598; GPT-5.5 0.414.
- Long context: Thomson 0.753, the highest score; Gemini 3.1 Pro 0.750; Opus 4.8 0.741; GPT-5.5 0.703.
The benchmark image identifies the best result in each row in green.
Thomson Reuters says instruction following combines the IFEval and FollowBench benchmarks. Reasoning combines GPQA Diamond, HLE, and MMLU-Pro; coding combines SWE-Bench Pro and Terminal-Bench 2.1; and long context combines Infinity Bench with internal Thomson Reuters evaluations.
The comparisons are not identical in configuration: Thomson-1-Large used test-time scaling, Gemini 3.1 Pro and Opus 4.8 used reasoning mode, and GPT-5.5 used non-reasoning mode. The source presents these as Thomson Reuters' evaluations; it does not provide third-party verification of the results.
Proprietary data and legal research testing
The company also tested Thomson on 53 legal research queries written by internal subject-matter experts. In that evaluation, models were connected to Thomson Reuters' Westlaw and Practical Law through an in-house agentic harness, while competing web-based research used the Brave search engine.
The evaluation measured two criteria:
- Completeness: whether an answer addressed every element listed in subject-matter-expert rubrics.
- Factuality: whether cited sources supported the claims made in each report.
Large language models acted as judges, with the metrics calibrated against subject-matter-expert scoring. Thomson Reuters argues that direct access to Westlaw, Practical Law, and Reuters gives Thomson an advantage in completeness and factuality over frontier models restricted to unrestricted web access.
The company says its broader training and evaluation process includes expert-written real-world queries, agentic workflows, deep-research training with human-calibrated judges, data-centric mid-training, and human and automated red-teaming. It did not disclose the model’s parameter count, training cost, or operating cost.
Thomson’s first deployment in CoCounsel Legal
Thomson’s first integration will arrive in August inside Tabular Analysis in CoCounsel Legal. The tool performs high-volume, structured document review against a measurable accuracy standard, which Thomson Reuters says makes it a clear test of the benefits of a domain-specific model.
Thomson will be the default model for Tabular Analysis. Over the following year, the company plans to integrate it across its legal and tax product portfolio. The announcement does not give a standalone price for Thomson or specify an exact August release date.
Thomson Reuters frames the model as the fourth part of its professional AI stack, alongside authoritative content, domain expertise, and software tools. Its stated standard is “Fiduciary-Grade AI™” for professionals with duties of care and accountability—a rationale for building a model trained on proprietary content rather than treating a general-purpose system as the final product.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via Hacker News


