Thomson Reuters announced the launch of its independently developed AI frontier model Thomson, which the company plans to deploy throughout its legal and tax portfolio over time, starting with Tabular Analysis inside CoCounsel Legal, where it now provides high-volume document review capacities.
Overall, Thomson was described as an AI with specialized capabilities within the professional domains where Thomson Reuters has unique expertise and content. It is not thought of as a standalone chatbot or desktop app, but as an embedded intelligence within Thomson Reuters products. It focuses on performing professional tasks such as reviewing documents, validating whether claims are supported by source material and checking citations. While it can operate as an agentic system and perform a broad array of tasks one would expect of any other model of its kind, it is not planned to be the top-level orchestrating model in every workflow, despite playing a critical role within the overall agentic system.
While a precise definition of "frontier models" can be somewhat slippery, they can be

Many AI models, even foundation models, are initially licensed from a larger frontier model such as ChatGPT, Gemini or Claude. This means that even if the AI has been rebranded and heavily modified for a specific purpose, it is still ultimately controlled by whoever owns the underlying frontier model, and all data passed through the derivative models ultimately winds up in the frontier model's servers. This is why changes to frontier models, such as going from ChatGPT 5.5 to 5.6, can create unpredictable changes in the derivative models attached to it: It may be a vital part of one's own infrastructure, but it remains under the control of a third party. This is part of why Thomson Reuters decided to make its own frontier model.
"AI sovereignty is about owning the layers of the stack that matter to you, but ownership does not mean exclusivity," wrote Thomson Reuters chief technology officer Joel Hron in a
During a webcast last week he compared it to renting versus owning a home. While renting a home might be cheaper and easier since it is the landlord who maintains the property, owning a home not only allows for complete control over the space but also builds equity in an actual asset that is owned long term, which compounds over time.
Beyond control issues, frontier model licensing can also introduce privacy issues, as data passed through the models can wind up on the servers of the companies that own them. Regarding the privacy structure of its own frontier model, Hron said Thomson Reuters engaged independent third-party security organizations to evaluate the hosting, serving and security controls surrounding the model. While the company has not disclosed the details of those assessments publicly, he said they discuss the applicable safeguards directly with customers as part of their security review.
Creating a frontier model is generally considered extremely expensive. While a standard foundation model might cost thousands to millions to develop, frontier models historically have
Hron, in a later email, said the $40 million represents broader development investment, covering talent and compute costs, including staff, domain-expert participation, research, data curation, infrastructure engineering, experimentation, evaluation, safety testing and vendor partnerships for the last several years that Thomson Reuters has been working on this initiative. A significant portion of that investment went into reusable capabilities and infrastructure, including the model-development pipeline Thomson Reuters can apply to future open-weight foundations. He specified the company did not purchase data to train Thomson, nor did it draw on its own customer data. Instead, the model was trained through Thomson Reuters proprietary content and work created by its own subject-matter experts who provided the training foundation. These experts recreated relevant professional tasks over thousands of hours during the training and evaluation process. The final training run for Thomson 1.0-Large cost less than $450,000 in GPU compute over approximately three weeks.
"The distinction is important: The $40 million reflects the people, research, infrastructure, experimentation and model-development capabilities built over a much longer period, while $450,000 represents the compute cost of the final training run itself," he said.
A
The result was something that was never meant to be all things to all people but, rather, a highly specialized model trained on the professional domains where Thomson Reuters concentrates, though Hron added that its training is designed to retain as much of this general capability as possible while gaining the specialized knowledge that they imbue into it. And so while Thomson is expected to handle more and more tasks across the company's products, it is not conceived of as a wholesale replacement of the models they currently use, and the company will continue to use leading third-party models where they are best suited to the task.
"Thomson will be applied where its professional specialization, trust, validation and efficiency provide the strongest advantage. Other leading models may be used for capabilities where they are better suited. The system can therefore route work to the appropriate intelligence for each task rather than relying on one model for everything," said Hron.
As the model's first deployment is embedded within CoCounsel Legal, it is being delivered as part of a Thomson Reuters product rather than sold today as a separate subscription or consumption-based API. However, Hron said the company has started to build an API that will allow people to access the model directly — people will be able to request API keys to the model, set parameters for how it operates and invoke it in code. He cautioned that this is in the very early stages, but ultimately it is a key part of their roadmap.
When asked whether Thomson Reuters plans to license its model much the same way that OpenAI, Google or Anthropic license theirs, Hron said the first priority is improving its use within the company's own products. But he added that they have heard some interest from customers about licensing the model directly, and are exploring those opportunities. But in the immediate term, the focus for this upcoming API is less direct licensing and more supporting developers and enabling external validation. He added that the company is also making Thomson-1.0-Small available as an open-weight model on Hugging Face under a noncommercial academic license.
Looking to the future, the company plans to continue training the model on house content; as much effort as it puts into training Thomson, the company said it has only used less than 10% of its available data, so there is much more they could teach it. But beyond static data, the company plans to also generate more training data based on how its products are actually used; to do this, they have already built reinforcement learning environments that trains the model on Thomson Reuters products to create more specialized knowledge to support it.
And that is just the plan for Thomson itself. During the webcast, lead researcher Andrew Bean said now that the company has a repeatable process for creating frontier-level models at scale, it will likely release more in the future.
"This is really meant to be a reusable pipeline," he said. "We've built on one particular base model. But really, as the frontier of open models continues to advance, we can continue to build on that, and so we see a lot of room to run following on bigger and better base models, more and more comprehensive use of our data, and finally closer reinforcement learning on things that we really care about."






