Legal AI Splits Between Owned Models and Frontier APIs
On Monday, August 24, Thomson Reuters disclosed a $40 million investment in Thomson, which it calls a legal AI frontier model. Training began from an open-source foundation and used less than 10% of the company’s proprietary content. Four days earlier, Harvey had introduced Tenet, a research model post-trained from Moonshot AI’s open-weight Kimi K3. The two announcements put a new expense inside legal software: owning more of the model layer.
That same four-day interval also showed why ownership is not a clean break from outside providers. Thomson Reuters launched the next version of CoCounsel Legal on August 20 using Anthropic’s Claude Agent SDK. Thomson is due to power Tabular Analysis first, while other models remain in the product. Harvey worked with Fireworks AI on Tenet. Neither company built every layer alone.
LexisNexis has made the opposing case in public. Its Protege product offers legal workflows grounded in LexisNexis content while its general AI mode gives users models from Anthropic, Google, and OpenAI. Chief Technology Officer Greg Dickason has described the company’s approach as technology-agnostic and multi-model.
Three vendors can therefore start with similar assets, including proprietary legal material, expert reviewers, professional customers, and recurring software revenue, then make different model decisions. One owns a new foundation model. One post-trains an open-weight base for selected workloads. One keeps several frontier providers in reach. All three still need retrieval, tools, citations, evaluation, security, and a lawyer who remains responsible for the work.
This choice reaches the P&L and the org chart. It determines where inference margin goes, which technical dependencies move inside, and who maintains evaluation data. It also changes the evidence a law firm or legal department should request before trusting a benchmark or signing a renewal.
An owned model can be cheaper per query and better fitted to a narrow workflow. It also brings training cost, refresh cycles, failure analysis, and obsolescence onto the owner’s books. A rented model can preserve access to rapid improvements and spread research cost across a provider’s customer base. It can also expose a product to pricing, architecture, and availability decisions made elsewhere.
A legal AI buyer needs to know which combination of model, content, tools, people, and contractual rights produces reviewable work at an acceptable total cost. This week’s announcements supply enough detail to begin that file, but not enough to close it.
Four days put Tenet and Thomson on different tracks
Harvey published its Tenet research preview on Thursday, August 20. The company said a team spent six months post-training Kimi K3 with synthetic data, public data, and examples produced by legal experts. It said no customer data entered the training set. Fireworks AI supplied training and inference infrastructure.
The word post-training matters. Harvey did not report pretraining a general model from an empty parameter set or assembling its own compute cloud. It started with an open-weight base that already contained broad language and reasoning capability, then changed its behavior for legal tasks and Harvey’s product environment. That route gives the company more control over weights and serving than an API-only integration without requiring the full cost of a general foundation model.
Harvey tested different Tenet variants in three settings: its Legal Agent Benchmark, high-volume document review, and work over a firm’s internal knowledge. The company reported that one Tenet configuration completed almost twice as many held-out benchmark tasks as the unmodified Kimi K3 base. On the contracts portion, it reported a gain of more than 20%.
The largest operating claims came from narrower experiments. In Review Tables, which can examine as many as 10,000 documents, Harvey reported a 3.6-point answer-quality gain, a 12.1-point citation-quality gain, and roughly one-tenth the model cost. In a firm-knowledge experiment, it reported a 58% reduction in tokens, nearly 10% more completed tasks, and a 90% reduction in cost per query.
Those are builder results, not an independent audit. The preview describes a research model, company evaluation sets, and LLM judges. It does not disclose a general-availability date, the distribution of errors, a law firm’s realized savings, or whether lawyers spent less time reviewing final work. Full technical reports remain forthcoming.
On Monday, Thomson Reuters announced Thomson. The company also began with an open-source foundation, then trained a model it says it fully owns and controls. It disclosed $40 million for talent and compute and said the training used less than 10% of its proprietary content.
Thomson sits one layer deeper than a specialized API integration, but its development network remains visible. The company’s engineering account names DatologyAI, Lambda, Together AI, Imperial College London, and internal subject-matter experts as contributors. Ownership of the resulting model does not mean solitary production.
The first deployment is deliberately bounded. Thomson Reuters says Thomson powers Tabular Analysis in the next CoCounsel Legal release. The workflow turns questions about a document set into a table with source-linked answers. CoCounsel can handle up to 10,000 documents and 100 questions in that surface. A model optimized for repeated extraction and citation can produce an immediate cost advantage there even if another model remains better for open-ended research or coding.
The four-day sequence exposes three decisions that vendor marketing often merges: who owns the base or adapted weights, who operates inference, and which model a product routes to for each task. A buyer needs all three answers.
Forty million dollars buys more model control
Thomson Reuters called $40 million an investment in talent and compute. That is the clearest public price attached to the model, but it is not Thomson’s total cost. It excludes 175 years of accumulated content, prior acquisitions, the time of hundreds of experts, product integration, sales, customer deployment, and future retraining. It also excludes the cost of errors that pass through a benchmark and reach a professional workflow.
The figure is easier to understand against Thomson Reuters’ existing economics. In its second-quarter filing, the company reported $1.954 billion in quarterly revenue, 82% of it recurring, and a 38.1% adjusted EBITDA margin. It forecast capital expenditures near 8% of revenue for 2026 and $680 million to $690 million of amortization for internally developed software.
Those company-wide figures do not establish Thomson’s return. They explain why a professional information provider can consider an investment that would be implausible for most law firms. Thomson Reuters already has paying distribution, content rights, product surfaces, expert labor, and enough query volume to spread a fixed model cost. The economic test is not whether $40 million is large. It is whether model ownership lowers variable cost or improves retention and product quality enough to beat the internal capital charge and continuing maintenance.
Thomson Reuters offers an initial quality case. On its published table, Thomson scored 0.823 on LegalBench, 0.352 on PRBench Legal Hard, and 0.857 on Harvey’s Legal Agent Benchmark. It scored 0.914 on an instruction-following measure. Its reported reasoning score of 0.684 and coding score of 0.399 were lower than some frontier alternatives.
The mixed profile is more informative than a single winning rank. A legal product does not need one model to dominate every general capability. It can route coding, open-ended reasoning, research, and repeated document analysis differently. Ownership becomes valuable when a high-volume task matches the model’s strength and when the product has a credible fallback for the weaknesses.
Another Thomson Reuters evaluation joined its model to proprietary research tools. On 53 legal research queries, the company reported citation factuality of 0.83, compared with 0.65 and 0.68 for two systems using frontier or open models with open-web tools. Subject-matter experts created a rubric and reviewed the evaluation design; an LLM judge scored outputs after calibration.
That 0.83 measures the whole research system. Proprietary sources, retrieval, query planning, citation controls, and the harness can account for part of the difference. The sample is small, the builder chose the tasks and tools, and the full technical report has not been published. A legal buyer cannot treat the figure as an expected accuracy rate for its own documents.
Stanford’s 2026 AI Index technical chapter provides a wider counter-signal. As of March, the reported performance gap between leading closed and open models on Chatbot Arena had narrowed to 3.4%, and the top four providers were separated by fewer than 25 Elo points. When broad model scores converge, cost, latency, reliability, data access, and domain optimization carry more weight.
That finding does not rank Thomson or Tenet, both announced after Stanford’s snapshot. Chatbot Arena is not legal work. The same chapter also records invalid-question rates as high as 42% in some widely used benchmarks. A precise score can overstate knowledge when questions are broken, leaked, ambiguous, or easy for the wrong reason.
Owning the model also means owning the benchmark burden. An API customer can challenge a provider, route elsewhere, or negotiate a service issue. An owner has to reproduce failures, maintain the evaluation set, decide when a new base has overtaken the current model, and fund the next training cycle. The asset and the aging risk arrive together.
Joel Hron, Thomson Reuters’ chief technology officer, described model control as a way to connect the company’s data and expert knowledge directly to training. The connection remains a forecast until production records show query cost, response time, correction rate, customer adoption, and renewal against an honest alternative.
Harvey trains Kimi K3 inside a legal harness
Harvey’s route makes the middle ground easier to see. Kimi K3 supplied the general capability. Harvey supplied workflow traces, legal examples, expert labels, product constraints, and an evaluation environment. Fireworks supplied infrastructure. The resulting Tenet variants can be served and tuned around Harvey workloads without asking Moonshot AI to change a closed API.
Product logs and reviewer corrections can become model-development evidence. A legal AI company sees which steps agents attempt, where tools fail, how documents are cited, and which long-context tasks consume too many tokens. If it can use those signals without putting confidential customer material into training, it can optimize for actual work rather than a generic chat leaderboard.
Harvey’s preview says its training set excluded customer data. That statement answers one question and leaves several others. Buyers still need to know what production prompts are retained, which telemetry supports evaluation, how experts’ examples are licensed, how model weights and logs are protected, and whether future opt-in programs change the boundary. Training, evaluation, retrieval, and inference have different data paths.
The Review Tables result shows the strongest economic logic. A user may ask the same set of questions across thousands of contracts. General reasoning matters, but so do schema discipline, citation placement, consistent tool use, and the ability to avoid rereading irrelevant text. A narrower model that uses fewer tokens can save money on every row.
Cost per query is still not cost per completed legal task. The denominator must include failed rows, human corrections, reruns, escalation, and the time spent checking sources. A tenfold model-cost reduction can disappear if a lower price encourages ten times as many low-value queries or if reviewers must inspect more exceptions. It can also understate value if higher citation quality removes substantial review work.
Harvey’s firm-knowledge experiment raises a second opportunity. Law firms possess precedents, memos, templates, and matter histories that rarely belong in a general training corpus. A model adapted to retrieve and reason across that material can make institutional knowledge more available. The same system can amplify an outdated clause, a local drafting habit, or a work product that was correct only for one client.
A usable answer needs provenance: document, date, jurisdiction, matter restriction, and authoring context. The model can shorten discovery, but it cannot decide that a prior matter’s conclusion transfers to the present one.
Calvin Qi, Harvey’s co-founder and chief product officer, framed Tenet as the result of building models inside actual legal workflows. That framing distinguishes a domain lab from a general model lab. It also defines the evidence the company should publish next: task distributions, reviewer agreement, error categories, latency, total cost, and results from customers who were not involved in model development.
Harvey can serve Tenet for repetitive, expensive workloads and keep external frontier models for tasks where their reasoning, context, or tool ecosystem remains stronger. It can replace the base in a later cycle. A credible internal route gives the company bargaining room without requiring independence from the model market.
LexisNexis keeps frontier providers in view
LexisNexis offers the countercase. Its current Lexis+ with Protege page describes legal AI grounded in LexisNexis content and Shepard’s citation treatment. In general AI mode, users can choose models from Anthropic, Google, and OpenAI. The product joins external model capability to an internal content and workflow layer.
In a March 9 article, Dickason and Anthropic General Counsel Jeff Bleich described the Anthropic Legal Plugin inside Protege. They also stated LexisNexis’ longer position: multi-model, multi-tool, and multi-agent rather than dependence on one provider.
LexisNexis has more than 2,000 technologists, data scientists, and subject-matter experts, yet external frontier providers can still carry useful research cost. They distribute pretraining across a much larger market and may improve general reasoning, multimodality, coding, context, and agent tools faster than a domain vendor can. A router can assign work to the model that fits it today.
Multi-model does not remove lock-in. It relocates it. The product depends on provider contracts, token prices, rate limits, regional availability, model retirement schedules, and changing behavior. Routing logic, prompts, and evaluations can become tied to provider-specific features. A provider update may require regression tests before a legal workflow returns to service.
An owned model has its own lock-in. Training pipelines, serving infrastructure, evaluation harnesses, and specialist teams become fixed commitments. Replacing an API endpoint may take weeks. Replacing an internal model program can take budget cycles and affect careers. A balance sheet can hide switching cost as readily as a vendor agreement.
The public evidence now supports three routes, each with a different missing field:
| Published route | Control gained | Dependency retained | Public proof | Evidence still missing |
|---|---|---|---|---|
| Thomson Reuters trains Thomson from an open-source foundation | Model ownership, training choices, inference deployment, closer connection to proprietary content | External compute, tooling, research partners, base-model advances, internal refresh budget | $40 million disclosed; benchmark tables; first Tabular Analysis placement | Independent production quality, total lifecycle cost, adoption, correction time, renewal effect |
| Harvey post-trains Kimi K3 into Tenet variants | Adapted weights, workload-specific serving, lower reported model cost, fallback leverage | Open-weight base quality, Fireworks infrastructure, future base selection, internal evaluation labor | Research preview; LAB, Review Tables, and firm-knowledge experiments | General availability, independent tests, reviewer effort, realized customer savings |
| LexisNexis routes among frontier providers | Access to several research programs, model choice, provider competition | API terms, pricing, availability, behavior changes, routing complexity | Named OpenAI, Google, and Anthropic choices; legal grounding; Anthropic integration | Route shares, per-workflow cost, fallback performance, customer outcomes |
No row wins by default. High query volume and differentiated data can support more ownership. A varied workload and low internal research capacity can favor external models. A large vendor can do both, placing an owned model under one table while keeping frontier APIs beside it.
For buyers, architecture should become contract evidence. Ask which model handles the purchased workflow, whether it can change without notice, who owns evaluation results, what happens when a provider retires a version, and how answers are reproduced after a route change. A product label alone cannot answer those questions.
Engineering roles move with the model layer
Move inference inside and the work moves with it. Thomson Reuters needs people who can curate training data, run distributed training, build evaluations, serve models, investigate failures, and join research to product releases. A multi-model product needs people who can integrate providers, design routing, manage regressions, monitor cost, and maintain fallbacks. Both need security engineers, product managers, legal experts, and reviewers.
July brought two employment signals in opposite directions. Reuters reported that Thomson Reuters confirmed cuts to a small number of global engineering roles. An employee who attended a company meeting told Reuters the number could reach 500. Using the company’s 2025 workforce, Reuters calculated that figure as about 1.8% of 27,100 employees and 5.2% of 9,400 Operations and Technology employees.
Thomson Reuters did not confirm 500. It told Reuters that it expected to hire more than 250 net-new engineering roles over two years, most of them senior and AI-native. One number is an uncertain estimate of a current restructuring; the other is a future plan, not filled positions. Subtracting them would create a false net cohort.
No public evidence assigns every cut or hire to Thomson. The overlap in timing still illustrates an organizational fact: moving deeper into the model layer changes the mix of engineering work. Some application and platform roles remain essential. Other work shifts toward model systems, data, evaluation, research infrastructure, and the translation of professional judgment into tests.
Hundreds of Thomson Reuters subject-matter experts helped evaluate and shape Thomson. Harvey says legal experts produced training examples. LexisNexis says experts help develop and validate Protege. Their work is not an annotation footnote. It is part of the model’s production system.
Expert labor can disappear inside a benchmark number. A higher score may depend on lawyers writing rubrics, resolving ambiguous questions, finding authoritative sources, classifying errors, and deciding whether an answer is useful in practice. If that labor is unpaid, episodic, or borrowed from customer work, model economics look better than they are. If it becomes a stable role with feedback rights and career credit, it can improve both the model and professional practice.
Junior professionals face a related change. Thomson Reuters’ Future of Professionals Report 2026 surveyed 1,816 people across 62 countries in law, tax, audit, accounting, compliance, risk, and trade. It reported that 74% used AI several times a week, 41% lacked access to tools with professional or verified content, and 34% used unsanctioned tools.
Nearly half, 48%, expressed concern about developing independent judgment, and 71% said early-career roles need structured support from experienced peers. Those are self-reported expectations in vendor-sponsored research, not a labor census or causal proof. They identify a design problem that model strategy cannot solve by itself.
Corporate clients add another pressure. In the same survey, 78% said they wanted outside providers to deliver AI-supported quality, yet only 6% said most or all of their providers were doing so. A model investment can therefore arrive as both a service promise and a pricing dispute. A client may expect faster review, stronger citations, and a lower bill while the firm is still paying for implementation and senior supervision.
The commercial model changes which benefit survives. Under a fixed fee, lower accepted-task cost can improve the provider’s margin or support a lower bid. Under hourly billing, removing review hours can reduce revenue unless the firm wins more work or changes its price structure. A vendor benchmark cannot decide how that gain is shared among the software provider, law firm, corporate legal department, and end client.
If a domain model completes first-pass review more cheaply, a firm may remove the repetitive task through which a junior learned document structure. Keeping the task unchanged merely for training would waste time and client money. Removing it without a replacement can leave the firm with efficient outputs and a weaker promotion pipeline.
The replacement should use the model’s failure surface. Junior lawyers can inspect disputed citations, compare model routes, diagnose inconsistent table rows, rebuild an answer from sources, and explain escalation to a senior reviewer. That work teaches judgment more directly than copying text, but only if the firm budgets supervision and does not treat all review as invisible cleanup.
Model operations can also create new professional-technical roles. Evaluation counsel can own task sets and error taxonomies. Knowledge lawyers can maintain provenance and jurisdiction rules. Legal engineers can connect product behavior to matter systems. Model-risk leads can decide when a regression blocks release. These roles require status and decision rights, not a temporary committee assembled after an incident.
The hiring test is concrete. A vendor claiming model ownership should be able to say which capabilities moved inside, how many people maintain them, which work remains with partners, and who can stop deployment. A firm buying the product should know who handles exceptions after implementation. Without those answers, ownership is a capital announcement rather than an operating model.
A build, rent, or hybrid decision file
Most legal organizations should not begin by asking whether they can train a model. They should begin with a workflow that carries enough volume, cost, or differentiation to justify a model decision. Contract review, research, drafting, discovery, and internal knowledge search have different error costs and different opportunities for reuse.
A build decision also needs a demanding definition. Training from an open model, post-training an open-weight base, hosting a third-party model, grounding an API in proprietary content, and building an agent around several providers are distinct choices. Calling all of them proprietary AI hides who controls weights, data, infrastructure, and change.
Scale changes the answer without removing the decision. A 50-lawyer firm is unlikely to finance a $40 million training program. It can still own the held-out matters used for evaluation, its acceptance rules, a source register, correction data, route logs, and the right to export them. Those assets make a rented product testable and reduce the cost of switching.
A global information vendor may own weights and inference while depending on outside compute and an open-source base. A law department may own no model yet retain tight control over data, tools, approval, and fallback. Model ownership is one field in the operating file, not a proxy for control over the entire service.
The decision file below keeps those layers visible:
| Decision field | Build or own more of the model | Rent frontier APIs | Hybrid or routed portfolio | Evidence to record before approval |
|---|---|---|---|---|
| Workflow and consequence | Repeated, differentiated task with a stable shape and enough volume to amortize fixed work | Varied or changing tasks that benefit from broad capability | Several task families with different volume and risk | Named task, user, baseline time and quality, error consequence, monthly volume |
| Content and data rights | Organization can lawfully use a durable corpus and keep it fresh | Sensitive material can remain in retrieval or approved provider controls | Restricted corpora stay local while models vary by task | Source ownership, licenses, retention, jurisdiction, consent, deletion process |
| Model control | Weights, adaptation, serving, and release schedule can move inside | Provider controls weights, upgrades, retirement, and much of serving | Internal models cover selected workloads; external models remain available | Model name and version, operator, change notice, reproduction window, fallback |
| People | Model, data, evaluation, infrastructure, security, product, and domain expertise become recurring functions | Integration, evaluation, security, procurement, and domain review remain necessary | Both skill sets exist, with clear route ownership | Named accountable lead, staffing plan, expert time, escalation duty, stop authority |
| Fixed cost | Training, infrastructure setup, data preparation, hiring, and evaluation arrive before scale | Lower entry cost, implementation and evaluation still required | Selective fixed investment plus vendor setup | Budget by phase, internal labor, partner fees, opportunity cost, depreciation or amortization |
| Variable cost | Inference, hosting, observability, support, refresh, and idle capacity | Tokens, tools, storage, rate tiers, support, and price changes | Per-route costs plus routing and duplicate testing | Cost per attempted task and per accepted task, latency, reruns, review time |
| Quality proof | Owner must construct and maintain tests and invite outside challenge | Buyer must test the provider in its own tools and sources | Every route needs comparable tests and regression thresholds | Held-out tasks, valid-question review, human agreement, citation checks, error categories |
| Security and resilience | More control can reduce some third-party exposure but adds internal attack and operations surfaces | Provider assurance, regions, subprocessors, outages, and contract rights matter | More fallbacks can improve resilience while increasing interfaces | Threat model, data-flow map, incident owner, outage plan, recovery test, regional limits |
| Product integration | Model can be optimized for one tool and high-volume execution | Mature APIs and agent tools may shorten delivery | Router selects by task, cost, availability, or quality | Tool permissions, retrieval source, audit log, route rules, override and replay |
| Professional accountability | Internal team owns failure analysis and release decisions | Provider evidence supports but cannot replace professional review | Responsibility stays stable even when a route changes | Reviewer role, source access, exception queue, sign-off, client disclosure, appeal path |
| Exit and refresh | Retire or retrain when a better base or lower-cost route wins | Export prompts, evaluations, logs, and workflow assets before provider exit | Shift volume while retaining a bounded fallback | Renewal date, benchmark date, migration test, portable assets, retirement trigger |
| Outcome | Better accepted work, lower total cost, defensible control, or strategic differentiation | Faster access to broad capability without internal research scale | Portfolio improvement after routing overhead | Accepted-task rate, correction minutes, matter cycle time, user coverage, client or business result |
The cost row needs one denominator: accepted professional work. Cost per token rewards shorter prompts. Cost per query rewards more queries. Cost per completed model task can ignore legal review. Cost per accepted task includes failures, retries, tools, retrieval, human correction, infrastructure, and support.
Quality needs a similar denominator. A benchmark score should disclose the number and type of questions, excluded items, model settings, tools, sources, judge method, human calibration, and uncertainty. A vendor should also show the error mix. A wrong citation and a verbose answer may receive similar average penalties while creating very different professional risks.
Buyers can run a route-off. Give the same held-out tasks to the owned model and at least one credible external alternative. Keep retrieval, tool permissions, and review standards as comparable as possible. Measure accepted-task rate, unsupported propositions, citation survival, correction time, latency, and total cost. Let reviewers mark which route produced each answer only after they score it.
Volume decides whether small unit advantages matter. A five-cent saving on a monthly research question cannot finance a model team. A five-cent saving on hundreds of millions of structured document operations might. The model owner should disclose the workload denominator behind its economic claim without exposing customer-confidential volumes.
The file also needs dates. A model decision made in August 2026 can age quickly when open and closed models are converging and providers release new versions. Set a review date before signing. Re-run the held-out tasks after a major model change, a price change, a content migration, or an error that reaches a client.
Ownership should earn its renewal like any vendor. An internal model can lose on quality, cost, or resilience. A team should be allowed to route away from it without treating the decision as a failed strategy or a threat to the people who built it. The evaluation assets and workflow knowledge retain value even when the model changes.
Renting should face the same discipline. A frontier model’s reputation cannot substitute for the buyer’s task results. Broad capability may be worth a premium, but only the workflow record can show where. Contracts should preserve model notice, data boundaries, audit evidence, fallback, and the ability to leave.
Hybrid systems add coordination cost. A router can cut inference expense or improve quality, but it can also make an answer hard to reproduce. Record the model and version for each material output. Keep a route-independent evaluation set. Test the fallback under realistic load before an outage, not during one.
Approval should wait until finance, engineering, security, domain experts, product owners, and professional reviewers can see their responsibility in the same record. Finance needs unit economics, security needs the data path, reviewers need the error record, and procurement needs a tested exit. Leave out one of those views and the model decision is unfinished.
Tabular review supplies the first live test
Thomson’s first product placement offers a practical place to apply the file. Thomson Reuters launched the next CoCounsel Legal on August 20. The product uses Anthropic’s Claude Agent SDK, while Thomson is assigned to Tabular Analysis. The company says the surface can examine up to 10,000 documents against as many as 100 questions and link answers to sources.
Tabular work has a countable shape. Take a held-out set of contracts that never entered training. Define questions with lawyers before the run. Record whether every row has the correct source, whether an answer survives review, how long corrections take, and which errors require escalation. Compare Thomson with the prior production route under the same tool and content conditions.
A general counsel receiving the finished table will not see which model generated each cell. The provider should preserve that route in the audit record. If a disputed answer returns months later, the team needs the model and version, source set, question, tool actions, reviewer change, and final approval. Reproduction becomes harder when routing changes silently.
Then add the cost record. Include inference, retrieval, storage, failed runs, product support, reviewer time, and the fixed model allocation. Report cost per accepted row and per accepted table, not only tokens. A cheaper model that increases completion and citation quality should improve those measures. If it does not, the claimed model advantage is not reaching the workflow.
Track which errors junior reviewers can resolve and which require senior legal judgment. Then inspect the queue. Does it teach why the model failed, or merely deliver unexplained exceptions? Better throughput and stronger professional development can coexist, but the product will not produce both automatically.
Clients need a result they can negotiate, not only inspect. A pilot record should state whether savings reduce a fixed fee, expand the scope delivered for that fee, shorten a deadline, or remain with the provider to cover investment. Leaving benefit allocation unstated invites one side to cite the model-cost reduction while the other cites unchanged professional risk.
Raghu Ramanathan, president of Legal Professionals at Thomson Reuters, and Jennifer Eng, the unit’s chief product officer, now have a rare internal comparison. The same product contains an owned domain model and an external agent foundation. They can show where each route works, what it costs, and when the system falls back.
Harvey can publish the equivalent production record when Tenet moves beyond research preview. LexisNexis can disclose routing outcomes without revealing proprietary rules. Independent researchers and customers can repeat bounded tests. None needs to prove that one architecture wins every legal task.
On August 25, the visible split is real but limited. Thomson Reuters owns more of one model, Harvey controls post-trained variants for selected work, and LexisNexis keeps frontier APIs in its published stack. Their products remain hybrid systems built from models, sources, tools, partners, and professional review.
The first verdict should arrive in a table, not a launch claim: accepted answers, surviving citations, correction minutes, total cost, model route, and the person responsible for release. That record will show whether owning more of the model created better legal work or merely moved the invoice and the risk inside.