On July 22, OpenAI published a number that can empty a timesheet.

NTT DATA Group had used five experienced engineers and three days to complete a complex incident analysis for a critical system. With Codex, the company said, the full process took 30 minutes. OpenAI presented the change as a 99.3% reduction in elapsed time.

The page did not describe the incident, the three-day workflow, the output that counted as complete or the review that followed. It did not say that five engineers worked continuously for all three days. It offered no external evaluation. This was a customer case published by the company selling the software, not a controlled industry benchmark.

Those missing details limit the claim, but they do not make the elapsed-time contrast irrelevant.

NTT DATA is a global provider of consulting, systems development and operations. Its work has historically scaled, in part, through the people assigned to client problems. OpenAI’s customer story says the company is trying to move from revenue growth tied to added headcount toward more value created with AI. A task compressed from days to minutes turns that ambition into a commercial problem.

If a client buys engineering time, faster work can reduce the invoice. If the client buys a fixed deliverable, the provider may keep much of the saving. If payment depends on an outcome, the two sides must agree on what caused the result and who carries the risk when the output fails. Each model gives the 30 minutes a different price.

The software does not remove every cost that sat behind the old delivery time. NTT DATA rolled out ChatGPT Enterprise, built an OpenAI Center of Excellence, distributed licenses, validated tools, wrote security guidance, ran events, monitored usage and trained employees. Experienced engineers still have to decide whether an incident analysis is sound enough to act on. A client may still expect a warranty, an explanation and a person who answers when the proposed fix creates another problem.

Yet the old unit has weakened. Three days once gave a buyer an observable quantity, even when that quantity was a poor proxy for value. The buyer could compare hours, rates and staffing levels. A 30-minute agent run removes much of that visible effort before the firm and client have agreed on a replacement measure.

This week, the Bank of England reported that some professional-services firms are already moving away from staff-time charges toward fixed fees, outcomes or value. Its business contacts also described lower demand for some graduate and junior roles, even as firms spent more on software, cloud services and specialist skills.

The two reports describe the same transfer from different sides. One records time disappearing inside a technology-services company. The other records firms beginning to change the price and workforce around that disappearance.

AI productivity has moved from internal efficiency dashboards into proposals, renewal calls, utilization targets, graduate hiring plans and the definition of acceptable work. The contract now needs a unit that can survive the change in delivery time.

July 22 compressed three days into 30 minutes

The NTT DATA case begins with a deployment much wider than one engineering team.

The company expanded Codex to about 9,000 active users across technical and nontechnical roles. It had already made ChatGPT Enterprise a primary internal tool and built a Center of Excellence to support adoption. According to the customer story, more than 96% of respondents to an internal survey were satisfied with ChatGPT Enterprise, and more than 95% reported productivity gains.

Those percentages are encouraging, but the page does not disclose how many employees answered, how respondents were selected or how productivity was measured. Satisfaction with a tool is also different from an audited business result. The distinction becomes important when a firm uses an internal success story to promise a client a lower price or faster completion.

The incident case is more concrete. It identifies an old elapsed time, a team size and a new elapsed time. It still leaves several questions open.

What work did the engineers perform during the three days? They may have collected logs, reconstructed a sequence, compared configurations, tested hypotheses, drafted a root-cause account or checked a remediation plan. The page does not say. It also does not state whether the Codex run covered every step, how many attempts were needed, which data the system could access or how much experienced review occurred before the analysis was accepted.

Without those details, the 99.3% figure measures a reported case, not a transferable production rate.

A buyer should resist two easy reactions. The first is to dismiss the case because it came from a vendor. The second is to multiply 99.3% across every engineering service. Both responses avoid the useful work.

A buyer can use the case to specify the denominator it needs.

For a comparable incident, a client would need the old cycle time, labor involved, severity, input volume, number of systems touched, review steps, rework, accuracy threshold and operational outcome. The AI-assisted case needs the same fields, plus model and tool cost, setup work, prompts or instructions, human interventions, failed runs and the evidence used to accept the result.

That comparison might show a dramatic saving. It might also show that the old three-day process included waiting for access, meetings or approval that the new demonstration excluded. A usable denominator makes that waiting visible and gives the buyer something specific to price.

NTT DATA’s own rollout shows why the tool cost cannot be reduced to the 30-minute run. The Center of Excellence created the conditions under which employees could use Codex. It handled technical validation, guidance, knowledge sharing and adoption support. Security rules defined usable data, connected systems, network controls, sandbox settings, automation limits and human review.

That operating layer is part of delivery cost even when it does not appear on a project timesheet.

The company also reported that weekly active Codex users increased to 1.4 times the prior level after it published a usage guide and conducted hands-on training. Adoption rose because people did work around the product. A pricing model that attributes the full saving to inference and treats training as overhead will understate what repeated delivery requires.

Hiroaki Sato of NTT DATA’s AI Technology Department said the incident result changed how people across the company thought about AI. The customer story described a division in which employees set direction and evaluate results while Codex carries defined work forward.

That division is commercially significant. Direction and evaluation become part of the product a client buys. If a firm removes hours spent producing an analysis, it may need to charge more explicitly for framing the problem, preparing the environment, reviewing the answer and accepting liability.

The experienced engineer moves from producing the analysis toward accepting it.

The Bank of England records a pricing shift

Two days after the NTT DATA case, the Bank of England published its July Agents’ summary of business conditions.

The report gathers intelligence from the central bank’s business contacts rather than presenting a statistically representative survey. Most of the July material was collected in the six weeks to the end of June. It describes subdued private-sector demand, pressure on margins, broadly flat employment intentions and more attention to operating efficiency.

Inside that wider picture, the Bank included a box on AI adoption, productivity, labor and inflation. Contacts said AI was automating routine tasks and accelerating knowledge work in software development, finance, administration, customer service, professional services and content creation.

Gains varied widely. Firms reported better results when skilled employees could validate and refine AI output. Implementation quality, training and the amount of human review affected what the tools produced. Some companies could raise output without matching employment growth. Others had not converted experiments into material operating change.

Costs moved to different parts of the service.

AI reduced the unit cost of some routine work and increased pressure to lower prices. At the same time, firms spent more on software, cloud services and AI licenses. Data-center demand raised costs around electricity, construction, hardware and specialized labor. Savings in one part of a service could be absorbed by a new bill elsewhere.

This is the commercial setting in which some professional-services contacts reported moving away from charges based on staff time toward fixed-fee, outcome-based or value-based pricing.

Each alternative creates a different contract problem.

Time-based billing lets a provider charge when scope expands, facts change or the client delays a decision. It also gives the provider revenue from inefficiency. A fixed fee rewards faster delivery, but the provider carries more scope and rework risk. Outcome pricing can connect the invoice to value when both sides define the result, observe it and agree on how much the provider influenced it.

An employment-law opinion, an audit, an advertising campaign and an incident analysis create different outcome problems. A client may want a lawsuit dismissed, a clean audit, sales growth or a stable system. The provider controls only part of each result. Courts, regulators, customers, market demand, client data and client action all intervene.

That is why “charge for value” is not a complete pricing method. Value must be bounded in a contract.

The Bank’s employment evidence adds another constraint. Some professional-services contacts reported reduced graduate recruitment and lower demand for administrative and junior staff. Routine document preparation, invoice processing and basic analysis had provided paid work as well as early-career practice. When AI compresses those tasks, the revenue unit and the training unit can disappear together.

Experienced judgment remains in demand, yet the routine cases that once developed it are thinning.

The report does not say how many contacts changed pricing or graduate hiring, which professions moved fastest or how permanent the decisions will be. Weak demand may explain part of the employment response. Firms also described targeted recruitment in IT, digital and specialist roles. The Bank’s evidence is directional intelligence, not a count of an industry conversion.

It is still unusually useful because pricing, technology cost and labor demand appear in the same official account. Many AI studies isolate productivity. Many workforce reports isolate jobs. The Bank recorded firms trying to reconcile both while clients resisted fee increases.

Keeping the old total requires evidence of better service or wider coverage. Applying a blanket discount ignores software, delivery risk and the human capability that makes the output acceptable. The fee has to account for both.

Clients want AI value before firms can measure it

Professional-services buyers are already asking firms to use AI. They are less clear about how the resulting value should be divided.

The 2026 AI in Professional Services Report from the Thomson Reuters Institute surveyed more than 1,500 people across legal, tax, accounting, risk, fraud and government work. Organization-wide use of generative AI rose from 22% in 2025 to 40% in 2026.

Measurement lagged. Only 18% of respondents said their organizations tracked AI return on investment. Another 40% did not know whether ROI was measured. Two-thirds of corporate respondents wanted outside firms to use AI, yet fewer than 20% required it.

That gap creates a difficult renewal conversation.

A general counsel may expect outside counsel to search, summarize and draft faster. The law firm may have bought professional AI tools, trained lawyers and created review procedures. Neither side may know how many hours a matter would have required without the tools. The client sees potential savings. The firm sees a new cost and a threat to revenue. Both can claim value without sharing a baseline.

Thomson Reuters published a second, broader Future of Professionals Report 2026 based on 1,816 professionals across 62 countries. It found that 74% used AI several times a week. Among corporate clients, 78% said AI-enabled quality improvements from outside firms were very important or essential. Only 6% said most or all of their providers delivered those improvements.

Thirty-two percent expected to reconsider relationships with firms that fell behind during the next 12 months. Among that group, one-third estimated that more than $1 million in annual work was at risk.

The reports come from a company that sells AI products to professionals. Its definition of professional-grade AI and its framing of client demand support that business. The survey populations, field dates and published methods make the findings more useful than a sales claim, but buyers should still read the commercial context.

Steve Hasker, Thomson Reuters’ chief executive, offers the strongest countercase to a simple revenue decline. He expects AI to increase the volume of professional work by accelerating business formation, transactions, restructurings, litigation and contracting. In that scenario, hours per task fall while total demand rises. A firm can complete more matters, serve clients that could not afford the old process or sell monitoring that was uneconomic when every check required manual work.

The countercase changes the timing, not the need for a pricing decision. A client still wants to know why a repeatable task carries last year’s fee. A provider still has to decide whether increased capacity becomes a lower price, more coverage or a larger book of work.

The pressure is uneven. Clients resist financing avoidable manual work, but a discount loses its appeal when it arrives with unreviewed output or hidden tool use. Firms want to keep the margin from faster delivery while still paying for expertise, data protection, training and correction.

The negotiation should begin with disclosure rather than an automatic discount.

For each material work type, the firm can state where AI is used, which systems or data it touches, what a professional reviews and which output the firm stands behind. The client can state whether AI use is permitted, required or restricted. Together they can choose a price unit that matches the work.

An hourly engagement can bill only human time and state whether AI fees appear separately. A fixed-price proposal identifies the assumptions that keep the fee valid. Outcome arrangements require a baseline, measurement window and treatment of external factors. A subscription sets capacity, response time and exclusions.

The proposal should state how the verified saving will be divided.

Suppose a repeatable review once cost $20,000 and a new process can reliably deliver the accepted output for $8,000 in labor, tools and quality control. A $14,000 fee would divide the saving. Keeping the old price would require more in return, perhaps faster turnaround, wider coverage or a warranty. For continuous demand, the same economics could support a subscription with more reviews inside the annual spend.

Those are commercial choices. None can be inferred from a model benchmark.

The strongest firms will bring evidence before procurement demands a blanket reduction. They will know which tasks became faster, which did not, where error and rework moved, how clients benefited and what new capabilities the margin funded.

An unsupported claim of “AI efficiency” invites the buyer to set the price.

Picture a panel renewal for a professional-services firm. Procurement has last year’s invoices and can see 1,200 hours billed for a recurring workstream. The practice leader proposes a fixed annual fee and says AI will improve speed and consistency. Finance asks how much cost has actually fallen. The client’s business owner wants faster answers, not a smaller team on paper. The firm’s talent leader knows those 1,200 hours included much of the work assigned to new professionals.

The negotiation will go badly if the firm brings only a license list and a productivity percentage. It needs task-level evidence, an acceptance standard and a decision about the early-career work that will be removed.

Fixed fees move risk back to the provider

The billable hour survived for reasons that extend beyond tradition. It prices uncertainty.

A professional may not know how many documents will arrive, how many systems are involved, whether the client has clean data or how many revisions a regulator will require. Time-and-materials billing lets the price expand with the work. The client carries much of the scope risk.

AI can make a defined task faster without making the surrounding engagement predictable. A model may review the first thousand documents quickly and struggle with the unusual files that matter most. An agent may produce an incident hypothesis in minutes while access, testing and approval still take days. A research draft may be fast, but a professional may spend longer checking citations after finding one fabricated source.

Fixed fees reverse part of that risk. The provider keeps the margin when delivery goes well and absorbs the cost when scope, quality or tool behavior moves against the estimate.

The Institute of Practitioners in Advertising made that trade visible in its 2026 Pricing Playbook. The guidance separates input-based, output-based, outcome-based and hybrid models. It asks agencies to examine risk appetite, resource flexibility, scope stability, commercial fluency and outcome measurability before choosing a model.

Those tests travel beyond advertising.

Pricing modelBest fitWho benefits first when time fallsRisk that movesWorkforce pressure
Time and materialsScope is uncertain and professional effort is observableClient pays less if billed hours fallProvider revenue falls with efficiency; client still carries scopeUtilization targets can discourage automation
Fixed feeOutput is repeatable and acceptance is clearProvider keeps delivery margin; client gains predictabilityProvider absorbs scope, rework and tool-cost varianceFewer production hours can reduce junior staffing
Outcome or value basedResult is measurable and the provider can influence itSaving can be shared around agreed valueAttribution, timing and liability become centralSenior judgment and commercial skill gain weight
Subscription or managed serviceNeed is continuous and volume can be boundedClient gains access and provider can reuse systemsVolume spikes, response commitments and renewal valueStable teams are possible if capacity is funded
HybridScope or outcome remains partly uncertainParties divide efficiency and uncertaintyContract complexity risesTraining and review can be priced explicitly

The appropriate model changes with the task.

An hourly arrangement may remain appropriate for a novel investigation with changing facts. A fixed fee can work for a repeatable filing or defined migration. An outcome fee may fit a recovery project when the baseline and collected value are observable. A subscription can support continuous monitoring or advice. A hybrid can combine a base fee with usage, milestones or a success payment.

The AI cost beneath each model is also unstable.

Providers can pay by seat, token, tool call, connected application, completed case or enterprise commitment. A reasoning-intensive task may consume more compute than expected. A model upgrade can change quality and price during a contract. A client may require a private environment, specific retention settings or a second model for validation. Human review may rise as the task becomes more consequential.

Passing every tool cost through to the client recreates time billing in another unit. Hiding the cost inside a fixed fee can destroy margin when usage changes. A workable proposal identifies the material cost drivers and sets a review point without turning the invoice into a model telemetry dump.

Liability determines how far outcome pricing can go.

A marketing agency can share upside from an attributable campaign metric. A law firm cannot guarantee a court decision. An auditor cannot sell a desired opinion. An incident-response provider may commit to response and analysis standards, but the client’s architecture and actions affect recovery. The price can move closer to the result without pretending the provider controls the world around it.

Jason Cobbold, who co-wrote the IPA playbook, argued that agencies return to familiar models because they feel safe even as the work changes. Ed Palmer of the IPA described firms blending models around client risk, technology and value.

Firms can blend models when the risks differ, as long as the contract makes the split explicit.

Graduate hiring falls before training is rebuilt

The first workers affected by a new pricing model may be the people who have not yet learned to price their own judgment.

The Bank of England heard that some professional-services firms had reduced graduate recruitment and demand for junior or administrative staff. Thomson Reuters found a related concern among people already in the professions: 48% feared that AI would harm the development of independent judgment.

Legal respondents expected the path to trusted judgment to lengthen by nearly two years. Tax respondents expected it to shorten by one. The effect on learning depends on the work, feedback and responsibility that remain.

Professional training has often been financed through client work.

A junior lawyer may begin with case research and a first draft, then watch a senior lawyer revise both the reasoning and the language. In accounting, reconciliation work exposes the exceptions that a clean training example omits. Analysts learn by defending assumptions in review. On an engineering team, tracing an incident shows which plausible hypothesis survives contact with production.

Some of that work is repetitive. It is also where people see many examples, make lower-stakes mistakes and learn the difference between a plausible output and a defensible one.

Hourly billing made the arrangement commercially awkward but legible. The firm could bill some junior time at a lower rate. The client paid for production and, indirectly, part of the profession’s training system. Senior review appeared as another line.

When AI removes production hours, a firm can celebrate the efficiency before deciding who funds the apprenticeship.

A fixed fee can create room to solve the problem. If the firm delivers faster and keeps part of the margin, it can reserve time for supervised practice. A subscription can support a stable team across matters. An outcome fee can reward experienced judgment while the firm treats training as an investment, provided leaders allocate the saving that way.

The margin can flow to partners, shareholders, software vendors or a lower client price. Utilization targets can remain even after the available billable tasks shrink. Juniors can receive polished AI drafts without seeing the wrong turns that produced them. Seniors can inherit more review work and have less time to teach.

The NTT DATA example offers a more constructive path. The company paired tools with guides, hands-on training, employee communities and a Center of Excellence. It encouraged internal teams to identify reusable use cases. That approach can broaden access to technical work, including for nontechnical employees.

The case does not show whether graduate hiring changed or how early-career engineers learned incident analysis after automation. Those are open questions, not grounds to assume harm or benefit.

Professional firms need a learning reserve beside the productivity split.

For each work type that loses junior hours, leaders can identify the judgment that the task once developed and create another route to practice it. A junior may compare the agent’s hypotheses with raw evidence, test a weak output, conduct the client interview, explain an exception or write the acceptance memo. The exercise should end with feedback from someone accountable for the work.

The reserve needs time, an owner and a budget. “Learn by using AI” is too vague. A firm should be able to state how many supervised cases a new professional completes, which decisions they can make at each level and what evidence supports promotion.

Paul Griggs, PwC’s U.S. senior partner, described his firm’s AI work in March as retraining people, embedding AI in methodologies and adding engineering capability. He also wrote that investment arrives before results and that progress moves team by team.

A firm can also lose experienced people when its tools and work design lag. In the Thomson Reuters survey, 24% of professionals experiencing an AI value gap said they were considering leaving within two years. The report estimated a $232,000 replacement cost per professional. Those are survey-based estimates rather than a forecast for every firm, but they put another cost beside the software decision.

Workers can be squeezed from opposite directions. A junior may lose the assignments that built fluency. A mid-career professional may leave because the firm provides weak tools and expects manual throughput. The training plan has to protect practice without trapping people in avoidable production.

Removing entry work before building the next training system may improve this year’s margin and weaken the future supply of reviewers. Preserving every old task has the opposite problem: juniors practice work the market no longer values. Firms can remove avoidable production and direct part of the saving to supervised judgment practice.

A contract for compressed professional work

AI capability will remain difficult to forecast. The contract still has to remain intelligible when delivery time changes.

The compressed-work contract matrix below is a pre-proposal artifact. The provider completes it for each material work type, and the client challenges the assumptions before selecting the price model.

Contract fieldDecision to recordEvidence before signaturePrice consequenceWorkforce consequence
Work unitTask, deliverable, service period or business outcomeCurrent statement of work and comparable mattersDetermines hourly, fixed, subscription or outcome baseDefines which team performs the work
BaselineOld cycle time, labor mix, quality and costComparable completed cases, with waiting time separatedSets the reference for any claimed savingShows which hours may disappear
Accepted outputWhat counts as complete and usableAcceptance test, required explanation and client sign-offLimits scope and supports a fixed feeNames the professional who can accept work
AI use and evidenceModels, tools, data access and retained recordsApproved architecture, usage policy and evidence sampleSupports tool charge or included-cost assumptionDefines skills and permissions needed
Human reviewDecisions that require named professional judgmentReview checklist, escalation rule and authorityPreserves a review component in the feeProtects reviewer capacity
Rework and warrantyWho pays when output fails or facts changeError classes, service levels and remedyPrices quality and scope riskDetermines incident and quality staffing
Productivity splitHow client and provider share verified savingsBaseline comparison and benefit measureSets discount, added coverage or provider marginDecides where released capacity goes
Learning reservePractice removed and replacement trainingCareer framework, supervised cases and feedback planFunds training inside margin or as a separate feeMaintains the early-career pipeline
Review dateTrigger for changing assumptions or priceVolume, quality, tool-cost and outcome thresholdsPrevents a temporary benchmark becoming permanentAllows hiring and team design to adjust

Start with the NTT DATA incident case.

The public evidence supplies an elapsed-time comparison and team size. A client contract would need more. It would define “incident analysis,” identify the inputs, state which findings and remediation elements are required, specify who reviews the analysis and record what happens if the result is wrong.

The baseline should separate active expert work from waiting. If the old three days included a 20-hour access delay, the agent did not remove that time unless the new process changed access as well. If five engineers joined at different stages, team size should not be treated as five full-time allocations.

The accepted output needs more than speed. It may require a reproducible sequence, evidence for each finding, rejected hypotheses, affected systems, confidence limits and a remediation plan. Those are examples of fields a buyer could require, not details that OpenAI disclosed about NTT DATA’s case.

Human review should name authority. A senior engineer may approve a low-risk internal analysis. A critical production change may require a service owner, security lead and client approver. The contract can state which decisions an agent may propose and which a person must make.

The productivity split then becomes a negotiation with evidence.

If repeated comparable incidents show that the provider can deliver accepted analyses faster and at lower cost, the client can receive a lower fixed fee, faster response, broader coverage or a combination. The provider can retain enough margin to pay for tools, the Center of Excellence, review and continued improvement.

The learning reserve keeps the commercial model from silently deleting the career model. It can require supervised incident reviews for junior engineers, rotations through root-cause work or a minimum training allocation funded by delivery margin. A client may care because a provider with no talent pipeline becomes a continuity risk.

The review date protects both sides from a single demonstration.

A 30-minute case should not set a permanent price for every incident. The parties can review after a defined volume, when quality stabilizes, when the model changes materially or when tool cost crosses a threshold. If performance improves, the benefit split can move. If rework rises, the price and process can be corrected before the service fails.

The matrix narrows disagreement from an abstract claim about AI to observable delivery choices. It also gives procurement a record of the assumptions behind the price.

Thirty minutes still ends with a named reviewer

Imagine the next incident after the case study becomes part of a client proposal.

At 9:00 a.m., the provider opens the investigation. Codex receives approved evidence and completes an analysis by 9:30. An experienced engineer reviews the sequence, rejects one weak hypothesis, checks the proposed remediation and signs the accepted report at 10:20. The client applies a change at 11:00 and restores stable service.

The old contract would ask how many people worked for how many hours. A better one records whether the analysis met the agreed test, which evidence supported it, who rejected the weak hypothesis and what the provider promised if the result failed. It also states how the verified saving is divided.

Those answers determine the invoice. The provider might bill the engineer’s review, charge a fixed incident-analysis fee or include the work in a managed response subscription. Value tied to avoided downtime is possible when the contract also recognizes the client’s role in recovery.

NTT DATA’s reported result makes the mismatch visible. The production clock can change faster than pricing, utilization and career systems. Waiting leaves clients, software costs and talent attrition to choose the replacement. Evidence gives the provider a chance to choose it instead.

At the next renewal, the most important line may no longer record days or hours. It may record the accepted output, the named reviewer and the result the client is willing to pay for.

The Codex run can stop after 30 minutes. Someone still has to sign below it.