Productivity Rose. Labor's Share Reached a Record Low.
At 8:30 a.m. Eastern on August 6, the U.S. Bureau of Labor Statistics released a pair of numbers that should interrupt any one-sided AI productivity story. Nonfarm business productivity had increased at a 1.4% annual rate in the second quarter. Output rose 1.7%. Hours rose 0.3%.
Lower on the same Productivity and Costs release, real hourly compensation had fallen at a 3.1% annual rate during the quarter and was down 0.1% from a year earlier. Labor’s share of output stood at 52.9%, the lowest level in a series beginning in 1947.
Twenty-four hours later, another BLS release reached the hiring meeting. U.S. nonfarm payroll employment changed little in July, with a decline of 23,000. The unemployment rate held at 4.1%. May and June payroll growth was revised down by a combined 103,000. Average hourly earnings still rose 3.2% from a year earlier.
None of those figures identifies AI as the cause. The productivity report covers a quarter, the jobs report covers a month, and each uses its own surveys and definitions. The labor-share series describes the nonfarm business economy, not the experience of an employee using a copilot or a company buying an agent. Treating the releases as an AI impact study would turn timing into evidence.
Yet the timing matters inside a company. On July 17, OpenAI CFO Sarah Friar published a scorecard for enterprise AI. She proposed measuring useful work, the full cost of each successful task, dependability, and the value produced by each AI dollar at scale. Employee time, human review, retries, and rework belong in the cost.
Seats, tokens, and demo quality make a poor renewal conversation. Friar’s measures are better. They still stop before one line in the BLS release: who receives the value after output rises faster than hours?
Consider a quarterly operating review. Finance has workflow volumes and software spend. The AI team has pass rates. Operations has a backlog. HR has vacancies, pay bands, and a training request. Employees know whether saved minutes became quieter shifts, extra cases, fewer openings, or a new review burden. Everyone can bring a true number while nobody explains the allocation.
First verify that the system produced reliable work at an acceptable cost. Then follow the resulting time and value into output, price, margin, compensation, shorter hours, training, redeployment, or reduced hiring. Labor share is not a firm-level AI KPI. Its 52.9% reading is a warning against measuring only the production side.
Two releases before the hiring meeting
Friday’s Employment Situation gave a subdued national backdrop. Payroll employment fell by 23,000 after averaging gains of 34,000 a month during the previous year. Local government education lost 50,000 jobs, retail trade lost 19,000, and financial activities continued a decline with 14,000 fewer jobs. Health care added 22,000, slower than its average monthly gain over the prior 12 months.
Those industry lines resist one technology story. A school calendar, consumer demand, public budgets, interest rates, insurance conditions, demographics, and company-specific restructuring can all move employment. BLS made no AI attribution. Information, manufacturing, professional and business services, and several other major industries showed little change in July.
Among households, labor force participation was 61.4%, down 0.7 percentage point since January. The employment-population ratio was 58.9%, down half a point over the same period. Temporary layoffs rose by 153,000, while permanent job losers changed little. Household measures and employer payroll counts answer different questions; merging them would create a synthetic workforce number that neither survey reports.
Pay also requires its own clock. Private nonfarm average hourly earnings reached $37.62 in July and were 3.2% higher than a year earlier. That is nominal pay across payroll employees. The productivity release’s real hourly compensation adjusts for consumer prices, includes employer-paid benefits, and covers the nonfarm business sector. A nominal earnings increase and a decline in real compensation can coexist without either series being wrong.
Now move from the national report to a company planning room. A CFO may see weak payroll growth as pressure to approve tools that promise capacity without hiring. A CHRO may see a smaller external market but persistent shortages in selected roles. A line manager may have an open supervisor requisition and no spare reviewer for an AI pilot. Employees may hear the word productivity after two rounds of vacancy cancellation.
One macro release cannot settle those views. It can, however, make the company’s questions more exact.
Did the business produce more with the same team because demand rose, the work mix improved, equipment changed, employees gained experience, or AI handled part of a workflow? Which hours disappeared from the task, and which hours appeared in checking, escalation, or customer repair? Did the firm protect a vacancy, remove one, retrain someone, raise a target, or share the gain through pay or time?
Without those answers, an AI savings target is a forecast attached to a software invoice. The employment release supplies context, not proof.
August’s data also arrive with revision dates. BLS plans to publish a preliminary annual payroll benchmark revision on August 28. Revised second-quarter productivity is due September 3, followed by the August jobs report on September 4. A board that treats the first release as a permanent baseline may end up evaluating a multiyear contract against a number that changed before the next meeting.
Put three columns in the operating packet. Paid invoices, accepted work, actual hours, payroll, and compensation belong under current facts. Controlled comparisons and matched baselines belong under company evidence. Expected hiring avoidance, future task volume, and vendor savings remain forecasts.
Labor’s 52.9% line
Labor productivity is output per hour. BLS divides an index of real output by an index of hours worked by employees, proprietors, and unpaid family workers. It is possible for productivity to rise because output rises quickly, because hours fall, or through a combination. Second-quarter output grew faster than hours, producing the 1.4% annualized productivity increase.
Annualized does not mean output per hour was 1.4% higher over three months. It describes the pace that the quarterly change would imply if sustained for a year. The year-over-year productivity increase was 2.2%. A renewal deck should label the period beside every rate; otherwise a quarterly pace can be mistaken for an accomplished annual gain.
Unit labor costs provide another relationship. BLS calculates them as hourly compensation divided by labor productivity. They rose at a 1.3% annual rate in the second quarter as hourly compensation rose 2.7% and productivity rose 1.4%. A unit-cost increase is neither a worker raise nor a profit figure. It measures compensation required for a unit of output.
Labor share goes one step further. BLS defines it as the percentage of output accruing to workers as compensation. The measure includes wages, salaries, and employer contributions to benefit plans. At 52.9%, the share was lower than at any earlier point in the published series. It is a level for the nonfarm business sector, not a quarterly percent change.
No employee can read a personal pay result from 52.9%. Many forces shape aggregate output and compensation, and the preliminary release does not decompose the record low. Business composition, capital income, prices, benefits, hours, and sector mix can all move. Its value lies in the distribution question that a task benchmark leaves unanswered.
Local workflow accounting need not reproduce the national formula. It needs an equivalent of the missing line. When accepted output rises or hours per case fall, finance should be able to trace the gain into a small set of destinations.
A lower customer price, a wider margin, more volume without overtime, a cancelled vacancy, and less contractor spend are five possible destinations. So are a bonus, a higher pay band, paid training, a shorter schedule, or investment in another team. Several can apply at once.
Each choice changes the employment bargain. A support agent who closes the same number of cases with fewer repetitive steps has experienced work redesign. A support agent who receives a higher case target has experienced work intensification. A team that uses released time for escalations has changed its skill mix. Calling all three outcomes productivity hides the decision that matters to workers and managers.
Firm-level allocation data will never roll neatly into the 52.9% macro figure. The practical link is modest: both ask where output value goes. BLS keeps the distribution measure visible while companies learn to count tasks in extraordinary detail.
Useful work from one side of the invoice
Friar’s scorecard begins in the right place: define completed work in the system where the work happens. A support team can count resolved issues. Engineers can count changes that pass tests. A legal team can count accurate, on-time contract reviews. Tokens matter only when they become an outcome someone can use.
Her cost formula also corrects a common procurement mistake. Model price is only one input. The OpenAI essay says a business should include employee time, human review, retries, and rework, then divide full cost by the number of tasks that met the quality bar. A cheaper attempt can cost more per successful outcome when it generates extra checking or repeated runs.
Dependability prevents a fast but unreliable workflow from looking efficient. Value at scale asks whether completed work grows faster than total cost while quality holds or improves. Together, the four measures can expose a pilot that produces attractive demos but poor operating economics.
Authorship sets the boundary. OpenAI offers a vendor’s CFO framework, not an independent return study. Its essay explains how customers can judge AI work and compute economics; wages, employment, and economy-wide productivity sit outside the claim. Its final objective is more useful intelligence per dollar.
Applying the framework to an accounts-payable team shows the remaining gap. If the team processes invoices with fewer manual touches, useful work means invoices completed accurately and on time. Full cost captures licenses, integration, employee time, review, retries, and corrections. Dependability covers exception and error rates. Value at scale shows whether the cost per accepted invoice falls as volume grows.
At that point, the company knows the workflow worked. It still has several choices. It can process more invoices with the same people, reduce external service spend, leave vacancies unfilled, move employees into supplier analysis, reduce peak overtime, or return some gain through compensation and time. A strong task score does not select among them.
Allocation affects whether a result persists. Concentrate review work on a few experienced employees and their load may become the scale limit. Remove junior work and the team may lose the practice that develops future reviewers. Turn every released minute into a higher volume target and people may stop reporting hidden repair. Future dependability will look stronger on paper than it feels in the queue.
Ownership also changes the measurement. Finance can validate cost. Operations owns volumes and queues. Risk or quality teams know which outcomes failed. HR holds pay, vacancy, training, and role-movement data. Workers can identify shadow work that never enters the system. No single function sees the whole workflow.
Vendors supply product evidence; operations and risk certify task results. Finance validates cost against observed workload. Executives who control staffing and pay own the allocation decision.
Large firms expect a different payroll
Company evidence already warns against one national AI budget assumption. A March Atlanta Fed, Richmond Fed, and Duke working paper analyzed responses from nearly 750 senior financial executives, most of them CFOs. More than half had invested in AI, but adoption and expected effects varied sharply by company size and sector.
Large companies anticipated workforce reductions related to AI in 2026. Smaller firms expected modest employment gains. After weighting the survey by firm size and sector, the authors estimated an aggregate employment decline of less than 0.4% due to AI in 2026. Respondents expected the workforce share in routine clerical roles to fall over the following three years while skilled technical work gained share.
CFO perception ran ahead of revenue in the survey. Executives’ reported improvements in output per worker were larger than gains implied by revenue per worker. Quality, capacity, or future demand may take time to reach revenue; optimism, measurement error, and attribution may also widen the gap. Reported and expected company effects are neither a randomized deployment nor a macro causal estimate.
A May follow-up from Atlanta Fed researchers put dollars around the dispersion. In their Survey of Business Uncertainty analysis, employment-weighted AI spending per employee rose from $1,358 in 2025 to an expected $2,068 in 2026. Multiplying the latter by private nonfarm payroll employment produced a rough $280 billion aggregate estimate.
More than half expected to spend no more than $200 per employee, while the top 10% planned at least $2,800. The researchers winsorized extreme values, and the raw 2026 mean was higher. Professional and business services expected $3,470 per employee; manufacturing expected nearly $900. A single average describes almost nobody’s budget.
Executives forecast that AI would reduce next-12-month hiring demand by 0.8% for workers with college degrees and 1.1% for workers without them. The authors cautioned that these estimates exclude jobs at firms that do not yet exist and extra demand generated by productivity and real-income gains. They described the employment estimates as tilted toward the downside.
Firm size belongs in allocation policy. A large company may buy AI to standardize a process across thousands of jobs and remove future positions. A small firm may buy its first reliable capability and hire around new demand. The same software category can sit inside opposite headcount plans.
At a small company, the record can live in one sheet owned by a founder or operations lead. Compare a phased rollout with the last stable month, then treat a contractor renewal or a new hire as a lumpy decision rather than a fraction of a theoretical full-time role. A ten-person business cannot usually cash 15 scattered minutes per employee, even when the task arithmetic looks attractive.
Company leaders should publish the local plan instead of borrowing an aggregate prediction. Is the approved return based on more revenue, fewer vendor hours, lower error cost, vacancy avoidance, layoffs, or faster product delivery? Which job groups bear the change? Which new technical or supervisory work enters the budget? A worker cannot respond to a blended productivity percentage.
Finance should notice the gap between attributed productivity and measured revenue. Keep experiments open long enough to observe output, cost, and workforce effects. Do not book a forecast as savings.
Who enters the design room?
Kaiser Permanente and the Alliance of Health Care Unions put allocation authority on a bargaining table before they had an AI outcome to report. A July 13 MIT Sloan case study by Thomas Kochan, Erin Kelly, and Arrow Minster documents more than 150 hours of observed negotiations during 2025.
Scale made the process consequential. Kaiser provides insurance and care to 12.9 million members and employs more than 240,000 people. The Alliance represents roughly 62,000 Kaiser nurses, technicians, professionals, and service workers. In February 2025, about 300 Alliance leaders and 80 Kaiser leaders met before bargaining to discuss expectations and experiences with AI.
Their agreement covers worker participation from problem identification through design, implementation, education, training, evaluation, communication, and transition assistance. A national task force will have ten members, split equally between labor and management. Kaiser also has more than 3,500 unit-based teams where frontline workers, managers, and often clinicians meet around local service and work problems.
The difficult word was “inception.” Labor representatives wanted participation when a problem and possible solution were first defined. Management raised confidentiality concerns and pointed to decentralized technology decisions across Kaiser. The agreement created structures for early engagement without pretending that every local purchase could pass through one central committee.
This is process evidence from a health system with a 29-year labor-management partnership, not proof of AI productivity, patient benefit, staffing change, or gain-sharing. The national task force still needs to establish working relationships, include executives with technology investment authority, support local teams, and decide how ideas travel through a large organization. No deployment denominator or financial return appears in the case study.
Its contribution is practical. Frontline employees can observe repair work, patient friction, missing training, and unsafe shortcuts. Managers can observe exception queues, coaching hours, span, and targets. Neither group automatically controls the software budget, staffing plan, or pay decision. A useful review names evidence owners, decision owners, and affected groups separately.
That separation also protects the middle manager. Finance may book a saving while a supervisor inherits more reviews and escalations. Recording those hours, the queue they enter, and the manager’s authority to change staffing shows whether the new work fits inside the job. Without that evidence, “manager capacity” can disguise a workload increase.
Kaiser’s agreement offers one formal route for worker input. A nonunion employer can still gather paid, role-level evidence outside individual performance scoring and record management’s response before renewal. Employees will underreport repair and risk if every observation can lower their own rating.
Senior judgment without a practice path
Value allocation reaches pay and careers before it reaches a national labor-share series. PwC’s 2026 Global AI Jobs Barometer analyzed more than one billion job advertisements in 27 countries and territories. Jobs requiring specific AI skills had grown 69% while the overall jobs market grew 9%. The reported average wage premium for AI skills reached 62%.
A 62% modeled premium does not reach every employee asked to use an AI tool. It compares postings across occupations, industries, geographies, and skill requirements; it is not a 62% raise for the same worker. Citing it in a strategy deck while adding review work to an existing analyst’s job does not mean the market has paid that analyst.
PwC also found that, in a 2.4 million-posting U.S. sample, the most AI-exposed entry-level roles were seven times as likely as the least exposed to ask for traditionally senior skills such as leadership and strategic thinking. Openings in those “seniorised” entry roles grew 35% from 2019, while other entry roles declined 10%.
Causation remains unproven, but the comparison exposes a design problem. Once routine work is automated, a junior employee may be expected to judge exceptions before the job has supplied enough examples to build that judgment. Production time can fall while a shortage of reviewers grows three years later.
SHRM’s June automation and displacement study offers a counterweight to direct task-to-job arithmetic. Based on a survey of 14,245 U.S. workers mapped across 830 occupations, it estimated that 21% of wage and salary employment was at least half performed using AI tools. Yet 60.4% of employment had at least one nontechnical barrier to automation displacement, such as client preference.
Only 5.1% of wage and salary employment was estimated to be at least half automated with no nontechnical barrier, down from 6% in the previous estimate. That still represented about 7.9 million jobs. These are modeled exposure categories assembled from worker responses, occupational similarity, and BLS employment values. They are not observed job losses.
Organizational choices occupy the gap between task exposure and displacement risk. A customer may demand a person, a regulation may require accountable review, and physical work remains in many roles. Tacit knowledge can resist encoding. An apprenticeship task may deserve protection because it trains judgment, even when a system can complete it cheaply.
Workers also hold mixed expectations. Anthropic’s June Economic Index survey linked about 9,700 Claude-user responses to sampled usage. More than a third expected AI to be capable of doing most or nearly all of their work tasks within 12 months. More than a third said significant responsibility changes were likely or very likely, and 10% rated loss of their own job as likely or very likely.
Its respondent mix is far from the working population. Computer and mathematical workers made up roughly 30% of respondents against 4% of U.S. employment; managers were 23% against 7%. Benefits appeared beside anxiety: 68% said they learned more with AI, and 57% felt it made their skills more valuable. People closest to the tools can anticipate learning and job risk at once.
A credible productivity bargain makes room for both. It states which tasks leave a role, which judgment moves in, how paid practice will occur, how pay or level will change, and what happens if the role no longer contains a viable path. “Employees can focus on higher-value work” becomes credible only when higher-value work has an owner, training time, a pay band, and an opening.
From removed minutes to a budget change
The renewal review can fit on one page. It belongs beside the ROI calculation before a contract is signed, with evidence ownership separated from the authority to change targets, staffing, and pay.
| Line | Evidence for each workflow | Evidence owner | Decision owner | Affected group and decision |
|---|---|---|---|---|
| 1. Baseline | Accepted output, paid hours, overtime, compensation, vacancies, contractor spend, error and rework cost before rollout | Finance and operations | CFO and business leader | Team and customers: establish whether any gain occurred |
| 2. Full AI cost | Licenses, compute, integration, employee setup time, review, retries, incidents, rework, and vendor services | Finance and technology | CFO and CIO | Users and budget owner: calculate cost per accepted outcome |
| 3. Work result | Volume, quality bar, pass rate, latency, customer outcome, and exception rate | Operations and quality | Business leader | Customers and team: verify useful work and dependability |
| 4. Time conversion | Gross task minutes removed, added review and repair, net released capacity, schedulable blocks, and changed output or cost | Line manager and employees | Operations leader | Each shift and role: prove where time moved |
| 5. Manager capacity | Review, escalation, coaching, queue, span, target change, and authority to adjust staffing | Manager and operations | Business leader and HR | Managers and team: test whether oversight can scale |
| 6. Workforce and ladder | Hires made or avoided, vacancies, contractors, redeployments, exits, entry seats, practice tasks, paid practice, readiness standard, and next-level openings | HR and line manager | CFO, CHRO, and business leader | Job groups: reconcile the tool with headcount and career paths |
| 7. Role pay | Scope, decision rights, failure liability, skill scarcity, pay band, and effective date | Compensation and manager | CHRO and business leader | Employees in changed roles: evaluate the job regardless of savings |
| 8. Gain allocation and review | Realized value, employee-share numerator and denominator, recipient group, period, other destinations, data gaps, stop rule, and renewal date | Finance, HR, and analyst | Accountable executive | Workers, customers, and owners: decide who receives realized value |
“Time saved” has four stages: gross task time removed, net released capacity after new work, schedulable capacity, and either cashable saving or redeployed output. Skipping a stage turns fragments into fictional headcount.
Take a hypothetical 40-person customer-service team. During a quarter, it handles 12,000 cases that meet its quality bar. An AI workflow removes an average of eight minutes from each accepted case. Multiplication produces 1,600 hours of gross task time, nothing more.
Human review adds 240 hours, escalations 120, paid training 80, and service repair 60. Net released capacity is 1,100 hours under those assumptions. The team must then identify whose minutes disappeared, on which shifts and weeks, how the time was observed, and whether it forms usable blocks.
If only 680 hours can be grouped into schedulable blocks, the other 420 may still improve queue latency or create breathing room. They cannot support vacancy avoidance by themselves. A cash saving appears only when an overtime, contractor, payroll, or other cost line actually falls. Redeployed value appears when the same paid time produces more accepted work, better service, or a capability the company had budgeted to buy elsewhere.
Double counting becomes visible at this point. Finance cannot book vacancy avoidance while operations uses the same hours to increase case volume. Removing junior openings while expecting current agents to become expert reviewers also carries a future cost: fewer entry seats, fewer practice cases, and a thinner reviewer pipeline. Those effects belong in the same decision, even if they arrive after the renewal quarter.
Manager time needs its own denominator. A tool may remove routine handling while adding disputed outputs, coaching, and exceptions. Capture those hours with the escalation queue, span of control, workload target, and authority to change staffing. A manager who owns evidence but lacks authority to alter the plan should not be named as the person who “delivered” the saving.
Role pay and gain-sharing are separate decisions. Changed scope, decision rights, failure liability, or scarce skill can trigger job evaluation before a pilot produces net savings. Gain-sharing begins later, using realized value for a named period. Its calculation needs a numerator, denominator, eligible group, and treatment of integration cost; otherwise “employee share” is a sentiment rather than an amount.
Explicit allocation does not promise that every gain becomes pay. First dollars may cover integration, risk reduction, or a price cut. Capacity may fund market entry or a backlog. The review shows what changed, which destination management chose, and whether that destination appeared in an output or budget line.
Procurement can ask a vendor for accepted-work and full-cost evidence while keeping allocation authority inside the company. Freeze the workflow definition and quality bar. Require exportable results. A model upgrade should not reset the baseline each quarter.
The same method works for a finance close, code review, claims process, or factory inspection without pretending those workflows share one productivity rate. Small firms can keep it in one sheet; large firms can connect it to finance, HR, and operating systems. In either case, the final step is observable: AI spend changes work, work changes time and quality, and management chooses where the resulting value goes.
Back in the hiring room
BLS may revise second-quarter productivity and labor share on September 3. July payroll estimates can change as more employers report, and a preliminary annual payroll benchmark estimate is due August 28. Those updates belong in the context for the next meeting, not in the cell that claims what one company saved.
Return to the hypothetical service team. Its opening slide says 1,100 hours saved. The team lead corrects the label: 1,100 hours of net released capacity, of which 680 formed schedulable blocks. Finance confirms that 300 of those hours reduced an outside-service invoice. Operations used the remaining 380 on a measured backlog. No vacancy disappeared.
HR adds that one entry seat stayed open and the 80 training hours were already counted among the new work. Compensation schedules a review because experienced agents now own more exception judgment and failure liability. Employees’ role-level report shows the repair queue falling after a process change. The September renewal date remains on the page.
One executive now has a decision to sign: 300 cashable hours, 380 hours redeployed to accepted work, one entry route preserved, and a pay review with a date. If the BLS estimate changes in September, none of those company facts has to be rewritten.
The first slide offered a productivity claim. The signed page says who received the time and value.