The Labor Department Turns to AI Vendors for Jobs Data
At a Forum Club event in Palm Beach on August 25, acting Labor Secretary Keith Sonderling described a new source for a question Washington has struggled to answer. The Labor Department, he said, had signed data-sharing memorandums with technology companies including OpenAI, Google, Meta, and Amazon. Their information would help the government track how artificial intelligence changes jobs and hiring. Findings would be made public.
Axios reported the arrangement the next day. The memorandums were not attached. No public document named the fields each company would supply or the people and businesses covered. Nor did it specify a time window, matching method, or date for the first release. What became public was an institutional move and a list of companies, not a production dataset.
One day later, another federal office placed AI inside the employment system it wants to measure. On August 27, the Office of Personnel Management told agencies to use artificial intelligence to help reach an average time-to-hire target of 80 days. The OPM memorandum allows uses such as drafting job content, screening resumes against source documents, checking a pre-offer record, and evaluating aggregate results. It also requires independent official review, quality assurance, privacy, accessibility, auditability, and a route for reconsideration.
Put together, the two actions create a practical collision. Private companies can see product use hours or months before a government survey is released. Employers can see vendor spending, job postings, and payroll movement. Federal agencies can add AI to a hiring workflow and observe its output. None of those records means the same thing as an employed person, a vacancy, a task, a wage, or a job lost because of AI.
ChatGPT can record a message while no job changes. A corporate card can show a purchase without revealing deployment depth. A posting can close without a hire, a public profile can lag payroll, and an exposure score can describe susceptibility without proving a layoff. Faster data narrows a blind spot only when those boundaries stay visible.
Here lies the value of the Labor Department’s move. Public statisticians could compare early signals with slower, representative measures. They could trace a sequence from product use to company spending, a hiring decision, and an eventual payroll record. Without a public method, the same memorandums could give vendors a federal stage for product-selected evidence.
For a CHRO or workforce planner, this is not a remote statistical dispute. AI budgets are being approved while finance, recruiting, and operations use incompatible measures of adoption. Managers hear that firms spending heavily on AI are hiring more. Early-career workers hear that employment in exposed occupations is falling behind. Both findings can be true. A public account must explain how.
A dramatic headline would be a poor first test. A reader should instead be able to reconstruct what each number observed, who was absent, what changed, and who retained the right to revise or reject the finding.
Fifty survey answers can move the unemployment rate
Pressure behind the vendor agreements starts with a household survey.
Every month, the Current Population Survey supplies the unemployment rate and a detailed picture of who is working, looking for work, or outside the labor force. Its strength comes from a probability sample, tested questions, documented weights, and methods that allow the answers of participating households to represent a much larger population. It is a public denominator rather than a count of people who happened to use one product.
Its collection system is under strain. In an April 2026 report to Congress, the Bureau of Labor Statistics said the survey response rate had fallen from the low 90% range to the upper 60% range over roughly 15 years. One participating person now represents about 3,500 people in the civilian noninstitutional population. A shift of fewer than 50 net responses between employed and unemployed can move the headline unemployment rate by 0.1 percentage point.
Fifty people do not arbitrarily determine the economy. The estimate uses a much larger sample, population controls, weighting, and seasonal adjustment. A lower response rate also does not prove that the published result is biased. Risk appears when the people missing from the sample differ in relevant ways from those who answer and the statistical adjustment cannot recover that difference.
Still, the figures explain the appeal of a supplementary signal. A lost response from a household with a young contract worker or someone moving between jobs may remove information that a product or payroll system sees quickly. Survey collection is expensive as well. BLS plans to add web self-response to the CPS in 2027, alongside further questionnaire and systems work. Modernization may improve participation, but a monthly survey will never become a live log of every AI-assisted task.
Private data can fill a timing or detail gap. It cannot supply the public denominator by default.
OpenAI observes people who use ChatGPT, not people who do not. Google, Meta, and Amazon each operate a different mix of consumer services, cloud products, advertising systems, productivity tools, marketplaces, and enterprise accounts. A technology company may know that use rose among its customers. It may have no idea whether the same work moved from a competitor, an internal model, a spreadsheet, or a person. Commercial coverage changes when prices, bundles, product interfaces, or sales campaigns change.
Picture the analyst asked to reconcile those feeds. One employer buys an enterprise account at headquarters and assigns seats across three subsidiaries. A consultant uses a personal account while working for a client. Ten people at a shop share one login, while another employee switches between two products. Account growth, user growth, and employer adoption separate before any record reaches a jobs table. The matching file needs a rule for each case and an unmatched category that remains visible.
Official survey coverage changes too, but its changes are documented and designed for statistical continuity. A public-private dataset needs the same discipline. A free tier, model rename, country restriction, or large contract can alter the vendor series. Government analysts must separate that product event from a labor-market change.
BLS already operates under confidentiality and scientific-integrity rules that set a useful floor. Its confidentiality pledge limits identifiable data acquired for statistical purposes. Its CIPSEA implementation report describes protections for information collected under a statistical pledge. Its scientific-integrity policy protects professional control over methods, scope, and release.
Existing policies do not tell us which office will hold the new vendor data or which legal authority governs each transfer. The memorandums may use aggregated reports, clean rooms, restricted microdata, or some other arrangement. Until the documents appear, the architecture remains unknown. One boundary is already clear: vendors may contribute observations, while public statisticians retain control of definitions, weighting, validation, revision, and publication.
CPS methodology does something no product dashboard can do alone: it asks who is missing. Judge a new feed by how much it helps answer that question, not by how many rows arrive overnight.
One adoption rate can be 18 percent or 78 percent
Artificial intelligence already demonstrates the denominator problem before the new agreements begin.
An April Federal Reserve research note compared three prominent surveys. One found that about 18% of firms used AI. Another suggested that firms employing roughly 78% of the labor force reported some adoption. A third found that about 41% of individuals used generative AI for work. None of the figures has to be wrong.
They answer different questions.
| Reported signal | Observation unit | Denominator | A defensible reading |
|---|---|---|---|
| About 18% | Business | Firms responding to a business survey | Share of firms reporting AI in at least one business function |
| About 78% | Employer, weighted by employment | Workers at firms whose respondent reports adoption | Share of the labor force attached to a reported-adopter firm |
| About 41% | Individual | Surveyed adults or workers within the study design | Share of people reporting work-related generative AI use |
A large employer and a two-person shop each count as one firm in the first row. The large employer accounts for many more people in the second. A company can report that it uses AI even if only one team does. A person can use a public chatbot at work even if the employer has not purchased an enterprise tool. The word “adoption” travels across all three rows while the thing counted changes.
Census makes the firm measure more useful by asking about business functions. Its 2026 working paper on AI diffusion found that about 18% of firms reported AI use from November 2025 through January 2026. Employment weighting lifted the estimate to 32%. Among adopters, 57% used AI in three or fewer functions. Sixty-six percent reported augmentation without automation. Two percent reported an AI-related employment decrease.
Those details spoil two easy stories. Eighteen percent does not mean AI touched only 18% of workers, because adoption is concentrated in larger firms. Nor does every employee at an adopter necessarily change work. A company with an AI marketing tool, a coding assistant, and no other deployment may count as an adopter while most staff never see the product.
That 2% employment-decrease figure comes from firm reports, not an audited causal estimate of displaced workers. A respondent may observe fewer positions while demand, financing, offshoring, automation, attrition, and ordinary restructuring move together. A company can also use AI, grow payroll, and eliminate a particular task or junior opening. Net firm employment hides composition.
Question wording matters. The Census survey revised its AI question in November 2025 to ask about use in any business function. A methodological change can move the series even if company behavior holds steady. A vendor series changes for similar reasons when its product taxonomy, account mapping, pricing, or event logging changes. Both need version labels.
Consider a regional retailer deciding whether to train 4,000 store employees. The 18% firm rate might make the project look early. The employment-weighted rate might make delay look reckless. The worker-use rate might imply staff already use tools outside approved systems. A sound plan does not choose the largest percentage. It identifies which population resembles the retailer, then measures use by function, role, location, frequency, and accepted outcome.
Publishing a denominator map beside every combined finding would improve the public record. Readers should see firms, workers, accounts, messages, tasks, occupations, vacancies, hires, and payroll records as separate units. When a chart combines them, the bridge between units belongs in the evidence rather than inside an ambiguous label.
Message logs see tasks before payroll sees jobs
Vendor telemetry is strongest when the question is close to the product event.
OpenAI’s July 27 Work at the Frontier study analyzed more than 800,000 work-related messages from U.S. users. It classified 16.8% of all work messages as crossing a traditional occupational boundary. Among nongeneric messages tied to occupation-specific work, the share was 43.5%. Human-resources users had a 69% crossover rate for occupation-specific messages.
This study catches work change that payroll may miss. An HR specialist can ask for code. A designer can request market analysis, while a founder drafts a policy without hiring a specialist for that single task. The message appears the moment the attempt occurs. Payroll may never change.
Telemetry stops well before an employment outcome. A message cannot show whether the answer was correct, accepted, copied, rewritten, or discarded. It misses the minutes saved and the review minutes added. Nor can it reveal whether the person gained responsibility, avoided a hire, earned more, or lost work. The dataset covers users of one product, and OpenAI has a commercial interest in evidence that its product expands capability.
Commercial interest does not spoil the data. It raises the standard for contribution and validation.
A Labor Department report using ChatGPT data should define its eligible message universe and explain how it identified work-related content. Readers also need to know how occupation was inferred, how enterprise and consumer accounts were treated, and which privacy transformations preceded the analysis. The report should disclose whether changing model behavior or product features could change the classification. Independent researchers need enough aggregated evidence or controlled access to reproduce the central result without receiving private conversation text.
Other named companies present different possibilities and risks. Google may observe enterprise productivity activity, search behavior, cloud model calls, or developer use. Meta may see business communication, advertising, creator activity, or open-model adoption. Amazon may see cloud consumption, marketplace activity, employer operations, or product purchases. The announcement left every category unspecified.
Do not fill that gap with a unified map of American work. The companies may supply four narrow extracts that cannot be joined. One may report accounts, another tokens, another businesses, and another aggregate trend. A useful release would name each contribution instead of hiding them behind “technology company data.”
Privacy is part of measurement quality. Workers may discuss health, performance, compensation, disputes, clients, or protected activity in an AI product. A program built to study jobs does not need the government to read those conversations. Data minimization can favor counts, classifications, uncertainty measures, and protected linkage keys over raw text. Retention and access should be bounded before transfer, with disclosure rules that prevent a small employer or rare occupation from becoming identifiable in a table.
For a worker, “aggregate” can still feel personal. A table that names one occupation in a rural county may describe only a handful of people. An unusual task combined with an employer’s public announcement may narrow the group further. Minimum cell sizes, suppression, and review for linked disclosures matter even when names and email addresses never leave the vendor.
An employee also needs protection from a second use. Data supplied for statistical research should not quietly become evidence for enforcement against the person, performance monitoring by the employer, or training by another vendor. The legal promise and the technical architecture should agree.
Message logs contribute speed and task detail. They do not confer authority to decide what happened to employment. Their strongest public role is as an early signal that can be checked against surveys, postings, employer records, and payroll over time.
Spending and payroll produce different verdicts
Two private datasets now support labor conclusions that can sound contradictory.
Ramp and Revelio Labs linked corporate spending with workforce profiles across more than 21,000 U.S. firms. Their June 30 analysis of AI investment and hiring defined an adopter as a company spending at least $100 on an AI vendor for three consecutive months. High-intensity adopters ended a 24-month window with employment roughly 10% above comparable firms that had not yet adopted. Their entry-level share was 1.15 percentage points higher.
This result challenges the claim that AI purchasing automatically shrinks employment. Buying AI still does not prove it caused the hiring. Firms able to spend heavily on new software may already have faster growth, better financing, more engineering work, or stronger demand. A $100 threshold detects a payment relationship, not deployment depth. Corporate cards and bill-pay systems miss some contracts, internal models, cloud discounts, and employee-paid tools. Revelio’s public-profile workforce data also differs from an employer’s payroll register.
Stanford’s Digital Economy Lab and ADP start closer to the worker. Their revised August 12 Canaries in the Coal Mine analysis uses millions of U.S. payroll records. It did not find economy-wide employment displacement. Within highly AI-exposed occupations, however, employment for workers aged 22 to 25 was about 19% below the path of less-exposed peers in the authors’ model.
Nineteen percent is a relative gap, not a claim that AI fired 19% of young workers. The estimate depends on occupation exposure, comparison groups, the payroll population, and the model used to separate age and occupation patterns. Demand, interest rates, industry mix, education, and employer behavior may move with AI exposure. The analysis offers a warning about early-career employment, not an individual causal record.
A 24-year-old applicant cannot experience a modelled relative gap. That person experiences three cancelled interviews, a junior role rewritten for two years of AI experience, or a long search that ends outside the studied occupation. Payroll can register the eventual employer and earnings. It does not capture the applications, changed expectations, unpaid training, or abandoned career route that came first. Linking aggregate payroll trends to posting and applicant data can illuminate the path, but the worker’s account is still evidence rather than decoration.
Both findings can appear inside one company. A fast-growing software firm might buy more AI, increase total payroll, hire experienced sellers, and reduce junior analyst openings. A large employer can expand while an exposed occupation contracts. Entry-level share can rise across selected adopter firms even as young workers lag within particular occupations across a broader payroll base. Firm growth and worker distribution are different outcomes.
| Dataset | Earliest event it sees | Outcome it can support | Outcome it cannot establish alone |
|---|---|---|---|
| Vendor telemetry | Account, message, feature, or model use | Product activity and task pattern | Job creation, displacement, wage change, or accepted value |
| Corporate spend | Payment to a classified AI vendor | Purchased adoption signal and spending intensity | Actual use, internal tools, worker coverage, or causal hiring effect |
| Job postings | Advertised vacancy and requested skill | Employer demand language and intended recruiting | Filled job, final wage, worker tenure, or net employment |
| Public profiles | Observed employment affiliation | Approximate firm and role movement | Complete payroll, hours, hidden workers, or exact separation reason |
| Payroll | Paid worker, earnings, and employer relationship | Employment movement for covered workers | Every task, product used, unpaid work, or causal motive |
| Representative survey | Reported person or business behavior | Population estimate within the survey design | Live product detail or automatic causal attribution |
A small manufacturer will not see itself equally in every row. Its AI experiment may run through an existing cloud contract that never appears as a named AI payment. Its workforce may be absent from public profiles. Payroll will record headcount but not whether a maintenance planner used a model to interpret a manual. A manager might answer a survey months later. The missingness is not random.
Testing intersections would make the private signals more credible. Does a first AI payment predict later self-reported adoption? Do task categories align with changes in job postings, and do the postings become hires? Do those hires persist on payroll? Which firms or workers appear in only one source? A failed match may be as informative as a successful one.
Publication should preserve disagreement. A dashboard that compresses all six rows into one AI labor score would remove the reason for combining them. A better release would show task activity, purchased adoption, vacancies, hires, employment, wages, and survey estimates on separate lines, with confidence and coverage for each.
OPM puts AI inside the hiring desk
OPM’s August 27 memorandum turns measurement into a feedback problem.
Federal agencies may use AI to draft job analyses and announcements, develop assessments, or review resumes against underlying documents. They may also find inconsistencies before an offer and evaluate hiring patterns in aggregate. OPM frames adoption as part of a push toward an 80-day average time to hire. It warns that agencies can compromise hiring quality by failing to adopt useful technology as well as by using it poorly.
Around consequential review, the memo draws a firmer line. An official should have access to the full application package, not only a generated summary. Agencies should sample outputs and conduct quality assurance while preserving veterans’ preference, accessibility, privacy, auditability, and reconsideration. Formal approval is weak if the reviewer cannot independently examine the source.
This is federal operating guidance, not a law for every private employer. Its measurement lesson travels farther.
Once AI helps write a posting, screen a resume, or check a hiring file, the employment data reflects both the labor market and the tool. A model may standardize job language and make apparent skill demand look more uniform. A screening system may accelerate some candidates and divert others into review. A faster workflow may increase recorded hiring without changing applicant quality. A system that produces more complete records may look more compliant than a manual process even when the underlying decisions are similar.
Government analysts will need a flag for AI-mediated hiring. Without it, a change in time to hire could be attributed to labor demand when it came from process speed. A change in selected qualifications could look like a worker-supply shift when a drafting assistant changed the announcement. An increase in rejected applications could signal weaker applicants, a changed model, or a new threshold.
From the candidate’s side, the feedback loop feels less abstract. Imagine a veteran whose resume contains the required experience under a job title the model does not recognize. The full source record and a reconsideration path can repair the miss. A generated summary shown to a busy official may bury it. “Human review” becomes real only when the person has time, authority, source access, and a reason to disagree.
Hiring offices should record model version, use case, source documents, recommendation, human change, final disposition, elapsed time, and any correction. Aggregate monitoring should include applicants who exited or were screened out, not only people hired faster. Measures should be split by occupation, grade, location, disability accommodation, veterans’ preference, and other lawful review categories. The purpose is to see where process gains and burdens land.
Vendor-derived labor research needs the same discipline. If OpenAI, Google, Meta, or Amazon contributes data from a hiring product, productivity tool, or business system, the Department should disclose whether the product influenced the outcome later counted. An instrument is not a neutral observer merely because it produces a timestamp.
OPM has already provided one operational standard worth carrying into the data partnership: keep the underlying record accessible and make independent review possible. The Labor Department should apply that standard to the evidence it receives from vendors, not only to the candidates agencies review.
A contract for public labor evidence
People do not need every private row. They need enough structure to know what a row means and enough independence to challenge the conclusion.
Call the structure a public-private AI labor data contract. It is a shared evidence specification, whether the legal instrument happens to be a memorandum, procurement agreement, statistical designation, or research license. Each provider completes the same fields before its data enters a finding.
| Contract field | Required disclosure | Decision protected |
|---|---|---|
| Provider and interest | Legal entity, product, commercial role, funding, and study authors | Reveals who benefits from a reported adoption or augmentation result |
| Observation unit | Message, account, worker, firm, payment, posting, hire, or payroll record | Prevents one unit from being described as another |
| Covered population | Eligible users or businesses, geography, industry, account type, and exclusions | Defines the denominator and who is missing |
| Time and latency | Event period, receipt date, publication lag, and revision window | Separates labor movement from delayed or backfilled data |
| Event definition | Exact rule for use, adoption, hiring, displacement, or augmentation | Keeps labels stable enough to compare |
| Product version | Model, feature, pricing, bundle, and taxonomy changes | Distinguishes product redesign from labor change |
| Fields and transformations | Raw fields used, derived classifications, matching keys, and uncertainty | Allows reconstruction without exposing private content |
| Legal purpose | Authority, statistical-use promise, permitted users, and prohibited secondary uses | Prevents research data from becoming personnel or enforcement evidence |
| Privacy and retention | De-identification, access control, minimum cell size, storage, deletion, and breach response | Protects workers and small employers |
| Sampling and weighting | Sample frame, probability or convenience design, weights, benchmarks, and calibration | Shows whether the result can represent a larger population |
| Missingness | Nonusers, unpaid tools, competitors, offline work, small firms, and unmatched records | Exposes systematic blind spots |
| Matching quality | Join method, false matches, unmatched share, and linkage evaluation | Prevents a clean chart from hiding a weak bridge |
| Causal boundary | Descriptive claim, comparison design, controls, confounders, and alternative explanations | Stops correlation from becoming a displacement verdict |
| Independent validation | Responsible statistical office, external review, benchmark source, and replication route | Keeps the vendor from grading its own evidence |
| Release and revision | Public tables, code or pseudocode, metadata, version history, and correction policy | Gives readers a stable record that can improve |
| Sunset and renewal | Review date, continued need, deletion status, and renewal criteria | Prevents a temporary feed from becoming permanent by inertia |
Completing the table would expose a useful difference among providers. A vendor might offer daily task classifications with a narrow user base. A payroll provider might offer slower employment records with rich worker continuity. A survey may offer representative coverage with limited product detail. The Department can value each contribution without pretending they are interchangeable.
Workers need a direct voice in the specification. A task classifier can label an employee’s work without capturing the review burden, stress, skill gain, or fear created by the tool. A payroll record can show that the person remains employed while hours, autonomy, or promotion prospects deteriorate. Survey modules and qualitative research can test those missing outcomes. Aggregate vendor telemetry cannot replace asking people.
Employer and worker outcomes may move in opposite directions. An AI assistant can reduce the time needed for a report while increasing monitoring, after-hours requests, or the number of cases assigned to one person. Finance may record a successful productivity gain. The employee may record a worse week. A labor study needs both measures before it labels the change augmentation.
Small employers need a place too. Large enterprise customers generate cleaner account records and larger sample cells. A local clinic, contractor, restaurant group, or manufacturer may adopt AI through a general subscription, a feature bundled into existing software, or an owner’s personal account. If the partnership calibrates only to named enterprise tenants, it can make AI adoption look more formal and concentrated than it is.
Shared accounts create another distortion. At a five-person contractor, the owner may pay for one tool and let an estimator, office manager, and project lead use it. The vendor sees one seat. A business survey may record one adopting firm. Three people’s tasks can change, yet no system holds a clean worker-level denominator. That messy case belongs in the coverage note because small employers should not disappear merely for having untidy software administration.
Vendors have legitimate constraints. Product logs can reveal security design, customer strategy, or private communication. Antitrust and contract duties may limit what can be shared. A credible program can use protected enclaves, standardized aggregates, synthetic validation data, or independent statistical agents. Confidentiality is compatible with scrutiny when definitions, coverage, tests, and limitations remain public.
Procurement should separate access from influence. A company may help engineers interpret a field or correct a documented error. It should not receive the power to delay an unfavorable release, choose the headline denominator, or veto an independently supported result. Any prepublication review should be limited to factual and confidentiality checks, logged, and time-bounded.
DOL should publish the memorandums or a detailed public summary for each one. Readers need the participating office, legal authority, duration, data category, statistical purpose, privacy structure, review rights, publication plan, and termination terms. Naming four companies without those fields creates attention, not accountability.
For enterprise buyers, the same contract is a usable diligence file. Before citing a government-vendor finding in a board memo, ask which row resembles the company’s workforce. Before cutting an entry-level budget, check whether the study observed postings, profiles, or payroll. Before promising productivity, look for accepted output, remaining review, wages, hours, and staffing. The public method can improve private decisions only if the units survive the trip into the slide deck.
September still belongs to the public denominator
In September, an economist at BLS will still prepare the next employment release on a fixed public schedule. Survey responses will be weighted, payroll reports will be processed, estimates will be reviewed, and revisions will be documented. At 8:30 a.m. Eastern, every reader will receive the same release.
Somewhere else, a vendor dashboard may update before breakfast. It may show more AI accounts, more messages crossing task boundaries, more corporate spending, or a shift in model calls. The speed is useful. The dashboard cannot see whether a 23-year-old found work or a small firm stopped recruiting. It also cannot decide whether a federal agency shortened hiring fairly, or whether a worker gained a task and lost a promotion path.
New memorandums can connect those clocks. A sensible first publication would be modest: identify each contribution, map its denominator, compare it with a benchmark, and list the results it cannot support. Readers should see which findings were reproduced independently and which remain vendor-reported. If matching changes a result, publish both versions.
Any partnership will face pressure to move faster. Officials want an answer about AI and jobs. Vendors want proof that their products help. Employers want a budget signal, while workers want to know whether a career path is opening or closing. Speed serves them only when the evidence survives disagreement.
Sonderling’s announcement gives the Labor Department access to companies that see work at unusual resolution. It does not yet give the public a method. The August 27 hiring memo shows that the government also plans to become an AI user, adding another source of data and another source of bias.
Publication should come before the forecast. Release the agreements or their full statistical specifications. Show the observation unit beside every percentage. Keep product use, purchasing, vacancies, hires, payroll, and population estimates on separate lines. Let independent statisticians decide how the lines can be joined and when they cannot.
When the next jobs release reaches the screen, private telemetry may explain a movement sooner or in greater detail. It may reveal task change that the headline cannot see. The headline should still belong to a method that covers everyone outside the product. That includes people absent from a vendor’s customer list and workers who never generated a convenient event log.
Fast data becomes public evidence only after the missing people are put back in the frame.