On July 27, OpenAI placed two workspace sizes beside the same work problem.

Among typical-volume ChatGPT users in workspaces with two to five seats, 18.9% of work-related messages involved tasks associated with an occupation other than the user’s own. In workspaces with more than 100 seats, the share was 16.3%.

The difference is modest, yet it resembles a familiar Tuesday at a young company. A salesperson opens a customer dataset and tests a renewal hypothesis before an analyst would have reached the request.

OpenAI calls the pattern “task crossover.” The company analyzed more than 800,000 work-related messages from U.S. users and published the first report in a new Work at the Frontier series. Across the sample, 16.8% of messages crossed an occupational boundary. After broad activities such as writing, summarizing, and scheduling were removed, 43.5% of the remaining occupation-specific messages did.

These are records of requests, not proof that the work was completed well. The study’s unit is a message rather than an hour, project, accepted deliverable, or job. Workspace seats do not equal company headcount. A two-seat workspace can belong to two people inside a larger employer, and a 101-seat workspace can still omit most of a company. The sample covers eight occupational groups and is not representative of the U.S. workforce. OpenAI did not observe whether a specialist reviewed the answer, whether anyone used it, how much time it saved, or whether the same work would have happened without ChatGPT.

The generalist’s request does not establish that a company can skip the specialist. It shows employees reaching across old role boundaries before job descriptions, pay bands, training plans, and acceptance rules have caught up. A missing handoff can feel like speed to the customer and like unrecognized scope to the employee. It can reduce a queue while increasing the chance that an answer reaches production without the person who knows where it breaks.

For a small team, the operating choice is rarely “AI or headcount.” It is whether a borrowed task should stay with the person who encountered it, pass through a specialist, return to a formal handoff, or become enough recurring work to justify a hire.

That choice needs evidence from the work itself. Message volume is an early signal. Accepted output, review load, error cost, customer consequence, and durable role expansion decide what the signal means.

Inside OpenAI’s July 27 work map

The full OpenAI report, written by Caroline Chin and Alex Martin Richmond, starts from a problem with most forecasts about AI and jobs. A forecast usually freezes an occupation into a list of tasks, then estimates which of those tasks a model could perform. The July analysis watches people request work that was never on their original list.

The researchers used a random sample of more than 800,000 work-related messages from individual ChatGPT accounts of U.S. users. Occupational information came from self-reported role data associated with ChatGPT Business. The eight groups were customer experience, design, engineering, finance, human resources, legal, marketing, and sales.

Messages were classified anonymously against work activities in the U.S. O*NET system. The researchers say they did not read the underlying messages. They first separated generic work from activities that could reasonably point to an occupation. Writing, summarizing, and scheduling appear across too many jobs to establish a boundary crossing. A message about financial calculation, software troubleshooting, contract interpretation, or marketing material has a more identifiable occupational home.

That mapping can still call legitimate multi-role work a crossover. An early-stage founder may already own sales, finance, hiring, and operations even if a self-reported occupation names only one of them. The classification captures distance from an occupational category. It does not prove that a worker entered a function the employer considered foreign.

The classification produces three denominators: 61.5% of all work-related messages were generic, 21.8% stayed within the user’s occupation, and 16.8% mapped to a task associated with another occupation.

The report’s larger 43.5% figure applies only after the generic messages are removed. It describes the outside-occupation share of the messages that could be mapped to a specific occupational source. It would be wrong to say that nearly half of every work message crossed a job boundary.

The denominator is a practical warning for employers. A dashboard that counts any use of writing or summarization as “role expansion” will exaggerate how far work moved. A dashboard that counts only formal job-title changes will miss the movement until much later.

A useful internal measure begins with the capability borrowed, the purpose of the request, the person who accepted the output, any specialist review, and whether the work recurred.

A salesperson may ask ChatGPT to summarize trends in an exported customer table. If the answer helps the salesperson decide which questions to ask an analyst, it is exploration. If it goes into an internal territory plan, it becomes a work product. If it changes a customer’s price or renewal terms, it supports a consequential decision. The prompt can look similar in all three cases while the acceptance path changes.

This makes the July report more useful as an early-warning instrument than as a productivity benchmark. It reveals where people are trying to remove a handoff. Whether that removal was economical, fair to the employee, or safe for the customer can be found only in the route from request to accepted work.

Inside the 43.5% denominator

Once generic activity is excluded, outside-occupation tasks account for 77% of occupation-specific messages from customer experience workers, 75% from designers, 69% from human resources workers, 56% from legal workers, and 53% from marketers.

Those percentages do not mean that 77% of everything a customer experience worker asks ChatGPT is outside the role. They describe the occupation-specific portion after generic requests have been removed. A customer service team may still spend most of its day on customer service. The data show that when its prompts point clearly to a specialized occupation, the source is often somewhere else.

The groups may also generate very different volumes of occupation-specific prompting. These shares show the mix inside each group; they do not rank which occupation produces the most crossover messages in total.

Marketing and engineering work travel especially far. OpenAI’s heat map shows marketing tasks appearing across every row, while engineering tasks show up frequently in design, finance, HR, legal, marketing, sales, and customer experience use. Financial calculation and technology troubleshooting rank among the three most common outside tasks for all seven of the other occupational groups.

The salesperson sees an unusual customer pattern before analytics schedules a study. The person closest to the symptom can now ask for the first layer of work without waiting for another queue.

Speed comes from collapsing coordination time as much as from generating an answer. A request no longer has to be written as a ticket, prioritized, clarified, and returned. For a small company with no dedicated analyst or counsel, the old handoff may have meant an outside invoice or no answer at all.

The report records a moved first attempt. The specialist’s contribution remains outside its field of view.

Financial calculation is a good example. A sales manager can use AI to explore the effect of a discount on annual contract value. That may be sufficient for a private scenario. The same calculation can become a finance decision when it changes a forecast, a customer commitment, revenue recognition, or board reporting. The arithmetic may be simple while the surrounding policy is not.

Software troubleshooting follows the same path. A marketer can inspect an error, compare likely causes, and repair a low-risk page in a test environment. Production access, customer data, security settings, and infrastructure changes create a different acceptance threshold. The ability to produce a plausible fix does not grant authority to deploy it.

The occupational source describes only one dimension of crossover. Employers also need to record whether the output remained a private aid, entered an operating document, reached a customer or production system, received another function’s acceptance, and became the employee’s responsibility to maintain.

The last question changes a borrowed task into a possible role change. A one-time website repair is neighborly work. Owning the page every week is an operating responsibility. Drafting one customer analysis is an experiment. Becoming the team’s default analyst without training, review time, title, or pay is a staffing decision made by neglect.

The July data reveal where that neglect may begin, leaving the employer to resolve it from the work that followed the prompt.

Small workspaces and the missing specialist

OpenAI finds more crossover among typical-volume users in smaller workspaces. The outside-occupation share is 18.9% for workspaces with two to five seats and 16.3% for workspaces above 100 seats. Because the published percentages are rounded, the difference is best described as roughly two and a half percentage points.

The comparison applies to users in the middle 50% of message volume. Among the heaviest users, the report does not find the same orderly decline as seat count rises. Heavy users may have established workflows that look similar across company sizes, or they may use ChatGPT more deeply inside their own occupation. The data cannot distinguish those explanations.

Seats are also a rough organizational proxy. They reveal access to a workspace, not the employer’s full org chart. A six-person startup might share a two-seat workspace. A regional team inside a multinational might do the same. OpenAI’s public article describes the pattern as more crossover in small businesses, but the report’s observable variable is workspace size.

With that boundary intact, the result still fits the economics of a small team. Carta’s H2 2025 startup compensation data put the median seed-stage team at four people. Its dataset also found that average Series D headcount had fallen 29% from its 2023 peak, while the median individual-contributor salary rose 6.4% and the median initial equity-grant size increased by nearly 11% over the prior two years. These measures do not establish an AI effect. They describe a market in which adding a specialist is expensive and many startups are trying to reach a milestone with fewer people.

In a four-person company, functions exist before jobs do. Someone has to price the product, inspect customer behavior, fix the site, reconcile cash, review terms, recruit, onboard, and answer support. A founder or early employee often carries several of those functions until volume, risk, or delay supports a dedicated hire.

AI changes the carrying capacity of that arrangement. It gives the generalist a faster first draft, a way to inspect unfamiliar material, and a lower-cost route to a question that might otherwise wait. The same tool that keeps a handoff from blocking a customer conversation can make an overloaded role look sustainable for longer than it is.

The U.S. Chamber of Commerce Foundation and Ipsos surveyed 1,070 workers at American small businesses in May. Their Main Street AI Monitor found that half used AI at work. Among those users, 64% primarily used it for personal productivity, 26% for recurring tasks, and 6% for workflows with minimal human involvement. When the tools saved time or improved quality, 59% said they used the gain to do more work or produce better output; 27% took on stretch assignments or new responsibilities.

Adoption was more often described as employee-driven, at 19%, than organization-driven, at 11%. Only about one in 10 respondents had been offered formal AI training. Privacy and security concerned 47%; unclear applicability and skill gaps each concerned 41%.

The Chamber data and OpenAI data measure different populations and behaviors, so their percentages should not be combined. Together, they describe an order of operations: employees experiment, tasks spread, and the employer’s support system arrives later.

A startup should track that lag in weeks, not wait for an annual workforce plan. Once a borrowed task becomes weekly, managers need to see its volume, review burden, consequence, and effect on the employee’s original work. Otherwise AI expands the job while the company keeps evaluating the old one.

Borrowed capability, borrowed risk

A handoff can disappear from the calendar while its review work remains. The generalist keeps moving instead of waiting for another function, but someone still has to check assumptions, identify the authoritative source, recognize exceptions, carry liability, and maintain the result.

Some borrowed tasks are cheap to review. A marketer can ask an engineer to look at a proposed change before it reaches production. The engineer may need five minutes to confirm that the fix is harmless. Other reviews reconstruct most of the original work. A finance lead who must rebuild a model, trace each input, and correct the policy treatment did not receive an efficiency gain. The company created an extra draft.

Review load helps show whether crossover is working. Track the specialist minutes needed to accept the output, the returns for revision, material errors, and the time between first request and accepted result. If review repeatedly consumes a large share of the time required to do the work properly, the apparent crossover has made the handoff harder to see.

Mandatory review can recreate the queue that AI was supposed to remove. It can also fragment the specialist’s own work as small checks arrive all day. Review capacity needs a service target and a place in the specialist’s goals. Treating it as free safety labor merely moves overload from one role to another.

Part of the transfer lands on the generalist, who has to decide whether an answer is safe enough to use without having the specialist’s pattern recognition. A fluent response can make that threshold difficult to locate. People may know enough to ask a good first question and still lack the experience to notice the missing fact that changes the answer.

Microsoft’s 2026 Work Trend Index provides a useful counterweight to any claim that crossover removes judgment. Microsoft surveyed 20,000 AI users across 10 countries and analyzed more than 100,000 Microsoft 365 Copilot chats. Half of surveyed users said quality control of AI output had become more important, 46% said the same about critical thinking, and 86% described AI output as a starting point. The self-reports are not audited behavior, but they identify the skill that a crossover metric omits: acceptance.

Gallup’s July 27 synthesis on AI and workplace productivity shows a similar distance between individual activity and organizational change. Within organizations that had implemented AI, 65% of employees said it improved their productivity and efficiency. Only 12% strongly agreed that AI had transformed how work gets done across the organization.

Frequent use was reported by 79% of employees who strongly agreed their manager supported AI use, compared with 46% among those who did not. Only 25% of U.S. employees said their organization had communicated a clear AI adoption strategy.

Gallup’s page synthesizes several studies and time periods, so those percentages should not be treated as one unified experimental sample. The relationships are also self-reported and associative. They do not prove that manager support caused productivity.

The association helps explain why a license cannot serve as a work design. A manager has to state where the generalist can act, where a reviewer enters, and who owns the final result. The boundary should follow consequence rather than occupational prestige. A public marketing claim, production change, employee decision, statutory filing, customer price, or contractual commitment deserves a named acceptance owner, while exploratory work can use a wider boundary.

When pay bands lag the work

The HR file may still describe last quarter’s job after the team has learned to depend on a broader one.

Task crossover can reverse that sequence. An employee solves a neighboring problem once, then again, then becomes the default person because the answer arrived faster than a formal hire. The organization gets a broader role before it decides whether the person has received broader authority, training, evaluation criteria, or compensation.

Early employees may expect some ambiguity. It becomes unfair when the company treats expanded scope as an informal favor while using it as a dependable operating system.

Managers can separate learning from role expansion with a short scope record. It should name the borrowed capability because “used AI” says little about whether the employee was doing customer analysis, financial modeling, software troubleshooting, policy interpretation, recruiting operations, or marketing production.

The record should also show the acceptance level. Exploration, internal draft, reviewed deliverable, and independently accepted deliverable represent different responsibility. A person who prepares a model for finance review is building a useful skill. A person who owns the forecast has a finance responsibility.

Recurrence completes the record. A stretch assignment can be intentionally temporary. Work that persists through a planning cycle, appears in performance goals, or displaces a material share of the original role belongs in a scope review.

The review does not automatically require a new title or a raise. It requires an explicit decision. The company might remove the task, pay for specialist support, train the employee, adjust objectives, recognize a new skill, or redesign the role. Silence leaves the cost with the person who kept the work moving.

Current labor-market signals make that silence more consequential. Indeed Hiring Lab reported on July 23 that U.S. job postings were tilting toward seniority: as of May 2026, senior postings were up 14.7% from a year earlier while entry-level postings were down 7.5%. Indeed did not attribute the movement solely to AI, and macroeconomic demand, industry mix, and employer caution can all affect the pattern.

The signal still raises a development problem for companies asking existing workers to cross role boundaries. Senior specialists learned judgment by doing work that could be checked, corrected, and gradually expanded. If AI gives a generalist an immediate first draft while hiring shifts upward, the company needs a deliberate route from attempt to reviewed competence. Otherwise it borrows specialist output without building specialist depth.

A reviewed task log can become that route. A junior employee starts with exploration, advances to a draft that a specialist annotates, then owns a bounded output after repeated acceptance. The log preserves the corrections that teach judgment and gives a manager evidence for expanded scope, rather than letting AI erase the apprenticeship between first attempt and independent ownership.

ManpowerGroup Talent Solutions and Everest Group published a small executive study on July 22. Among 80 senior leaders in the United States and United Kingdom, 34% said their strongest productivity gains came from augmented roles, compared with 8% who pointed to fully automated roles. Only 3% described leaders as highly prepared for AI-enabled work.

The vendor-commissioned sample is far too small to establish market prevalence. Caroline Pfeiffer Marinho of Talent Solutions used it to argue that organizations now face an adaptation problem after deployment. Sailesh Hota of Everest Group emphasized work organization rather than tool installation.

For a founder, the useful implication is modest: augmentation changes the employee’s side of the bargain. If a role can now reach into analysis, engineering, finance, or marketing, performance review and compensation should account for the accepted value and responsibility of that reach. Counting only the original job description makes expanded work free on paper.

A handoff-or-hire matrix for small teams

A team should make the routing decision before a borrowed task becomes invisible infrastructure.

The following matrix is an operating template, not an empirical rule. Each company should set thresholds that match its product, customers, regulation, cash position, and tolerance for error.

Borrowed taskKeep with the generalist whenRequire specialist review whenHand off or hire whenEvidence and accountable owner
Marketing copy or market researchThe output is exploratory, uses public information, and stays inside a draftIt contains a product claim, customer comparison, pricing promise, or brand-sensitive statementCampaign volume is recurring, review queues block launches, or positioning needs durable ownershipSource links, claim check, revision history; marketing owner accepts publication
Customer or sales data analysisData is approved, aggregated, and used to form internal questionsThe result changes segmentation, forecast, discount, renewal, or customer treatmentRequests recur across teams, models require maintenance, or review rebuilds the analysisDataset version, query or method, assumptions, validation sample; analytics or finance owner accepts use
Software troubleshootingWork stays in a local or test environment and has an easy rollbackA change touches production, permissions, integrations, customer data, or reliabilityIncidents repeat, on-call load rises, security exposure appears, or no qualified reviewer is availableError record, proposed change, test, rollback, reviewer; engineering owner accepts deployment
Financial calculationThe calculation supports a private scenario with documented inputsIt enters a budget, board material, customer quote, payroll decision, or cash planStatutory, tax, audit, financing, or repeated planning work requires professional ownershipSource values, formula, scenario, reconciliation; finance owner accepts the number
Contract or policy summaryThe employee is identifying issues and preparing questionsThe interpretation affects an employee, customer, vendor, or company obligationNegotiation, dispute, regulated advice, or repeated legal work exceeds internal competenceDocument version, cited clause, open issues, counsel notes; authorized legal owner accepts the position
Job description, leveling, or pay analysisThe output is an internal draft based on current approved frameworksIt changes scope, grade, compensation, performance criteria, or candidate representationPay equity, employment law, workforce scale, or recurring calibration needs dedicated expertiseFramework version, benchmark source, decision log; HR and the hiring manager accept the action

Use the matrix with the destination and review burden in view. A private analysis can tolerate uncertainty that a customer commitment cannot. A usable draft should reduce the reviewer’s effort; repeated reconstruction is a reason to repair the process, develop the generalist’s skill, or restore the handoff. Recurring work needs an owner, and that owner needs authority to stop or correct it.

A customer whose price, renewal, eligibility, or support route changes also needs recourse. The acceptance owner should know how the person can challenge an input, correct the record, and reach someone authorized to reverse the outcome.

The economic comparison should include more than salary.

For work kept with a generalist, count tool cost, user time, training, specialist review, rework, delay to the original role, and expected error cost. For a formal handoff, count specialist time, queue delay, coordination, and outside fees. For a hire, count loaded compensation, recruiting, ramp, management, and the value of capacity the person creates.

That produces a simple decision frame:

keep cost = tool + generalist time + review + rework + role displacement + expected risk

handoff cost = specialist time or fee + coordination + queue delay

hire case = durable demand + strategic ownership + capacity value - loaded employment and ramp cost

The formula is a checklist, not a promise of precision. Expected risk will often be a range. Capacity value may depend on a product milestone or sales plan. Its purpose is to keep “the model did it” from erasing the human work around the model.

A founder can fill the frame with a real week. Suppose a sales lead spends five hours on recurring customer analysis, a finance lead spends three hours reviewing it, and corrections consume another two. The company is using 10 internal hours before counting interruptions to either person’s main work. A fractional analyst who can deliver and explain the same accepted output in six hours, with one hour from sales for context, may be the cheaper handoff. A full-time hire still needs several durable workstreams, enough urgency, and enough strategic ownership to justify the ramp. The hours are illustrative; the team should replace them with its own time and rates.

Record the route as keep, review, handoff, or hire. Reopen it after a material correction, customer escalation, new data class, regulatory change, sustained recurrence, or a review queue above the team’s service target. Those are internal triggers, not market benchmarks.

Training and pay belong in the same record. If the route stays with the generalist, specify what competence the employee is expected to develop, who coaches it, how accepted work will be evaluated, and when sustained scope enters title or compensation review.

The matrix leaves room for generalists to grow without hiring for every unfamiliar prompt. It also prevents a company from stretching one capable employee across every missing function simply because ChatGPT makes the first attempt cheap.

A Friday customer analysis reaches finance

Consider a Friday at a six-person software company. A sales lead has a new analysis of customer usage.

The lead exported approved account data, asked ChatGPT to find patterns, checked several rows, and found that one segment appears to adopt faster and renew at a higher rate. The result could change Monday’s pipeline meeting. It could also change which customers receive a discount and how next quarter’s revenue is forecast.

The old route was slow. The sales lead would send a request to a fractional analyst or ask the founder to build a spreadsheet over the weekend. The new route produces a table before the day ends.

The message has crossed an occupational boundary, while the decision still needs an owner.

If the table is a hypothesis for Monday, the sales lead can keep it. If it changes a forecast, the finance owner needs the source data and assumptions. If it becomes a recurring weekly analysis, the team should decide whether sales owns a reviewed workflow, an analyst takes the handoff, or the company now has enough demand for a data hire.

The useful part of OpenAI’s July report is that it makes this moment visible before a title changes. One in six work-related messages in the sample already points outside the user’s occupation. Among the occupation-specific portion, the share is much larger. Smaller workspaces show somewhat more of it among typical-volume users.

The report leaves the Friday analysis’s accuracy, the sales lead’s pay, the review threshold, and the return on a possible hire unanswered. It can still show the founder where to look: the task that stopped entering a queue, the employee who quietly absorbed it, the person who accepted the answer, and the review or rework between them. Those observations are enough to name a route.

A two-to-five-seat workspace and a 101-plus-seat workspace may produce different amounts of crossover, but neither size removes the need for an accountable owner. AI can move the first attempt across an org chart in seconds. The company still has to decide where judgment, authority, learning, and pay should land.

The sales lead sends the table to the finance owner with the source file, method, assumptions, and one explicit question: safe for Monday’s forecast, or useful only as a hypothesis?

The finance owner notices that mid-quarter expansions were counted as renewals. The segment still looks promising, but the reported renewal rate cannot enter the forecast. She marks the table hypothesis-only, writes the cohort definition beside it, and schedules a corrected query for Monday.

The review takes minutes and preserves much of the shorter route. It also shows what the prompt data leave out: the specialist judgment that keeps a plausible table from becoming a false commitment.