An employee studies a difficult case at a workbench while an AI conveyor delivers finished pages under the headline JUDGMENT NEEDS PRACTICE.

AI-generated editorial illustration.

On September 21, IBM released a workforce study with a problem hidden inside two percentages. Among the chief human resources officers it surveyed, 71% put skills for evaluating, validating and overriding AI outputs near the top of their workforce priorities. Only 38% of surveyed employees ranked those same skills as important for their own near-term success.

The study reached 1,500 CHROs and equivalent executives and 8,800 full-time employees. It also found that 60% of employees worried automation could cause their skills to atrophy. Employers were asking for more human judgment at the same moment workers feared they were getting fewer chances to exercise it.

IBM CHRO Nickle LaMoreaux framed the company response as work redesign: routine tasks move to AI while people focus on areas where they add more value. That is an executive prescription from the company that commissioned the study. It does not establish that workers receive the time, authority or career credit required to make the prescription real.

The evidence does not support a clean story of AI making people less capable. IBM measured perceptions, priorities and reported work conditions rather than cognitive decline over time. Other research has found that AI can help novices absorb expert practices, improve immediate learning and produce better work. The direction changes with the task, the interface and what the person still has to do.

Output can rise while practice falls. If an assistant writes the first draft, finds the anomaly, chooses the formula and proposes the answer, the employee may finish sooner. They may also perform fewer repetitions of the steps that once built pattern recognition. A dashboard can record the saved minutes. It rarely records the missing rehearsal.

That gap matters to employees deciding what to learn, managers assigning work, HR teams designing careers and finance leaders evaluating an AI budget. The task is the useful unit. Broad claims that AI either augments or replaces people blur what the system does, what the person must still notice, and whether the person can handle the next difficult case without being carried by the tool.

This article compares survey, field and experimental evidence available as of September 23, 2026. It does not diagnose skill loss in any individual or claim that one workflow fits every occupation. The skill-retention ledger below is an editorial operating framework, not a validated assessment or legal standard.

A 71-to-38 evaluation gap

The full IBM CHRO study describes two related gaps. Executives expect people to supervise, validate and override AI. Employees are less likely to see those evaluation skills as an immediate career priority. At the same time, 57% of the executives and 49% of employees prioritized critical thinking and problem framing. Workers have not rejected judgment. Their ranking of the specific practices around AI is lower than leadership’s demand for them.

A CHRO sees systems spreading across functions and imagines a workforce that can catch bad outputs. An employee sees today’s queue, response-time target and approved-tool dashboard. Both views make sense from where each person sits. If the performance system rewards completion while a leadership memo celebrates judgment, the completion metric wins each afternoon.

Unclear task ownership muddies the incentives. IBM found that only 26% of organizations clearly define which activities should remain human-led, which should be AI-assisted and which can be AI-executed. Fifty-two percent of employees said AI had changed the tasks assigned to them during the previous year. A person may therefore be held responsible for a result produced through a workflow that nobody has formally assigned.

Two executives quoted in IBM’s release described what they want instead. Amit Das, CHRO of Bennett Coleman & Co., put the speed problem in ten words: “Fluency without judgment simply helps an organization make mistakes faster.” Hyatt CHRO Kristin Oliver said her company asks whether AI improved decision quality, employee experience and business performance, rather than counting completed work alone. These are practitioner positions, not audited outcome studies, but they make the proposed shift concrete.

Accountability does not disappear with the work. Forty-three percent of employees in the study said they expected blame to fall on them when AI fails. Forty-one percent of CHROs said employees might not feel safe challenging AI outputs. These findings do not show how often an error occurred or whether a worker was punished. They expose a design problem: responsibility can remain human while the authority and time needed to exercise it become uncertain. Fear of challenging the system makes that mismatch harder to surface.

Verification vanishes from view. Four in five CHROs told IBM that AI creates invisible work. Forty-two percent of employees said the technology either increased their work or created work that went unrecognized. Checking citations, recovering a missing source, comparing two recommendations and explaining why an AI suggestion was rejected can all improve the final result. None is visible in a count of generated documents.

An employer that wants judgment has to decide whether those actions are work. If verification is an unofficial extra, employees learn that careful review is a private tax. If finding a model error delays a delivery target without earning credit, silence becomes the efficient response. The 71-to-38 gap joins a training problem to an incentive problem: leadership requests one capability while the operating system rewards another behavior.

Self-reported erosion is a warning, not a diagnosis

The phrase “skill erosion” can sound more settled than the evidence permits. In IBM’s September 21 release, 60% of employees expressed concern that automation could atrophy their skills. Concern is important evidence about trust and career risk. It is not a before-and-after test of what they can do.

A widely discussed Microsoft Research and Carnegie Mellon study has a similar boundary. Hao-Ping Lee, Advait Sarkar and their co-authors surveyed 319 knowledge workers about 936 examples of using generative AI at work. Higher confidence in the AI was associated with less self-reported critical thinking, while higher confidence in one’s own ability was associated with more. Participants described thinking moving away from producing material and toward verifying, integrating and supervising it.

Verification is work. A plausible recommendation may require more expertise to check than a routine first draft requires to compose. Review can still become shallow if the person accepts fluency as accuracy. The study captured how workers described their effort across examples. It did not test whether their reasoning ability declined after months of use.

The remedy depends on what changed. Treat every use of AI as deskilling and a company may protect outdated tasks while denying employees useful support. Count every completed output as augmentation and it may miss the practice that prepares people for exceptions. Neither label tells a manager whether the next analyst can investigate an unfamiliar discrepancy or whether a recruiter can recognize when a ranking criterion is wrong.

Skill is also plural. Recall, procedure, diagnosis, communication and judgment can move in different directions. An AI assistant may reduce the need to memorize a command while increasing the need to frame a problem. It may teach a novice the standard response while hiding the rare condition that an expert notices. A single employee can gain one capability and lose practice in another during the same workflow.

Some employees will welcome that trade. A worker may reasonably prefer to stop memorizing syntax, rewriting routine emails or searching a knowledge base by hand. A retention program that treats every fading procedure as a loss would turn yesterday’s busywork into tomorrow’s test. The organization must name the capabilities it will retire as clearly as the ones it expects people to retain. Otherwise “protecting skills” becomes an argument for keeping low-value work.

Periodic sentiment surveys stop here. They can reveal fear, confusion and perceived workload. They cannot establish which capability changed, in which task, over what period, or whether the change transferred to a new case. A credible diagnosis needs a baseline and a later performance sample that does not let the same assistant supply the answer being tested.

Teams can create a baseline without turning it into surveillance. They can use anonymized scenario exercises and sampled quality reviews instead of recording every prompt. The organization needs to observe the work capability it says it depends on. Counting AI sessions is a poor substitute for measuring whether people can detect, explain and recover from a failure.

Productivity can rise while practice disappears

The strongest counterargument to a skill-loss narrative is that AI can carry expertise to people who did not previously have it. In the NBER field study Generative AI at Work, Erik Brynjolfsson, Danielle Li and Lindsey Raymond followed 5,179 customer-support agents using an assistant trained on successful interactions. Access increased productivity by about 14% on average. The gain was much larger for novice and lower-skilled workers, while the most experienced agents saw little change.

The tool surfaced patterns from stronger performers during a live conversation. Newer agents could handle more issues and appeared to adopt language associated with experienced colleagues. In that setting, AI distributed organizational knowledge that coaching had not spread as quickly. The assistant could supply practice material as well as save time.

The study did not remove the assistant and test what remained. Yet it shows why a blanket ban on AI for novices would be costly. People can learn from examples, feedback and structured suggestions while doing real work. The relevant question is whether the workflow asks them to interpret the suggestion or merely routes it through them.

A 2026 randomized learning experiment by Zara Contractor and Germán Reyes offers another counterexample. Participants with AI access scored 0.27 standard deviations higher on an immediate knowledge test, and some gains persisted one week later. The pattern depended on behavior. Users who sought explanations retained learning; users who focused on automating the answer produced better short-term work but their advantage faded on the delayed measure.

That study involved students and a short follow-up, rather than employees navigating production systems. Its value lies in the mechanism it isolates. Tool access alone does not determine learning. Asking for an explanation and constructing an answer can turn the same model into a tutor. Delegating the task end to end can turn it into a substitute.

Two paths circle an AI workbench: one sends finished pages directly to an output tray, while the other passes through human inspection, correction and an unaided practice station.

AI-generated editorial illustration.

The longer path represents practice and verification, not a requirement to add friction to every low-risk task.

Managers often see only the final column of this tradeoff. Output per hour goes up. Backlogs shrink. More customers receive a response. Those are real benefits, and an employee may reasonably prefer to stop doing repetitive work. The missing measure is whether the faster workflow leaves enough exposure to variation, mistakes and recovery for people to learn the underlying system.

Routine cases finance readiness for rare ones. Junior accountants learn by tracing reconciliations before they own an unusual close. Support agents learn the product through ordinary tickets before a novel incident arrives. Recruiters learn a labor market while reading imperfect resumes before they have to challenge an automated ranking. Removing every routine repetition may save time now and create an experience shortage later.

Work design has more settings than full manual work or full automation. A company can automate stable steps while reserving a sample for unaided practice, require the worker to form an initial view before seeing the model, or rotate responsibility for diagnosing failed cases. It can let the tool explain a recommendation and ask the person to compare alternatives. These designs spend a small part of the productivity gain on keeping the capability that handles exceptions.

Inside the jagged frontier, confidence becomes a risk

AI performance is uneven across tasks that look similar. At that irregular boundary, human confidence can become as consequential as model capability. A person who receives ten good answers may be less prepared to slow down for the eleventh, especially when the system presents both with the same polish.

Fabrizio Dell’Acqua and his co-authors tested 758 Boston Consulting Group consultants in the jagged technological frontier experiment. On tasks inside the model’s capability frontier, access to GPT-4 helped participants complete more work, faster and at higher quality. On a task outside the frontier, participants using AI were 19 percentage points less likely to reach the correct answer than those without it.

The consultants were operating in a familiar professional domain, yet some accepted a persuasive wrong path. The experiment cannot establish that current models fail at the same rate or that consultants lost skill. It used GPT-4 in 2023 and bounded exercises. The management problem remains: past success with an assistant can increase the chance that a person relies on it in a nearby task where it is weak.

Traditional training often teaches a tool’s features. Jagged-frontier work requires something harder: recognizing when the problem has changed. That depends on domain knowledge, access to original evidence and permission to reject the output. A checklist can confirm that a citation exists. It cannot decide whether the cited population is comparable to the customer in front of the employee.

The employee also needs time. A target built around the AI-assisted average can remove the margin required to examine an exception. If the model produces a report in five minutes, a manager may set a ten-minute expectation for the completed task. The human reviewer then has less time than before to inspect sources, reconstruct assumptions and document an override. A nominally human-in-the-loop process can exist on a diagram while the schedule makes meaningful review impossible.

Error recovery is the missing practice in many deployments. Teams record that a person can approve or reject an answer, but not who investigates after rejection. When the model suggests an invalid discount, misclassifies an applicant or hallucinates a policy, someone has to find the original record, correct the case, assess similar outputs and decide whether the workflow should pause. That sequence is where institutional knowledge grows.

Blame can concentrate in the same place. IBM’s employee findings suggest that workers expect responsibility for AI failures even when system boundaries are unclear. A useful override mechanism includes escalation and credit alongside a red button. It records that catching a failure protected an outcome. Otherwise employees learn that accepting the answer is fast, while challenging it creates unpaid investigative work.

Training changes outcomes only when work changes

More courses will not repair a workflow that removes every chance to apply them. IBM reports that 80% of organizations have a reskilling roadmap, yet its study distinguishes a roadmap for filling skill gaps from the ongoing risk of skills atrophying. A course can introduce evaluation techniques. The production environment decides whether employees can use them.

One Microsoft Research field experiment illustrates the difference. Alex Farach, Alexia Cambon and their co-authors studied 388 employees at a Fortune 500 retailer using the same generative AI tool. A behavioral protocol that prescribed a sequence of human and AI actions did not improve the assessed outputs and reduced production. Training that framed AI as a cognitive thought partner was associated with more documents near the top of the quality distribution.

The authors disclose important limitations. The training conditions ran at different times of day, attrition differed, and the language-model grader was sensitive to document length. The paper is a preprint. It does not prove that a brief cognitive course will improve every deployment. It does suggest that teaching people how to reason with a tool is different from prescribing a fixed click path.

Judy Hanwen Shen and Alex Tamkin ran another set of randomized skill-formation experiments in which participants learned an asynchronous programming library. AI access harmed conceptual understanding, code reading and debugging on later tests, even though some participants completed the initial task more quickly. The pattern varied across six ways of using the assistant. Full delegation was harmful to learning; asking for explanations and preserving active problem solving could protect it.

The narrow programming task, reported in a preprint, cannot carry a general claim about workplace deskilling. Paired with the positive learning experiment, it supports a more useful point: interaction design and user strategy change the result. “AI access” is too coarse a treatment category. Two employees can use the same product and receive opposite learning effects because one constructs and tests an answer while the other accepts a completed one.

The OECD’s 2026 AI and Skills report adds a labor-market view. Advanced AI development skills are required in fewer than 1% of jobs, while analytical thinking, resilience and digital literacy appear alongside AI exposure much more broadly. More than half of workers who reported using AI said they had received employer-funded training. Training was associated with better self-reported performance and working conditions.

Much of the OECD evidence draws on data collected before the latest model cycle, and association does not establish that training caused the improvement. Still, it corrects a common planning error. Most employees do not need to become model engineers. They need domain-specific practice in framing a task and checking evidence, followed by enough authority to handle uncertainty and recover when the system fails.

Current worker behavior shows what happens when employers leave that curriculum implicit. An iCIMS survey of 1,000 U.S. job seekers found that 47% said they had built AI skills in the previous six months, while the share relying on self-teaching rose from 22% to 30%. About one in six reported employer-provided training. The vendor survey is not representative of every employee, and it does not measure skill quality. It does show people responding to market signals with uneven support.

A stronger program joins instruction to task assignment. Employees learn a concept, use it on a real but bounded case, receive feedback, explain an override and return later for a transfer exercise. Managers reserve some work for learning instead of sending every easy case to automation. HR recognizes review and error recovery in workload and career evidence. The process builds the capability the company claims to value.

A skill-retention ledger for AI-assisted work

Organizations already maintain model inventories, risk registers, training records and productivity dashboards. None answers a basic workforce question: can the people accountable for this task still perform the judgment that the task requires? A skill-retention ledger can sit beside those systems without becoming another exhaustive compliance file.

Organize the ledger around a task or workflow, not a person-level score. It should expose the tradeoffs before a team automates every repetition. A useful record has twelve fields.

FieldQuestion to recordEvidence at review
Task and decisionWhat result is being produced, and which decision can change a customer, worker or financial outcome?Named workflow, decision owner and affected group
Unaided baselineWhat could a trained person do before this AI step was introduced?Historical sample or bounded scenario, not a memory of past performance
AI roleDoes the system draft, recommend, rank, decide or execute?Actual production configuration and fallback path
Human judgment stepWhat must the person notice, compare or decide?Observable action, not the phrase “human review”
Original evidenceCan the reviewer reach the source material behind the output?Source link, record or calculation available in the same workflow
Known failure and overrideWhich credible error has the team rehearsed, and can the person stop the process?Test case, override route and escalation owner
Practice intervalHow often does the role perform a representative task without a completed AI answer?Scheduled sample based on risk and learning need
Transfer testCan the person solve a new case rather than repeat the training example?Delayed scenario with different surface features
Error recoveryWho investigates, corrects related cases and updates the process?Named owner, time allocation and incident record
Review laborHow many minutes of checking, comparison and documentation does the workflow require?Sampled time, including rejected outputs
Career creditHow does verification, teaching or a justified rejection appear in evaluation?Work evidence tied to role expectations
Business outcomeDid quality, speed, cost, safety or customer experience change?Outcome measure with baseline and observation window

The fields separate four quantities that are often blended: output, learning, control and employee cost. Output measures whether the team produced more or better work. Learning measures whether capability transfers to a new case. Control measures whether people can see evidence, challenge the system and recover. Employee cost includes review time, interruptions and the career effect of doing work that dashboards ignore.

Practice intervals should vary by task. An employee need not prove unaided skill for low-risk text formatting each week. A role that approves payments or ranks candidates needs more frequent and realistic checks. The interval should reflect consequence, rate of change and how quickly a skill fades without use.

Give the ledger a deletion rule. If a task becomes reliably automated and the organization no longer needs the human capability, leaders should say so, redesign the role and plan the transition. Pretending that judgment remains essential while removing time, access and career credit for it transfers the risk to employees. Honest elimination is more manageable than ceremonial oversight.

There is a fair objection to adding any ledger. It can become another HR form, another productivity score or a pretext for testing workers on tasks the employer already automated away. The defense is strict scope: the record belongs to the workflow, it collects the minimum evidence needed for a named capability, and it expires when the organization retires that capability. A person-level score requires a separate purpose and safeguards. The default should be team learning, not individual surveillance.

For capabilities the company chooses to retain, the ledger turns training into a budget question. How much of the saved time will be reinvested in practice, review and feedback? IBM found that 42% of CHROs said productivity gains were primarily reinvested in innovation and reskilling. That is a reported allocation pattern, not a verified return. The ledger makes the reinvestment visible at the workflow where capability can actually be observed.

The weekly review that keeps judgment visible

A company does not need to test every employee every week. It needs a small operating rhythm that reveals whether the workflow is producing both reliable output and capable people. A manager, a domain specialist and one employee who uses the system can review a sample in 30 minutes.

Start with one accepted output, one corrected output and one case where the employee rejected or bypassed the AI. Ask what evidence the system exposed, what the person noticed, how long verification took and what happened after a problem was found. If the team cannot produce a rejected case over several reviews, that may indicate excellent performance. It may also mean nobody feels able to record disagreement.

Next, choose one transfer case. Remove the completed answer and ask a rotating team member to frame the problem, identify the evidence and explain the likely failure modes before using the assistant. The exercise should resemble real work but remain separated from a live customer or employment decision. Record the capability at the team level unless an individual assessment is genuinely necessary and governed.

Then compare the four measures. Did output improve? Did a new-case test show retained judgment? Could the person reach original evidence and override the system? How much invisible review or recovery work was required? A productivity gain with a falling transfer score calls for a workflow change. A stable transfer score with excessive review time may mean the automation is not economical. A strong result on both supports broader use.

The review should produce an assignment, not a slogan. A product owner may need to expose sources. A manager may reserve two cases for unaided practice. HR may add error recovery to the role rubric. Finance may include verification minutes in the business case. A learning team may replace a generic course with a scenario drawn from a real failure.

Employees need a visible return for the effort. Catching a model error, teaching a colleague or documenting a sound override should count as contribution. The person who preserves quality should not appear slower than the person who accepts every draft. When judgment becomes legible in staffing, workload and promotion evidence, the skill priority that CHROs describe can become rational for employees too.

AI will continue to remove some practice, create other forms of practice and make certain old skills less valuable. Leaders can let obsolete manual steps go. They still need to notice when a workflow removes a capability for which people remain accountable.

At the next budget review, ask for two lines. One records the output gained from AI. The other records the practice, verification and recovery required to keep human judgment real. If the second line is blank, the organization has not eliminated the cost. It has made the cost harder to see.