The Factory Floor's 62-Point AI Gap
On July 23, Siemens and HD Hyundai announced plans in Washington, D.C., to create a repeatable model for an AI-powered digital shipyard. Their nine-figure agreement bundled ship design, lifecycle data, manufacturing operations, production simulation, systems engineering, and industrial AI into a digital thread intended to run from design through vessel service.
Seven days earlier, manufacturing software company Parsec had published results from a February survey of 1,200 manufacturing leaders. Seventy-two percent said their organizations had adopted AI in some form. Ten percent said they had deployed it at scale across operations.
The distance between those figures is 62 percentage points.
It is simple subtraction, not a standardized maturity score. Parsec sells manufacturing operations software, and the survey relies on executive, operational, and technical self-reporting. It did not audit plants, publish a productivity counterfactual, or show results by country and manufacturing subsector.
The two releases describe different things: one is an intended shipyard architecture; the other is a cross-industry adoption survey. They meet at a management problem. Buying access to AI is different from making it dependable when an unfamiliar batch arrives, a sensor drifts, a night shift has fewer specialists, maintenance is behind, or a worker hears something in a machine that the model has missed.
The Siemens announcement did not name a U.S. deployment site or disclose the exact contract value, training budget, workforce count, production baseline, or measured result. Those omissions leave the operating work open: connecting old and new equipment, deciding which production decisions a model may influence, training the employees who will challenge it, and paying for review and recovery after rollout.
One budget buys models, sensors, integration, and software. Another pays for changed jobs, supervised practice, shift coverage, and decision authority. Parsec did not test whether fragmented budgets caused its reported gap. But a manufacturer that approves the technology line while scattering the operating costs across maintenance, quality, IT, production, and training has no single owner for the passage from pilot to shift.
July 23 puts nine figures behind a digital shipyard
The Siemens and HD Hyundai agreement begins with a capital commitment and a product architecture. Its harder work starts when the architecture reaches a live yard.
A ship exposes the limits of a clean software-deployment analogy. Design changes can affect materials, suppliers, sequencing, fabrication, testing, and later maintenance. Work moves between engineers, planners, skilled trades, inspectors, contractors, and program owners. A late configuration error can create physical rework rather than a corrected screen.
Siemens said the digital thread would connect design, engineering, manufacturing, production, planning, simulation, shipyard execution, and lifecycle management. Each connection changes who sees information and when.
An engineer may receive production feedback earlier. A planner may simulate a sequence before it reaches the floor. A supervisor may see a constraint that used to live in a separate system. A quality team may trace a defect to a design revision or supplier record. A worker may read updated instructions on a connected device rather than wait for a paper package.
Siemens described the collaboration as a reusable blueprint for earlier decisions and fewer disruptions during construction. Those are proposed benefits. The release did not report a completed U.S. pilot, reduced rework, or a workforce result, so a buyer would still need to identify the first yard, baseline workflow, affected employees, and evidence required for expansion.
Delivery pressure is arriving on a separate track. On July 17, the U.S. Department of Transportation announced a $2 billion investment for two military ships at Hanwha Philly Shipyard, with the first vessel scheduled for delivery by June 2030. Hanwha and HD Hyundai are different groups, and public records do not connect the order to the Siemens agreement. The juxtaposition matters only as a sector timetable: ship orders and digital programs are advancing while the workforce to execute them develops on a slower schedule.
The U.S. Navy’s May 2026 shipbuilding plan describes the same operating connection. It calls for a digital-first data environment linking shipyards, suppliers, and program offices. Workers would use dynamic 3D models, digital twins, tablets, or connected devices at the workstation. The plan also calls for workforce programs using virtual reality trainers, computer-controlled machinery, and advanced manufacturing techniques.
The plan requests a $19.3 billion increase in the PB27 budget over PB26 for the shipbuilding industrial base. That figure covers a broad set of industrial-base investments, not a disclosed training allocation. Infrastructure, wages, supplier development, workforce programs, and technology all compete inside the plan.
A digital thread may reduce rework only if data remain synchronized and people act on the right revision. A tablet shortens the route to an instruction but raises the consequence of serving the wrong configuration. A simulation can improve sequencing while giving planners and supervisors a new responsibility: recognizing when its assumptions no longer match the yard. The capital purchase moves these duties into different hands; it does not fund or staff them automatically.
Seventy-two percent adopt; ten percent scale
Parsec’s 2026 State of Manufacturing release separates manufacturers into several stages. Ten percent reported AI at scale, 22% were actively implementing it, and the rest of the 72% adoption group remained in pilots or early use. Twenty-eight percent had not started.
The survey also found that 65% had begun adopting generative AI, up from 48% in Parsec’s 2024 report. Adoption was broad enough to create urgency: 60% of respondents worried more about moving too slowly than moving too aggressively.
Scale depends on the systems beneath the model. Sixty-nine percent of respondents said their factories operated a hybrid mix of legacy and modern equipment. Thirty-seven percent reported a unified, data-driven strategy, while another 60% were implementing or planning one.
These categories can overlap in practice. A company can have a corporate data strategy while one plant still extracts records from an old controller. It can run a successful visual-inspection model on one line while maintenance data remain incomplete elsewhere. It can buy a generative AI assistant for engineers without allowing the assistant to touch production instructions.
The reported barriers make that difference visible:
| Barrier to wider AI adoption | Share of Parsec respondents |
|---|---|
| High implementation cost | 40% |
| Data privacy and security | 39% |
| Integration with existing systems | 38% |
| Internal skill gaps when integrating new systems | 45% |
The skill figure appears in Parsec’s labor section rather than its ranked AI barrier list, so it should not be merged into that ranking. It describes a related implementation problem: nearly half of respondents saw internal capability as a major challenge when new systems entered the plant.
Quality control was the most common AI use case, cited by 50%, followed by IT operations at 46% and supply chain management at 45%. Each use carries different operating dependencies.
A visual-inspection pilot can begin with archived images and a narrow defect class. Scaling it may require stable lighting, camera maintenance, traceable labels, rules for ambiguous output, and a quality inspector authorized to stop the line. The model’s accuracy in a test set is one input. False escapes, false rejects, inspection time, scrap, customer claims, and reviewer load determine whether it improves production.
Predictive maintenance has another dependency. A model can rank likely failures from sensor history. The plant still needs sufficient data from the equipment, a technician who can interpret the signal, parts and time to act, and a process for recording whether the recommended intervention was correct. If a sensor has drifted or maintenance codes are inconsistent, the system may learn the plant’s recording habits instead of its failure modes.
Supply chain AI reaches outside the site. A system can estimate delay risk or recommend a supplier action, but procurement and operations must know which data are current, which contract terms matter, and how to handle a recommendation that protects inventory while increasing cost or quality risk.
A license count cannot distinguish these situations. A pilot can optimize for model performance; a production deployment has to work inside the plant’s maintenance, quality, staffing, and recovery routines.
Nor does the gap mean that 62% of manufacturers failed. Some programs are new, some use cases should remain local, and some plants should pause because integration costs exceed the expected gain or the process lacks stable data. What matters is whether each pilot has an explicit path to a production decision. Without one, “pilot” becomes a permanent category that protects a promising demo from an operating test.
More than eighty jobs across seven functions
The World Economic Forum published its Human-Machine Collaboration Framework on June 23. It maps more than 80 jobs across product development, planning, production, maintenance, quality, logistics, and supply chain management.
The Forum expects three in four of those jobs to evolve over the next decade. It classifies about 40% of future industrial skills as new or emerging and places jobs into four categories: elevated, expanded, emerging, and consolidated.
Several of its proposed roles sound unfamiliar because they combine operational work with oversight of automation: Supply Chains Intelligence Analyst, Quality Automation Technician, Control Tower Governor, Autonomous Logistics Specialist, Autonomous Warehouse and Fulfilment Operator, and Robotics Engineer or Orchestrator.
The names are planning devices, not observed labor-market categories. The framework draws on more than 40 consultations, over 10 workshops, and visits to Global Lighthouse Network sites. It does not show that employers have hired these roles at scale, that older roles disappeared, or that a particular job title will spread between countries.
Its value lies in the workflow view. A company can look at the work changing around a system before deciding whether the change deserves a new title.
Consider quality automation. A model may handle the first pass over images or sensor readings. The quality inspector’s role can expand toward threshold design, exception review, root-cause investigation, and model-performance monitoring. A technician may own camera calibration and data integrity. An engineer may decide when a process change invalidates the old model. A manager may own the cost of false rejection and the customer risk of a missed defect.
Calling all of that “AI skills” would be too vague to train or staff. The job map needs to specify the decision, the tool, the production context, the evidence required, and the authority granted.
The National Institute of Standards and Technology offers a second taxonomy. Its June 2 Manufacturing USA competency analysis identifies 132 entry-level occupations connected to 235 knowledge, skill, and ability items. The researchers organize the material into 13 competencies and 68 subcompetencies across biomanufacturing, digital and automation, electronics, energy and processes, and materials.
NIST’s framework uses 2025 data and looks at advanced manufacturing through 2030. It is not an AI adoption study. It gives employers and training providers a shared language for the capabilities attached to entry-level work.
The Forum begins with future workflows and asks how jobs change. NIST begins with occupations and asks what a worker must know and demonstrate. Used together, the two views let a manufacturer trace a moving task into a trainable capability without pretending that either framework settles the staffing decision.
A maintenance technician working with a predictive system may need the same mechanical and electrical foundation as before, plus sensor-data interpretation, model-limit recognition, and a documented override process. A supervisor may need enough statistical judgment to distinguish normal variation from a change that deserves escalation. A production engineer may need to understand how a data pipeline and a physical process fail together.
These combinations resist a simple “old role versus new role” story. An existing role may absorb exception review; two jobs may meet at a shared control point; a specialist may become necessary once review volume or system complexity grows. Routine checks can decline even as troubleshooting consumes more skilled time.
Kiva Allgood, a World Economic Forum managing director, framed industrial AI around the multiplication of human judgment. Vidya Gubbi, Western Digital’s chief of global operations, emphasized investment in the next generation and continuous upskilling.
For a plant manager, those aspirations become a schedule: which shift gets practice time, who observes it, and what evidence shows that an employee can act safely when the automated path is wrong.
Hiring pressure near supervision and maintenance
July hiring data capture a different kind of pressure: manufacturing employers on the ICIMS platform were posting more openings for roles that keep production running.
ICIMS analyzed activity from more than 3 million users on its platform for its July 2026 workforce report. U.S. job openings across its data were 19% above the prior year, while hiring remained roughly flat for a third month. Application volume was 5% below the June 2025 baseline.
Manufacturing demand was especially strong for several roles. Openings for first-line supervisors of production and operating workers rose 59% year over year. Industrial engineers rose 39%. General maintenance and repair workers rose 30%.
The figures describe ICIMS customers, not every U.S. employer. Job openings are not hires. The report does not establish that AI caused the increases.
The mix still complicates a common automation assumption. ICIMS manufacturing employers continued to seek people who coordinate shifts, repair equipment, and translate process requirements into operating changes. Those jobs are relevant to AI scale because they sit close to the point where a model recommendation becomes a physical action; the data do not show that AI created the openings.
Trent Cotton, ICIMS’ head of talent insights, described employers as making narrower bets on jobs tied to growth and operations. In manufacturing, that can raise the cost of leaving the workforce plan until after a pilot.
Parsec found that IT and technical specialists were the hardest roles to fill, cited by 60% of respondents. Quality assurance staff followed at 49%, then management roles at 39%. A factory that plans to scale quality inspection AI may already be short of the people needed to integrate, challenge, and supervise it.
ManpowerGroup’s July 14 manufacturing workforce outlook adds another view. Its research reports that 72% of manufacturers have difficulty finding skilled workers. Its accompanying commentary by Louise Ramstedt says more than half of manufacturing workers received no skills training during the prior six months, while nearly half worried that AI or automation could replace their role within two years.
ManpowerGroup is a workforce provider, and the report aggregates several surveys and field periods. Worker concern is not evidence of displacement. Employer difficulty is not a count of unfilled jobs.
Together, the surveys expose an awkward operating condition. Manufacturers report scarce skills while many workers report no recent training. Anxiety can rise just as the plant needs employees to take on more technical responsibility.
An employer can make that conflict worse by speaking about AI as a headcount program while asking technicians and inspectors to supply the judgment that makes it work. The worker hears replacement. The implementation plan assumes cooperation, exception reporting, and knowledge transfer.
Operations leaders need to make the employment bargain explicit. If AI changes a role, the plan should identify which tasks decline, which responsibility grows, how the employee learns it, who reviews early work, and how competence affects progression and pay.
Training cannot be limited to the people building models. The supervisor deciding whether to pause production, the technician deciding whether an alert indicates a real failure, and the inspector reviewing an uncertain defect all sit inside the AI system’s practical boundary.
The scale gap can therefore become a staffing loop. A company delays training because a pilot has not proved value. The pilot cannot prove value because the people needed to run it across shifts have not been trained or assigned. The project remains local, and management records another adoption without scale.
Breaking that loop requires a smaller claim than “transform the workforce.” Choose the next production decision, identify the roles around it, and create enough supervised practice to test whether the process works.
Everyday AI confidence stops at the factory stack
Autodesk’s 2026 AI Jobs Report examines roughly 4 million global job postings across architecture, engineering, construction, product design, manufacturing, media, and entertainment. GlobalData collected the postings across three rolling May-to-April periods.
AI-related roles across these Design and Make industries grew 147% over two years and 33% in the latest year. The fastest-growing titles shifted toward applied and creative work. AI UX Designer grew 145%, AI Creative Technologist 123%, and AI Consultant 90%.
The headline spans several industries, including media and marketing. It should not be presented as a manufacturing-only hiring rate. Autodesk and GlobalData use a report-specific definition of AI-related roles, and job listings do not prove that employers filled the positions.
The report’s workforce survey is more directly useful for the factory scale gap. In May, Autodesk surveyed more than 1,000 U.S. students ages 14 to 23 and over 500 Design and Make professionals.
Eighty-two percent of students felt confident with everyday AI tools such as ChatGPT and Claude. Thirty-six percent felt confident with AI tools specific to their field. Among professionals, the comparison was nearly 80% versus 49%.
Students understood that field knowledge mattered. Sixty-five percent ranked field-specific AI skills as important for landing a good job, ahead of the 46% who selected general AI tools. Yet 56% were unsure that they were learning the right AI skills, and fewer than 10% felt ready for emerging jobs in their fields.
Practical exposure was thinner. Eighty percent said they were developing job-relevant skills through self-teaching on platforms such as YouTube. Nineteen percent reported gaining real-world experience through internships or projects.
Self-teaching can explain a tool. It cannot reproduce the consequences of a live production decision. A learner needs access to equipment, representative data, safety rules, experienced review, and the chance to see a recommendation fail.
This does not mean every student needs a factory internship. It means employers and training providers should distinguish general AI familiarity from competence in a particular work system.
A person can write a strong prompt and still be unable to diagnose a vibration pattern, interpret a tolerance, recognize a process hazard, or understand which revision controls a build. General tools can help someone learn those domains. They do not replace the supervised work that turns information into judgment.
The student findings also challenge the idea that young people have abandoned physical work. Sixty-six percent said they wanted careers where they make things or work with their hands. Fifty-two percent found physical-world design and building careers appealing as AI changes work, compared with 23% who preferred mainly digital or online careers.
There is interest in the work. The gap lies between that interest and a credible path into the field-specific stack.
One federal program shows how a yard can put a price on that transition. The U.S. Maritime Administration’s FY 2026 Small Shipyard Grant Program has $35 million appropriated for capital improvements and maritime training. Eligible yards have no more than 1,200 production employees, and federal funds can cover up to 75% of a project.
The program asks applicants to quantify expected benefit in hours saved, dollars, percentages, or another meaningful measure. It also allows preference for projects using innovative technology to improve efficiency, safety, or resilience.
As of July 29, the page does not list 2026 award recipients or completed results. The program serves here as a budgeting model, not proof that training has closed the gap.
Equipment and training can appear in the same application. The applicant has to name the operational benefit and provide a method, timeline, cost, and evidence. Manufacturers scaling AI need the same discipline even when they are spending private money.
A pilot-to-shift scale map
A scale decision should follow the production decision, not the software catalog.
The table below is an operating template. It does not prescribe universal thresholds. A plant should replace each example with its own baseline, safety requirements, labor agreements, customer obligations, and economics.
| Use case | Legacy and data dependency | Roles and supervised practice | Acceptance evidence | Scale or stop trigger | Budget owner |
|---|---|---|---|---|---|
| Visual quality inspection | Camera position, lighting, defect labels, product revision, disposition records | Inspector reviews uncertain cases; technician maintains the image path; quality engineer tests new defect classes | False escapes, false rejects, review time, scrap, customer claims, traceable disposition | Scale after stable performance across products and shifts; pause after drift, untraceable labels, or material escape | Quality leader with plant operations and IT |
| Predictive maintenance | Sensor coverage, calibration, failure history, work-order codes, parts records | Maintenance technician compares alerts with physical condition; reliability engineer reviews failure logic | Avoided downtime, unnecessary interventions, lead time, missed failures, technician override record | Scale when alerts create enough actionable lead time; stop when bad data or false alarms consume the gain | Maintenance or reliability leader |
| AI-assisted work instructions | Approved design revision, configuration control, device access, language and accessibility | Operator performs the task under trainer observation; supervisor verifies revision and exception route | Completion quality, rework, instruction errors, time to proficiency, worker-reported ambiguity | Scale after workers can identify the current revision and recover from an exception; pause on version mismatch | Production leader with engineering and training |
| Production scheduling | Order data, routing, labor availability, machine constraints, maintenance windows | Planner tests recommendations against known constraints; supervisor records reasons for override | Schedule adherence, changeovers, overtime, work in process, late orders, override causes | Scale after the model survives demand and equipment variation; stop if gains depend on hidden overtime or unstable inputs | Operations planning and finance |
| Autonomous material movement | Facility map, traffic rules, load data, equipment telemetry, pedestrian zones | Operators and safety staff rehearse normal and degraded modes; maintenance owns recovery | Near misses, blocked routes, manual interventions, cycle time, damage, recovery time | Scale only after safe recovery works on every shift; stop on repeated blind spots or unclear authority | Logistics leader with safety and facilities |
| Supplier and inventory risk | Supplier master, lead times, quality history, contract terms, demand forecast | Buyer reviews high-impact recommendations; operations tests shortage and substitution scenarios | Forecast error, expedite cost, stockouts, excess inventory, supplier-quality result | Scale when recommendations improve total operating outcome; pause when opacity blocks contractual or quality review | Supply chain leader with procurement and finance |
Before it governs a rollout, each row needs six plant-specific fields: a baseline, a target, an observation window, one accountable individual, an escalation time limit, and a change that triggers reassessment. “Stable performance” might mean two product families across three shifts for eight weeks in one plant; another site may require a longer run because a failure is rarer or more severe. The plant should write the threshold before reading the pilot result.
The acceptance record also needs signatures from the function owner, shift supervisor, trainer, and employees doing the changed work. Where a union or works council represents those employees, its role belongs in the rollout plan rather than in a consultation after the software arrives. Practice time should be paid and covered on the schedule. The record should state how a good-faith override affects performance review, discipline, and liability when throughput, safety, and quality point in different directions.
The map makes hidden labor visible: inspector review for a quality model, technician judgment for predictive maintenance, supervisor overrides for scheduling, and safety practice for an autonomous vehicle. Each row needs a baseline before the pilot. Inspection starts with existing escape and rejection rates, review time, scrap, rework, and claims. Maintenance starts with downtime, failure lead time, unnecessary work, parts delay, and missed failures. Scheduling starts with adherence, overtime, changeovers, inventory, and late orders. Without those measures, a project can report model accuracy or user activity without showing whether production improved.
Training evidence should also match the decision. Course completion can show exposure. A supervised task shows application. Repeated correct action across normal and exception conditions gives stronger evidence. A named assessor, using the same rubric across shifts, should decide when an employee can work independently. The training plan needs protected practice hours, enough qualified trainers to cover every shift, and recertification after a material change to the model, sensor, process, or instruction. NIST’s competency language can organize the curriculum; the plant’s workflow determines the practical assessment.
Shift coverage is part of the test. A pilot that works only when the vendor engineer and the plant’s best specialist are present has proved a narrow operating condition. It has not proved that the system can run on Tuesday night, during vacation coverage, or after the product mix changes.
The financial case should include the human system. Over a named period, the plant can compare cashable operating gains and risk-adjusted avoided losses with one-time integration and changeover costs; recurring model, data, security, and support costs; paid training and backfill; review and recovery work; and the cost of replicating the change on another line, shift, or site.
That is a cost ledger, not an accounting formula. Finance should separate one-time and recurring items, set the payback hurdle before the pilot closes, and name the operating leader responsible for realizing each benefit. Quality risk may be better represented by scenarios than by one expected value. The discipline prevents training, review, and recovery from being treated as free.
The map also creates a stop decision. A good pilot can reveal that a use case should stay local, wait for better data, or end. Stopping is not evidence that the company is behind. Scaling a brittle system across plants can be far more expensive than leaving it in one bounded workflow.
Ownership should stay close to production. IT can own platforms and integration. A data team can own models. The plant function affected by the decision must own acceptance. Quality accepts inspection changes. Maintenance accepts failure interventions. Operations accepts schedules. Safety owns the conditions under which autonomous equipment may run.
Workers need a place in that decision. They see sensor failures, workarounds, ambiguous instructions, and conditions that were absent from training data. A reporting channel, response deadline, and appeal path should be visible on every shift. Reporting a defect or making a good-faith override should improve the system, not become evidence that an employee resisted adoption.
At the end of each pilot shift, the crew can record which recommendations were used or overridden, what took longer, which skill was missing, and whether the model exposed an old process problem or created a new one. That shift log becomes the evidence for the next scale review.
A night-shift handoff after the pilot
Consider a planning scenario from the first week after a predictive-maintenance system leaves its pilot cell.
At 1:40 a.m., the system flags a rising failure risk on a machine that is still producing within tolerance. The model recommends maintenance before the next scheduled changeover. The night supervisor has an order due in the morning. Taking the machine down could delay it.
A technician still has to make the decision. She checks the sensor history and notices that the reading changed after a recent replacement. She compares it with vibration at the machine, reviews the work order, and sees that the new sensor was installed with a different calibration setting. Yet vibration is also slightly higher than its usual range. The measurement path may be wrong, the machine may be changing, or both may be true.
If the technician has access to the data path, understands the model’s input, and has authority to challenge the recommendation, she can escalate the uncertain case instead of treating the alert as an instruction. A workable policy names the engineer or maintenance lead who must respond, the time allowed, the production cover available while the case is checked, and the person authorized to stop the machine. Only a later review can determine whether the event calls for calibration repair, maintenance, a model change, or no change at all.
If she has only been trained to acknowledge alerts, the supervisor faces a worse choice: stop production on an uncertain recommendation and miss the morning order, or keep running without knowing whether the risk is real. A throughput target that punishes every cautious stop silently decides the issue before the alert arrives.
No job disappeared and no new title appeared, yet the employee’s access, judgment, and authority now determine whether the system is safe to use.
The pilot may have produced a strong accuracy score. The scaled system depends on a technician who combines machine knowledge, sensor judgment, data access, and escalation authority on a shift when the pilot team is not present.
Parsec’s 72% adoption figure can include the first model. The night-shift scenario illustrates the repeatability implied by its 10% at-scale category.
The Siemens and HD Hyundai agreement may eventually connect ship design, planning, production, and lifecycle data inside a digital yard. The public announcement gives the software architecture and capital signal. It leaves the shift-level apprenticeship open.
Before expanding the next pilot, a plant leader can ask one concrete question: who handles the 1:40 a.m. recommendation when the model is plausible, production is waiting, and the specialist who built the demo has gone home?
The answer identifies the integration work, the training, the authority, and the budget that stand between adoption and scale.