# One Employer Reached 141 AI Solutions. Another Kept Three.

> MIT researchers found 141 organization-wide AI solutions at one employer and three at another. The gap points to protected time, shared review, technical support, and career credit.

- Published: 2026-09-10
- Author: Gene Dai
- Canonical: [https://digidai.github.io/2026/09/10/ai-sandbox-141-solutions-hidden-work/](https://digidai.github.io/2026/09/10/ai-sandbox-141-solutions-hidden-work/)
- Topics: Enterprise AI, Organizational Design, Future of Work, Employee Experience, AI Skills

---

On September 9, MIT Sloan published a comparison that should make every enterprise AI steering committee reopen its
pilot list.

One organization had 141 organization-wide AI solutions in use. Another had three. In the researchers' framing, these
were applications intended to create value for a broader set of organizational stakeholders, not merely private prompts
saved by individual workers.

Both had given employees a secure sandbox in which to test generative AI. Both had provided training. Both encouraged
domain experts to develop applications that could be used beyond one person's desk. The two-year field study covered an
academic medical center, identified as NE Health, and a corporate law firm, identified as LegalCo.

The result was not a simple contest between medicine and law. According to
[MIT Sloan's account of the research](https://mitsloan.mit.edu/press/why-some-organizations-turn-ai-experiments-business-value-while-others-quietly-fail),
the researchers reported no meaningful difference in the organizations' AI readiness, access to technology, regulatory
constraints, or suitability of the problems being considered. Their explanation for the divergence centers on the
support around the people doing the experiments. The public material does not provide a causal estimate for that
support.

NE Health offered continuing education, shared evaluation methods, technical help, risk screening, integration support,
and career recognition. LegalCo's support was thinner after initial training. More than 80 percent of its participating
domain experts eventually disengaged from the innovation work, the MIT release says. Many initiatives were scaled back,
leaving three organization-wide solutions in use.

Those counts demand attention, but they remain incomplete.

A solution "in use" is not necessarily a profitable solution, a safe solution, or a solution that improves a patient or
client outcome. The release does not name the organizations, list all 144 solutions, report their active users, disclose
their costs, define a common value threshold, or provide a quality-adjusted return. A portfolio of 141 small workflow
aids could create less value than three deeply embedded systems. The study is a qualitative comparison, not a randomized
trial.

Its practical value lies elsewhere. It shows the labor between giving someone an AI account and maintaining a tool that
colleagues can trust. Domain experts have to find a useful problem, test where a model fails, assemble examples, ask
colleagues to review outputs, satisfy risk owners, connect the tool to a real workflow, explain it to users, and revise
it when the model or policy changes. That work can be treated as part of the job. It can also be left for evenings,
lunch breaks, and whatever time an enthusiastic employee can hide from a full calendar.

Leaving that labor with volunteers looks cheap right up to the week they stop volunteering.

An executive choosing an AI platform can sign a contract and provision seats. Turning employee experiments into shared
capability requires a capacity decision: whose time will be protected, who will review the work, what evidence is good
enough to proceed, and how the contribution will count in pay and promotion.

## September 9 put 141 beside three

The underlying paper has a more careful title than the adoption headline:
["Experimentalist Intensification Governance: Managing Worker Negative Consequences Associated with Generative AI Innovation Work"](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7088100).
Arvind Karunakaran of Stanford, Katherine Kellogg of MIT Sloan, and Batia Wiesenfeld of New York University wrote the
54-page working paper in May 2026. It was posted to SSRN in July and remains a preprint rather than a peer-reviewed
publication.

Its subject is not ordinary individual use, such as asking a chatbot to draft an email. The researchers examined medical
and legal workers using generative AI to create organizational solutions. Such a solution must survive contact with
colleagues, standards, systems, and people who did not build it.

The paper's abstract names three forms of experimental work: distributed trialing, collective review and revision, and
aligning and hustling. The short definitions used here are interpretations of those labels. People explore possible uses
and limits across their areas of expertise, recruit other experts to judge and improve results, and gather the support
needed to move a promising idea through the organization.

Those phrases describe work, not software features. A sandbox can make experimentation safer by isolating tests from
production systems. It cannot clear a clinician's calendar, persuade another department to review examples, decide
whether a mistake is tolerable, or secure an integration owner. A model can produce another version in seconds. It
cannot make a colleague accept responsibility for that version.

MIT's release adds the longitudinal result. At LegalCo, participation did not end with a public refusal. People
gradually stepped away. The extra work remained burdensome, so the everyday trial, review, and refinement slowed. The
firm eventually had three organization-wide solutions in use.

NE Health moved in the other direction. More people joined the experimentation over time, and the organization reported
141 solutions in use with more under development. Its teams included work on patient-friendly discharge summaries.
LegalCo's participants included work on legal research.

Those examples make the comparison tangible without making the sectors equivalent. A discharge summary and a legal
research product carry different error consequences, review practices, data rules, and users. The researchers report
that the organizations did not meaningfully differ on the broad constraints they assessed. That finding does not erase
every difference between the tasks or institutions.

An anonymous field study also limits outside inspection. A reader cannot compare the two organizations' budgets,
staffing, leadership changes, existing digital systems, or exact definitions of "solution." The public release gives the
portfolio count and reported support practices, but not a solution-level dataset. The contrast supports the researchers'
account of how the structures operated in these cases. It is not a universal conversion rate for sandboxes.

This boundary matters for buyers. The wrong response is to copy 141 into a target and ask every business unit for more
ideas. The useful response is to ask what employees had to keep doing for an idea to remain alive long enough to become
shared work.

That question changes the unit of adoption. Seats measure access, and prompts record activity. Neither tells a leader
whether a test survived evaluation, whether a released tool acquired an owner, or whether continued use produced a
business outcome worth its maintenance.

Once those states appear in one dashboard, 141 becomes the start of an audit, not the last slide in a celebration.

## Two secure sandboxes began from similar ground

Secure sandboxes solve a real problem. Employees need somewhere to test a model without casually placing confidential
material into an unmanaged consumer tool or changing a production workflow. Isolation can control data access, preserve
logs, restrict integrations, and give technical teams a place to observe behavior.

The two organizations in the study had that starting point. They also provided training and invited employees to develop
uses with organization-wide potential. That is already more structure than an employer that merely reimburses a chatbot
subscription.

Yet a sandbox contains the experiment, not the employee's workload.

Suppose a clinician wants to make discharge instructions easier for patients to understand. The first demonstration may
be quick. The organizational solution is slower. Someone must decide which clinical facts may be simplified and which
must remain exact. People must collect representative examples, including difficult cases, and compare the generated
text with the source record. Reviewers must agree on what counts as a harmful omission. A technical team has to
determine where the tool receives information and where the result appears. A workflow owner must specify who signs off
before a patient sees it.

Now consider a legal research tool. A plausible answer can still miss controlling authority, use a superseded source,
confuse jurisdictions, or state a proposition more confidently than the evidence allows. Lawyers must design the tests,
inspect citations, argue about acceptable scope, and decide which matters are suitable. They also need a route for
withdrawing a result when the evidence is weak.

The model produces text in both settings. Domain experts make the text usable in each one.

Initial training can teach people where to click, how to frame an instruction, and why confidential data requires care.
It rarely settles the local questions discovered during use. How recent must a source be? Which exceptions matter? Who
owns the final wording? What happens when a platform update changes a previously stable output? When does an assistant
become part of a regulated or professionally accountable process?

The demonstration usually ends before these questions appear, and the answers sit across organizational boundaries.

An employee experimenting alone can choose an example, judge the draft, and try again. An organization-wide solution
needs a shared definition of quality. It also needs technical access, risk acceptance, documentation, release support,
user help, and continuing ownership. Every additional dependency makes the experiment less likely to fit inside spare
time.

That is why the MIT comparison should not be reduced to "healthcare supported AI and lawyers resisted it." LegalCo's
experts participated. More than four-fifths later disengaged from innovation efforts, according to the release. Quiet
withdrawal is compatible with interest in the technology. It can be a rational response to work that carries no
protected capacity, visible credit, or reliable route through institutional review.

The distinction affects employee surveys. A worker may say that AI is useful and still stop contributing to a shared
project. A manager may see regular chatbot use and assume that innovation is spreading. Both observations can be true
while organization-wide development is shrinking.

An adoption survey therefore needs at least two questions. Are employees using AI for their own tasks? Are domain
experts still willing and able to help build, evaluate, and maintain tools for other people?

The capacity problem is hiding in the answer to the second question.

It also separates three employees often collapsed into one adoption rate. The user gets help with a task. The
experimenter tests whether a shared tool could work. The reviewer lends judgment and may never touch the final
interface. A manager can have high user adoption while the experimenter and reviewer pools are shrinking. Reporting only
logins makes their withdrawal invisible.

## Domain experts paid the experimentation bill

Organizations often describe domain experts as a source of ideas. The role is heavier than ideation.

A useful expert can identify a task where the model might help because that person understands the exceptions, failure
costs, and existing workarounds. The same knowledge makes the expert necessary during evaluation. When the first output
looks convincing, the expert is the person most likely to see what it missed.

The result is a recurring bill paid in hours.

Trialing consumes time before anyone knows whether the idea will work. Example selection consumes time because easy
cases make a weak test. Review consumes time from colleagues whose agreement gives the result credibility. Alignment
consumes time with security, legal, technology, data, operations, and business owners. Revision begins again when a
model, policy, source system, or user population changes.

None of those hours appears in the license price.

They may not appear in the project plan either. An innovation program can invite voluntary proposals, celebrate a
promptathon, and count prototypes while leaving participants responsible for their normal targets. The most
conscientious employee then faces a choice. Cut the evaluation short, let ordinary work slip, or absorb the experiment
outside recorded hours.

The SSRN abstract calls the result work intensification. That phrase is important because it directs attention to
consequences for workers, not only to lost innovation. Extra effort can be temporarily exciting. People may enjoy
solving a problem and learning a new tool. Repeated invisible effort is harder to sustain, especially when other
colleagues receive credit for measurable business work while experimentation remains classified as enthusiasm.

The burden is not evenly distributed. Experienced employees are more likely to be asked whether an output is correct.
People who already perform informal coordination may become the bridge between a prototype and several departments.
Employees from underrepresented groups may be asked to review bias or accessibility without having that responsibility
recognized. A high performer can attract more experimental work precisely because the person is trusted to rescue it.

Early-career employees face a different risk. They may receive access to the sandbox but lack authority to recruit
reviewers or challenge a senior sponsor's preferred result. If only senior experts perform the consequential checks,
junior workers also lose the supervised practice through which they could learn those judgments. A capacity plan must
preserve apprenticeship, not merely protect the current experts.

Managers may not see the accumulation. Each request looks small: review ten outputs, attend one workshop, introduce the
team to security, explain one exception, rerun a test after a model change. Across several projects, the expert becomes
a shared dependency with no queue and no budget.

The bottleneck can produce misleading conclusions. When a prototype stalls, leaders may call the tool immature or the
business unit resistant. Sometimes it is waiting for two protected hours from the person qualified to decide whether the
result is safe. Technology readiness and organizational readiness meet inside that calendar invitation.

Career systems can make the problem worse. A law firm may reward client work, revenue, and recognized professional
contributions. A medical center may recognize publications, improvement projects, or formal innovation roles. The MIT
release says NE Health incorporated AI experimentation into job responsibilities, performance evaluations, promotions,
publications, and career advancement. At LegalCo, participants often reported that the AI work remained invisible in
reviews and compensation.

Line managers sit inside that conflict. They can release an employee for a workshop while leaving the employee's monthly
target unchanged. The calendar then records support, but the workload still charges the time back. Protected capacity
means adjusting priorities or staffing, not placing a learning block beside the same volume of expected work.

This is not a recommendation to reward every prototype. It is a reason to record the work required to reject one
responsibly. An expert who discovers that a model fails on a critical case may create more value than a colleague who
produces an attractive demonstration. If performance systems credit only launches, employees learn to hide negative
evidence or avoid difficult evaluations.

A fair recognition rule should therefore cover at least four contributions: finding a viable problem, documenting a
failure, improving an evaluation method, and carrying a safe solution into use. It should also distinguish contribution
from ownership. A reviewer deserves credit without inheriting permanent maintenance for someone else's project.

The bill becomes clearer when leaders stop calling all of this "AI training." Training is one input. Experimental work
is production work performed under uncertainty.

## Support turned private trial and error into shared work

NE Health's support did not consist of one central AI team taking every project away from employees. Its reported
practices helped domain experts continue participating.

NE Health kept education going through workshops and promptathons, maintained documentation, and assigned technical
experts to high-priority projects. An employee still had somewhere to take the second question after the introductory
course ended.

Shared evaluation rubrics replaced "looks good" with criteria another person could inspect. Knowledge forums let one
team's failure become another team's starting point. Together they reduced the rediscovery hidden inside private
experiments.

Formal risk screening gave projects a route through institutional concerns. Technical support helped connect approved
uses to existing systems and workflows. Review became a designed stage, while integration became an assignment instead
of a favor requested after the prototype attracted attention.

Experimentation also entered formal responsibilities and career processes. Employees could connect a contribution to
performance reviews, promotions, publications, and advancement. The institution asking for the work could finally see
it.

LegalCo, in the MIT account, offered initial training but less continuing support. That contrast helps explain why the
same word, "sandbox," can describe different operating systems.

| Support practice                    | Question it answers                                                    | Work it prevents from remaining private               |
| ----------------------------------- | ---------------------------------------------------------------------- | ----------------------------------------------------- |
| Ongoing workshops and documentation | Where does an employee take a new failure or changed model behavior?   | Relearning the same lessons alone                     |
| Dedicated technical experts         | Who helps a promising use become a reliable application?               | Unfunded integration and troubleshooting              |
| Shared rubrics and review forums    | Which outputs are acceptable, and who agrees?                          | Personal judgment disguised as organizational quality |
| Formal risk screening               | Which uses may proceed, pause, or require controls?                    | Last-minute escalation after effort is already spent  |
| Workflow integration support        | How does the solution enter real work without breaking accountability? | Manual copying and unofficial workarounds             |
| Performance and career recognition  | Why should an expert keep contributing after the novelty fades?        | Invisible effort outside the employee's evaluated job |

Few projects need a large committee. Support should follow consequence. A low-risk internal formatting aid may need a
light review and a named maintainer; a system affecting patient communication or legal advice requires deeper evidence,
professional judgment, and clear release authority.

Shared support can lower the cost of that distinction. A risk intake form can route a simple experiment quickly while
identifying sensitive data or consequential decisions. A reusable evaluation set can keep each team from starting with
five convenient examples. An integration queue can expose whether technical capacity, rather than model quality, is
stopping progress.

Central support also needs limits. If every decision waits for a small AI office, the organization replaces invisible
domain work with a visible central bottleneck. Domain experts still have to judge local meaning. Technical teams still
need priorities. Risk owners still need a proportional standard.

The people receiving the changed service need a route into that standard. A patient who cannot understand a summary, a
lawyer asked to rely on a research result, or an employee required to use a new internal assistant may see failures that
the project team missed. Complaint, correction, and opt-out signals belong beside model tests. Otherwise the
organization can scale a tool while shifting the repair work to its users.

A workable arrangement resembles a joint production line. Employees bring the domain problem and examples. A shared team
supplies tooling, documentation, evaluation help, and integration patterns. Risk specialists set routes based on
consequence; a business owner accepts the workflow change; a named maintainer watches the solution after release.

The budget question then becomes more useful: "How many experiments can our review, integration, and maintenance system
support without borrowing unpaid capacity?" A larger sandbox population says nothing about that limit.

The answer may be lower than the number of ideas. That is healthy. An organization can reject weak uses early, preserve
evidence about why, and direct scarce expert time toward the next candidate. A smaller, maintained portfolio is more
credible than a large prototype gallery.

It also protects the meaning of 141. Without lifecycle records, a solution count can grow while abandoned tools remain
on the list. Support is not only what helps a project launch. It is what lets the organization retire one honestly.

## Training access missed the time constraint

Evidence outside the MIT comparison points to the same gap between offering learning and creating capacity.

The UK Department for Work and Pensions and Skills England published
[research on AI upskilling](https://www.gov.uk/government/publications/skills-for-ai-what-works-for-ai-upskilling-in-the-uk/executive-summary-what-works-for-ai-upskilling-in-the-uk)
based on 23 workshops involving about 150 organizations, ten case studies, and an employer survey with 536 responses.
More than 44 percent of surveyed organizations reported daily AI use. The report says the main problem was not lack of
interest but capacity, including time pressure, cost, unclear provision, and training that was too generic or
disconnected from real roles.

The evidence base has limits. Workshop participants are not a statistically representative workforce sample. The survey
is organization-level self-report, and the methodology says respondents were recruited through a paid online platform
after screening for UK location and senior or decision-making roles. Results capture employers already engaged enough to
answer questions about AI. They do not measure the productivity of each training program.

Even within that engaged sample, access did not settle effectiveness. The accompanying
[employer guide](https://www.gov.uk/government/publications/skills-for-ai-what-works-for-ai-upskilling-in-the-uk/employer-guide-what-works-for-ai-upskilling-in-the-uk--2)
says 97 percent reported providing AI training, while 51 percent identified a gap in flexibility and accessibility and
34 percent identified a gap in practical, contextual learning. Thirty-five percent cited missing clear skills
frameworks, 29 percent insufficient ethics and governance, and 22 percent weak leadership or organizational support.

The guide explicitly includes paid or protected time in its criteria for reachable training. It also asks employers to
connect learning to real tasks, provide repeat practice, integrate it with systems and standards, refresh it as tools
change, and monitor whether confidence, quality, and decision-making improve.

That design closely matches the work revealed by the two-organization study. An introductory course creates a common
vocabulary. Continued practice discovers local failures. Peer review turns private judgment into a standard. Technical
and risk support connect learning to production. Protected time prevents participation from depending on who can donate
the most personal capacity.

The
[OECD's June policy brief on AI and skills](https://www.oecd.org/en/publications/ai-and-skills_f843b352-en/full-report.html)
provides a broader synthesis built partly from older evidence. It says 40 percent of surveyed employers in manufacturing
and finance cited skills as the main barrier to AI adoption, while more than half of workers using AI reported
employer-funded training. The brief estimates that fewer than one percent of workers need advanced AI-specific skills
such as programming or model development. For most people, digital capability, data interpretation, management, problem
solving, creativity, and innovation remain more relevant.

Those figures aggregate studies conducted at different times, including data that predated the latest generation of
models. They should not be used to estimate the current training rate at NE Health or LegalCo. They do make one
distinction useful: scaling AI does not mean turning the whole workforce into model developers.

Most domain experts need enough technical understanding to test a tool, recognize weak evidence, protect sensitive data,
and communicate with specialists. They need deeper knowledge of their own work to design a meaningful evaluation. The
organization needs a smaller group that can build integration, monitor performance, and update shared infrastructure.

Treating everyone as if they hold the same role wastes time, while a short literacy course alone leaves the difficult
work ownerless.

A practical learning architecture can use three layers:

1. All users learn permitted uses, data boundaries, output checking, and escalation.
2. Domain experimenters receive protected practice time, evaluation methods, peer review, and a path to technical
   support.
3. Builders, risk owners, and maintainers receive deeper access, implementation responsibility, and workload allocation.

Movement between the layers should be possible. An employee who becomes a strong evaluator may enter a formal expert
role. A builder may return a use case to local experimenters when the problem is poorly defined. Recognition should
follow the actual contribution, not the prestige of the most technical title.

This design also makes inclusion measurable. Who attends training during paid hours? Who gets invited to high-value
projects? Whose review labor is recorded? Who receives promotion credit? If experimental capacity belongs only to people
with flexible schedules or powerful managers, the program can reproduce workforce inequalities while claiming broad
participation.

Access is a procurement result. Capacity is an employment design result.

## A capacity plan for employee-built AI

An AI experimentation capacity plan begins before a promptathon and continues after release. It connects each stage to
work, ownership, time, evidence, and a decision.

| Stage                       | Hidden work                                                                                  | Named owner                                    | Paid or protected capacity                                | Decision evidence                                                         | Stop or scale signal                                                        |
| --------------------------- | -------------------------------------------------------------------------------------------- | ---------------------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------- | --------------------------------------------------------------------------- |
| Problem selection           | Observe the workflow, define the user, identify failure cost, and check whether AI is needed | Business process owner with a domain expert    | Scheduled discovery time, not an after-hours idea request | Baseline volume, delay, error, cost, and affected users                   | Stop if the problem lacks an owner, baseline, or acceptable use             |
| Controlled trial            | Assemble representative cases, test limits, document prompt and model conditions             | Domain experimenter                            | A fixed test block and access to a secure environment     | Test set, failure examples, data classification, and reproducible setup   | Scale testing only when important cases and known exclusions are documented |
| Collective review           | Recruit peers, compare judgments, resolve disagreements, revise criteria                     | Evaluation lead                                | Reviewer allocation with a queue and due date             | Rubric, inter-reviewer disagreement, corrections, and unresolved cases    | Stop if reviewers cannot define acceptable output or consequence            |
| Risk screening              | Assess privacy, security, legal, professional, labor, accessibility, and safety concerns     | Risk owner appropriate to the use              | Proportional review service with an escalation route      | Risk tier, required controls, prohibited uses, and accountable approver   | Proceed only when controls and residual risk have named acceptance          |
| Integration and release     | Connect systems, design human handoffs, train users, prepare rollback                        | Product or technical owner plus workflow owner | Engineering capacity and user-release time                | Acceptance test, access controls, workflow map, user notice, and rollback | Release when the workflow, not only the model output, passes                |
| Operation and maintenance   | Monitor use, investigate incidents, refresh tests, handle model or policy change             | Named maintainer with business accountability  | Recurring maintenance budget and expert review reserve    | Active use, quality sample, incident log, update history, support load    | Expand when value survives review cost; retire when use or quality falls    |
| Recognition and progression | Record contributions, distribute credit, avoid permanent volunteer labor                     | Manager and people leader                      | Review time and explicit goal allocation                  | Role expectations, contribution record, promotion or reward decision      | Continue participation when effort is visible, fairly allocated, and useful |

The plan makes several decisions harder to avoid.

Start a proposal with a baseline. If the team cannot describe the current volume, delay, error, or user problem, it will
struggle to show what changed. The measure can be small and practical, but it belongs in the record before participants
fall in love with the prototype.

Give the experiment a time budget. "Spend up to 20 protected hours over four weeks" is a capacity decision. "Explore
when you can" transfers that budget to the employee. The exact allowance will vary; writing it down exposes how many
parallel experiments a team can actually support.

Review depth should follow consequence. A formatting error in an internal note differs from a missing medication
instruction or an invented legal authority. Sample size, reviewer seniority, and release control should reflect the harm
a failure could create.

A release needs technical and business ownership. One person may hold both in a small project. The responsibilities
remain distinct: maintain the system, accept the changed workflow, and stay accountable for the service delivered to
patients, clients, employees, or customers.

Set retirement criteria for the portfolio. A project should leave the "in use" count if nobody relies on it, the owner
departs, the source system changes, quality falls, or maintenance exceeds the remaining value. Retirement can show that
the organization is managing a living portfolio instead of accumulating demonstrations.

Finance has a role before renewal, not only after the portfolio becomes expensive. The capacity plan lets a CFO compare
another software commitment with the review and integration queues already funded. It also makes a trade explicit:
adding ten experiments may require more domain release time, technical support, or risk-review capacity than the
operating plan contains.

Cost should include more than software and engineering:

`total experiment cost = domain hours + peer review + technical support + risk review + integration + user change + maintenance + opportunity cost`

The equation is a checklist, not a promise of precise accounting. Some benefits and opportunity costs will remain
estimated. Recording the categories still prevents a license invoice from masquerading as the full investment.

Benefits also need separate fields. Time saved, fewer errors, shorter cycle time, improved comprehension, increased
access, and employee experience are different outcomes. A project can improve one while worsening another. Gross minutes
saved should be reduced by review, correction, support, and maintenance time before leaders call the difference new
capacity.

Career evidence belongs in the same plan. Managers should record who originated the problem, who found important
failures, who built the evaluation, who carried the integration, and who maintains the result. That list can enter a
performance conversation without assuming that all contributions deserve the same reward.

Finally, the plan should limit unpaid dependency. No employee should become the permanent reviewer for a growing
portfolio because that person once volunteered. When demand exceeds allocated hours, leaders can add capacity, lower the
portfolio target, narrow the risk appetite, or stop projects. They cannot honestly call the queue free.

## Next quarter exposes which solutions lasted

The MIT comparison gives leaders a useful warning and an unfinished measurement job.

NE Health's 141 solutions show that employee-led experimentation can travel far beyond personal chatbot use when the
surrounding institution supports it. LegalCo's three show how quickly a portfolio can contract when people stop doing
the extra work. Neither count reveals enough to rank business value.

The next quarterly review should start with the portfolio itself. For every solution marked "in use," ask:

- Did anyone use it in the past 30 and 90 days?
- Which workflow owns it, and which technical person maintains it?
- How many outputs were accepted, corrected, rejected, or escalated?
- What user or business outcome changed against the baseline?
- How many expert, review, support, and maintenance hours did it require?
- Did an incident, policy update, or model change force rework?
- Which employees contributed, and where did that work appear in goals, workload, pay, or progression?
- Does the evidence support expansion, continued observation, redesign, or retirement?

These questions protect both sides of the comparison. They stop leaders from dismissing LegalCo's three without knowing
whether any are valuable. They stop leaders from celebrating NE Health's 141 without knowing whether they remain active,
maintained, and useful.

They also reveal where the next dollar belongs. If promising projects wait for integration, another platform license
will not clear the queue. If review disagreement is high, the organization needs better examples and criteria. If
maintenance absorbs the expected saving, the use may need redesign or retirement. If experts withdraw despite strong
results, workload and recognition deserve attention before another innovation campaign.

For employees, the quarterly review should answer a plainer question: did the organization make space for this work, or
did it merely notice who was willing to carry it? Participation data should be examined by role, level, team, and access
to protected time. A broad invitation is not broad opportunity when only some managers can lower ordinary workload.

Imagine the review meeting in December. The slide still says 141. The head of AI begins to move on.

A clinical leader stops the presentation at solution 37. Active use fell after a source-system update. Two nurses have
been correcting its output manually. The original builder changed roles, and nobody received maintenance time. The tool
saved minutes in its first month, but the current review burden is unknown.

No apology slide will repair it. Pause the solution, assign an owner, rerun the evaluation, count the correction time,
and decide whether the workflow still deserves support.

Then open the performance records. Find the people who noticed the problem, tested it, and kept patients from receiving
a weaker result. If that work is absent, the organization has repeated the mistake the research makes visible.

A secure sandbox can hold a model experiment. Only the employer can make room for the people who turn it into dependable
work.

## Continue reading

- [Botsitting Takes Back the AI Workweek](https://digidai.github.io/2026/07/06/botsitting-ai-workweek-operating-cost/): Measure the context, checking, and correction work that AI usage reports miss.
- [Inside Workday's $1.1 Billion AI Learning Bet](https://digidai.github.io/2026/07/23/workday-sana-ai-learning-bet/): Compare course access with practice, manager feedback, and career movement.
- [Inside NEC's AI Department, Where People Still Decide](https://digidai.github.io/2026/08/31/nec-ai-department-four-management-layers/): See where human decisions remain inside an AI-staffed department.
- [After 300,000 Copilot Seats, HR Needs the Audit File](https://digidai.github.io/2026/06/15/enterprise-ai-seat-rollout-audit/): Turn enterprise AI access into an evidence file for renewal.
