# Three AI Providers Reported Outages in One Workday

> OpenAI, Anthropic, and xAI reported overlapping AI outages on September 3. A work-continuity ledger shows where model failover and human recovery begin.

- Published: 2026-09-06
- Author: Gene Dai
- Canonical: [https://digidai.github.io/2026/09/06/ai-providers-one-workday-outage/](https://digidai.github.io/2026/09/06/ai-providers-one-workday-outage/)
- Topics: Artificial Intelligence, Operational Resilience, OpenAI, Anthropic, xAI, Business Continuity, Enterprise AI, Deep Investigation

---

At 13:26 UTC on September 3, Anthropic began investigating elevated errors across several Claude models. Four minutes later, xAI recorded the start of a model outage. At 14:43 UTC, according to a later account from OpenAI, a routing error began affecting ChatGPT and Codex.

Three major AI providers had acknowledged trouble inside the same 77-minute window. Anthropic reported that the impact ended at 16:16 UTC. xAI restored traffic at 17:05 UTC, three hours and 35 minutes after its incident began. OpenAI moved its incident through investigation, mitigation, monitoring, and resolution on a status record that named 15 ChatGPT components and four Codex components.

Those timestamps establish overlap. They do not supply a postmortem.

The public records do not establish one common failure. OpenAI attributed its trouble to routing. xAI and SpaceX pointed to a compute-center outage in Memphis. Anthropic said it had identified a cause but did not publish it. No provider disclosed a customer-by-customer impact count, a business-loss estimate, or enough technical detail to prove that the incidents were connected or fully independent.

The distinction matters because the outage arrived after companies had moved AI from optional chat into coding, research, support, recruiting, legal work, and internal knowledge systems. OpenAI's August enterprise report said active Codex users had grown 108 times in legal teams, 41 times in sales, 41 times in recruiting, 26 times in marketing, and five times in engineering between February and June. Those figures measure usage growth inside OpenAI's customer base. They do not show how many business processes failed on September 3. They do show why a model endpoint can now sit inside work that has a deadline.

An enterprise cannot read a provider status page and infer its own recovery time. Provider availability, contractual service levels, application health, and business continuity are separate measurements. A model may be reachable while a required tool is not. A chat interface may recover while queued work remains stranded. A second model may answer prompts but fail the tool calls, data controls, or quality checks that make the first model useful. A human may know the old process but no longer have the access or time to run it.

September 3 offers a bounded lesson. The evidence supports three overlapping provider incidents, with different public explanations and incomplete root-cause disclosure. It does not support a claim that all frontier AI failed, that one supplier caused all three events, or that every customer lost service for the full incident period.

The operational task begins at that boundary. Identify the work that depends on each system and decide how long it can wait. Test a technical alternate where one is justified, keep a staffed manual route where the consequence demands it, and plan how to reconcile the queue after service returns.

## September 3 produces three different incident clocks

Status pages look comparable because they use the same words: investigating, identified, monitoring, and resolved. The clocks underneath those labels can measure different things.

Anthropic's [public status history](https://anthropic.statuspage.io/) recorded an investigation into elevated errors for multiple models at 13:26 UTC. The incident listed Mythos/Fable 5.1, Mythos/Fable 5, Opus 5, Opus 4.8, and Opus 4.6, and said the impact ended at 16:16 UTC. Anthropic also recorded a shorter Sonnet 5 incident earlier that day, from 12:37 to 12:56 UTC. Combining the records into one long "Claude outage" would erase a distinction the provider itself preserved.

xAI's [model outage record](https://status.x.ai/grok-in-x/INC429d651a) starts at 13:30 UTC and ends at 17:05 UTC. It reports that Grok was experiencing issues and later that traffic was healthy. The page gives a duration of three hours and 35 minutes. It does not say what share of requests failed, which customer workflows stopped, or whether every interface was impaired for the entire period.

OpenAI's [ChatGPT and Codex incident record](https://status.openai.com/incidents/2rm6gqeh) identifies affected components rather than publishing one customer impact rate. Fifteen ChatGPT components and four Codex components appear on the record; the list does not say that every component had the same error rate for the same duration. The company also warns on its status site that availability metrics are aggregated across tiers, models, and error types. A customer on one plan, in one region, using one model can experience something different from the aggregate.

Independent reporting added detail without filling every gap. OpenAI spokesperson Kathleen Chaykowski told [WIRED](https://www.wired.com/story/nobody-is-saying-why-openai-and-anthropic-had-outages-today/) that a routing error ran from about 7:43 to 8:17 a.m. Pacific. That converts to 14:43 to 15:17 UTC, inside the Anthropic and xAI windows. The explanation is narrower than saying "OpenAI was down." It names an internal mechanism and an approximate period, while the status page records the wider customer-facing incident process.

A comparison that stays inside the public evidence looks like this:

| Provider | First confirmed time | Publicly named scope | Recovery marker | Public cause at review time |
| --- | --- | --- | --- | --- |
| Anthropic | 13:26 UTC | Elevated errors across several Claude models | Impact ended at 16:16 UTC | Cause identified but not disclosed |
| xAI | 13:30 UTC | Grok model issues | Traffic healthy at 17:05 UTC | xAI and SpaceX cited a Memphis compute-center outage |
| OpenAI | Routing event reported from 14:43 UTC | 15 ChatGPT and four Codex components on the incident record | Status record progressed to resolved | Routing error |

These are incident clocks, not business clocks. The first confirmed time may follow the first failed request. A recovery marker may mean error rates returned to normal, not that every queued task completed. An affected-component count says where the provider observed impact, not how many invoices, code changes, candidate messages, or support cases missed a deadline.

[Ars Technica's account](https://arstechnica.com/ai/2026/09/four-major-ai-models-suffer-rare-overlapping-downtime/) also discussed user reports involving Google services. Google did not post a corresponding incident on its own AI status dashboard. That is why Google is absent from the three-provider count in this article. User reports can reveal trouble before a provider confirms it, but they cannot substitute for a provider incident record when the claim is that the provider reported an outage.

Rolling uptime creates another tempting shortcut. On September 6, Anthropic's status page displayed about 99.4% availability for claude.ai and roughly 99.5% for its API over the preceding 90 days. Ars reported a 99.63% 90-day figure for ChatGPT on September 3. Different products, observation times, measurement rules, and customer routes make those numbers unsuitable for a simple league table or a claim about a contract. Even a precisely measured 99.5% says nothing by itself about whether the unavailable minutes arrived during payroll, a product release, an applicant deadline, or an idle Sunday.

The incident table is useful because it preserves the limit of the evidence. An enterprise needs a second table for its own work.

## Overlap does not prove a common failure

Three incidents in one workday invite one story. The available evidence supplies at least three.

WIRED reported that OpenAI traced its incident to a routing error. The same report said xAI and SpaceX attributed Grok's trouble to an outage at the Memphis compute center. Anthropic had identified its cause but did not disclose further detail. WIRED found no matching public incident at Amazon Web Services, Microsoft Azure, or Cloudflare that explained all three events when it checked. That observation limits one hypothesis; it is not proof that every shared upstream service was healthy.

That evidence weakens a confident claim about one shared infrastructure failure. It does not prove that the events were unrelated. Providers can share data-center suppliers, network routes, identity services, open-source components, safety services, traffic patterns, and customer-side integrations without those relationships appearing on public status pages. A later postmortem could add a dependency that was not visible on September 3.

As of September 6, the defensible classification is "overlapping incidents, shared cause unconfirmed." A later provider postmortem could change it. This cutoff leaves room for new evidence without filling the current silence with a theory.

The same problem appears inside companies. An operations team can observe that its research assistant, coding assistant, and customer bot failed near the same time. That does not tell the team where the common point sits. All three may call one provider. They may use different providers behind one internal gateway. They may share identity, retrieval, a vector store, a cloud region, an observability service, or a network egress path. The provider models may be healthy while the internal dependency is not.

A dependency map should separate five layers:

1. The employee or customer interface, such as a web app, IDE extension, help center, or messaging bot.
2. The company's application layer, including the orchestration service, prompt store, policy engine, retrieval system, and tool registry.
3. The model route, including provider, model family, region where selectable, rate limit, and account.
4. The business systems touched by tools, such as applicant tracking, source control, CRM, ticketing, document storage, or payment services.
5. The network, cloud, identity, logging, and secrets services that every route may share.

Without that map, "multi-provider" can describe two logos sitting behind the same fragile gateway. A routing rule may send work to OpenAI or Anthropic, but both paths can still require the same single sign-on service, the same retrieval index, and the same cloud function. Buying a second model does not remove a shared choke point that nobody recorded.

Cause classification also affects what a company asks from a supplier. A model error, regional routing error, account quota, content-policy refusal, tool timeout, and customer-side authentication failure need different fixes. One calls for model routing. Another calls for quota management. Another may require a manual approval queue. Treating every bad response as "the AI is down" hides the information needed for recovery.

The evidence package for an internal incident can stay small:

- the first failed transaction and its timestamp;
- the interface, model route, region or account, and tool involved;
- the provider incident identifier if one exists;
- request IDs and sanitized error classes, without copying confidential prompts into a broad incident channel;
- the last known successful transaction;
- the fallback decision, who authorized it, and what entered a queue;
- the provider recovery marker and the company's own recovery marker;
- any residual work that required replay, review, or customer contact.

That package lets an incident owner update the classification later. It also prevents a status-page screenshot from becoming the entire postmortem.

There is a reasonable counterargument: a two-hour AI interruption is ordinary software trouble, and not every assistant deserves an elaborate continuity plan. That is correct. Continuity controls should follow business consequence, not the cultural attention surrounding AI. A tool that helps employees rewrite meeting notes can wait. A tool that sends customer advice, changes production code, ranks applicants, or prepares a regulatory filing may need a bounded queue and a tested alternate. Treating every use case as critical wastes money. Treating every use case as a chat subscription ignores the work already delegated to it.

September 3 does not prove that frontier models fail together. It proves that a company can encounter several provider incidents during one operating window and still lack enough public detail to build its recovery from provider explanations alone.

## A chatbot subscription can hide a business service

An employee loses an assistant for an afternoon and sets the draft aside. No queue forms and nobody else waits. A second employee loses an assistant and misses the input promised to a customer team. Both may use the same subscription, but only one workflow has an immediate service dependency.

Work accumulation, downstream waiting, a missed customer response, or a moving deadline marks that dependency. It can exist even when procurement still records the product as a set of seats.

OpenAI's [enterprise usage report](https://openai.com/index/how-enterprises-put-ai-to-work/) offers a view of that shift. As of June, the company said Codex accounted for 64% of output tokens across Codex and ChatGPT in the enterprise data it studied. Active Codex users since February had risen 108 times in legal, 41 times in sales and recruiting, 26 times in marketing, and five times in engineering. These are provider telemetry and growth ratios, not measures of completed work or outage damage. Still, legal, sales, recruiting, and marketing are not experimental departments waiting outside normal operations.

The same endpoint can carry very different tolerances:

- A software engineer asks for a refactor suggestion. The task can wait four hours with little consequence.
- A release workflow asks an agent to diagnose a production failure. A 20-minute delay may extend a customer incident.
- A recruiter uses a model to draft candidate messages. The queue can wait, but automated retries must not send duplicates after recovery.
- A support bot answers account questions. The safe fallback may be a staffed queue, not a different model that lacks the same account context.
- A legal team uses an assistant to locate clauses. Work can continue manually only if people retain access to the source documents and the search method.

Provider uptime cannot encode those differences. The company has to define an impact tolerance for each business service. That tolerance is the maximum interruption the business is prepared to accept before it changes route, reduces the service, or works manually. It is a business decision, not a promise extracted from a status page.

Recovery time has at least three markers:

1. Provider recovery: the provider reports that the affected service is healthy.
2. Application recovery: the company's own route passes health checks and a representative transaction.
3. Work recovery: the queue is processed, duplicates and stale outputs are handled, downstream records agree, and the affected team can meet its next obligation.

The third marker can be hours after the first. Suppose a recruiting assistant generated 600 messages before an endpoint began timing out. The provider returns after 90 minutes. Some requests have an unknown outcome: the application did not receive a response, but a downstream action may have occurred. Blind replay could contact the same candidate twice. The service is not recovered until the team classifies uncertain transactions, reconciles the applicant system, and releases the remaining queue.

Customer-facing work adds a communication clock. An incident may fall below a provider's broad disclosure threshold while still mattering to one enterprise account. The company needs its own rule for when to tell customers that a response is delayed, when to remove an automated feature, and when to correct work produced during degraded operation.

Internal work has people on the other end too. An employee assigned a time-sensitive analysis needs to know whether to wait, switch tools, or use the manual route. Telling everyone to "try again later" transfers the continuity decision to each worker. Some will stop. Some will paste sensitive material into an unapproved service. Some will repeat requests until a rate limit makes the incident worse. A named decision owner and one status channel reduce that improvisation.

An IBM Institute for Business Value study conducted with Oxford Economics gives scale to the dependency concern. The [June 2026 release](https://newsroom.ibm.com/2026-06-17-ibm-study-limited-control-and-rising-dependencies-leave-enterprises-exposed-in-the-age-of-ai) covers 1,000 senior executives across 16 countries and 17 industries. Seventy-one percent said changing a primary AI vendor or model would be difficult. Ninety-one percent said they did not fully understand their AI dependencies. Respondents reported an average of six AI-related disruptions over two years, and 81% said a seven-day vendor outage would cause severe or critical disruption.

IBM sells technology and frames the findings around AI sovereignty. The survey relies on executive reports rather than observed incident logs. It cannot tell us what September 3 cost or whether a multi-vendor design reduced loss. It does identify a management mismatch: senior leaders say a long vendor outage would hurt, while most also say their dependencies are not fully understood.

The first correction is semantic but useful. Inventory business services, not AI products. "ChatGPT Enterprise" is a contract. "Prepare the daily exception report before the 9 a.m. risk meeting" is a service. Only the second description tells an operator what deadline, input, output, owner, and fallback must survive.

## Multi-model failover changes the work

When a team adds a second model, it inherits a second implementation. The controls, inputs, outputs, and operating limits may all differ.

AWS [production architecture guidance](https://docs.aws.amazon.com/prescriptive-guidance/latest/gen-ai-lifecycle-operational-excellence/preprod-architecting.html) describes an AI gateway that can route across regions, providers, or self-hosted models while centralizing quotas, cost, and observability. The pattern is sound. The diagram does not remove the application work needed to make each route safe.

Several differences tend to appear during a switch.

Prompts are behavior coupled to a model. An instruction that reliably produces a short structured answer on one model may produce commentary, omit a field, or interpret priorities differently on another. The alternate needs its own evaluation against the actual task.

Tool contracts vary too. Providers use different schemas, validation behavior, streaming events, error classes, and parallel-call conventions. A model that can answer a question may still fail the tool sequence that updates a ticket or retrieves an applicant record.

Context does not automatically move. Conversation state, retrieval results, cached summaries, uploaded files, and provider-managed memory may live inside the primary route. A fallback needs a permitted, current representation of the state, not an improvised copy from an employee.

Safety policy, identity, and audit evidence can change with the route. A backup may refuse a permitted task, answer a task that the primary route blocks, or apply different filters to names, health information, source code, and personnel data. It may also use a different account, retention setting, regional configuration, encryption path, or logging system. If the company cannot connect an output to a user, policy version, source, and approval, it may recover throughput while losing accountability.

Quality remains use-case specific. A benchmark average does not establish that the backup can classify this company's support intent, preserve a contract citation, or call its internal tool correctly. The minimum test set should contain ordinary cases, known edge cases, refusal cases, and examples where the safe result is to stop.

Restoration then creates state conflict. Requests may have succeeded after the caller timed out. A queue sent to the alternate may still be waiting on the primary. When both routes recover, idempotency controls and a reconciliation owner matter more than model eloquence.

These differences produce a ladder of fallback options:

| Fallback level | Action | Best fit | Main risk |
| --- | --- | --- | --- |
| Wait and queue | Stop new execution, preserve requests, publish a delay | Low-consequence asynchronous work | Deadline breach or uncontrolled backlog |
| Retry primary | Use bounded retries with jitter and a stop rule | Brief transient errors with safe idempotency | Retry storm or duplicate side effects |
| Change primary route | Select another approved model or region from the same provider | Model-specific or regional issue | Shared provider control plane still fails |
| Route to second provider | Use a separately tested provider path | Material work with portable state and policy | Behavior, data, tool, and audit mismatch |
| Degrade the service | Remove generation or tool actions, offer search, forms, or read-only access | Customer service where a smaller function is still useful | Users mistake reduced output for normal service |
| Work manually | Give a trained person the source data, authority, and queue | High-consequence work that cannot wait | Capacity, skill, access, and fatigue limits |

The order need not be fixed. A customer bot might degrade immediately to a contact form. A coding assistant might queue. A regulated decision may stop rather than switch models. The route belongs to the workflow owner because the consequence belongs to the workflow.

Multi-provider design has a cost before an incident occurs. Teams must negotiate and review another supplier, connect identity and logging, maintain evaluations, adapt prompts and tools, reserve quota, monitor drift, train operators, and run exercises. The second route can also expand the data boundary and the number of failure combinations.

That cost supports the counterargument against universal failover. A low-impact assistant may be cheaper to pause for a day than to maintain a fully tested alternate. The decision should compare the expected consequence of interruption with the recurring cost of recovery capability. It should not assume that every AI minute has equal value.

Correlated risk remains after the second provider is live. Both routes may depend on one cloud, one network, one source database, or one internal gateway. They may also fail under the same unusual input even if the infrastructure differs. Independence has to be demonstrated at the dependency being protected. Two model contracts provide vendor diversity. They do not automatically provide network, data, identity, or process diversity.

The smallest credible test is a forced switch in a non-production environment using a representative queue. Measure the time from the incident decision to the first accepted alternate output, then inspect correctness, tool effects, audit records, cost, and restoration. If the test stops at "the backup model returned text," the company tested model access rather than business recovery.

## Human fallback loses skill while it waits

Architecture diagrams often label one box "manual fallback" without showing the people, access, or hours behind it. That route may be unavailable when the incident arrives.

AWS [incident-response guidance for agentic AI](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-security/best-practices-incident-response.html) recommends safe fallback systems and staff for critical processes, with recovery objectives and defined recovery methods. It also identifies a less visible risk: human operators can lose the skills needed to take over when automated systems handle the work for long periods.

Skill decay is only one constraint. The manual worker also needs source access, decision authority, current instructions, a place to record the result, and time. An employee whose full workload assumes automation does not become spare capacity when automation fails.

Consider four manual routes:

- A support agent takes over from a customer bot. The agent needs the conversation, account context, approved response policy, and a queue that prevents the bot from sending a later duplicate.
- A recruiter takes over candidate communication. The recruiter needs the approved template, candidate status, time-zone rule, consent boundary, and a record of messages that may already have been sent.
- An engineer works without a coding agent. The engineer may write code manually, but the release process can still depend on an AI review or incident-analysis step that nobody designated as optional.
- A legal reviewer replaces clause extraction. The reviewer needs the original documents and a manual search path. If the source was uploaded only into a provider workspace, "do it manually" is not a route.

A continuity plan should name capacity in units of work. "Operations will handle it" is not a capacity statement. "Two trained operators can clear 80 cases per hour for four hours, after which a second shift is required" can be tested. The plan should also state which cases receive priority and which are allowed to expire.

This is where AI continuity becomes a workforce-design issue. Automation often removes routine repetitions first. Those repetitions may be how junior employees learn the exception patterns that experts use. If people no longer perform or review enough of the base task, a manual takeover months later can be slower and less accurate than the old process.

Maintaining the fallback does not require preserving every old step. It requires preserving the outcome, evidence, and safety boundary. A quarterly exercise might give a team a small real or synthetic queue without the primary AI route. Operators use source systems, record decisions, and measure completion time. Reviewers then compare the manual and automated paths for omissions and policy errors.

The exercise should reveal labor consequences rather than conceal them. Who was pulled from other work? Which deadlines moved? Did the fallback depend on unpaid extra hours? Did a specialist become a single point of failure? Did the queue expose people to unusually repetitive, distressing, or sensitive material? A plan that succeeds only through emergency overwork has transferred the outage cost to employees.

Human review and human fallback are also different controls. Review assumes the automated output exists and asks a person to inspect it. Fallback asks a person to create or complete the work when the output does not exist. A team staffed for sampling 5% of outputs may not have capacity to produce 100% of them.

An outage can create an awkward accountability gap. Leaders may say a model made the error before the interruption, while employees own every decision during manual operation. The business should assign responsibility consistently. A person needs authority to reject a stale queued output, halt replay, or keep a service degraded after the provider says it is healthy.

Restoration deserves its own role. Someone must compare work completed manually, work completed by the alternate route, uncertain primary transactions, and new incoming demand. That person decides what to replay, cancel, review, or communicate. Without reconciliation, recovery can produce duplicate candidate messages, conflicting ticket updates, overwritten document changes, or code suggestions based on an older branch.

The manual route becomes credible when an operator has used it recently, can access it without an emergency permission scramble, and knows the point at which capacity runs out. A document that says "human in the loop" provides none of those facts.

## Price the recovery before buying more tokens

The September 3 overlap can trigger a quick procurement response: add another provider. That may be correct for a few workflows. It is an expensive substitute for deciding which work deserves recovery.

IBM's survey found that 73% of respondents described their organizations as intentionally multi-vendor, but only 7% placed themselves at an advanced level of control over AI infrastructure, models, data, security, and skills. Seventy-two percent said they would accept a 20% cost increase in exchange for more flexibility and control. The figures come from a vendor-sponsored executive survey and measure stated willingness, not signed budgets or proven resilience.

The cost of a second route has at least six parts:

1. Supplier cost: minimum commitments, reserved capacity, support, legal review, and security assessment.
2. Build cost: gateway changes, adapters, tool schemas, state transfer, identity, logging, and user controls.
3. Assurance cost: task-specific evaluations, red-team cases, privacy review, audit tests, and repeated checks after model changes.
4. Operating cost: monitoring, incident ownership, quota management, documentation, and exercises.
5. Workforce cost: training, protected drill time, manual capacity, and the work displaced during takeover.
6. Restoration cost: queue reconciliation, customer communication, correction, and investigation after service returns.

Against that total sits the cost of waiting. A team can estimate it without pretending to know a perfect outage probability. Record the business deadline, work arrival rate, value or consequence per delayed unit, maximum queue, contractual or regulatory exposure, and manual clearing rate. Compare a two-hour, one-day, and seven-day interruption. The ranges will be more honest than one return-on-investment percentage.

Regulation can narrow the acceptable range for some firms. The UK Financial Conduct Authority's [PS26/2 policy statement](https://www.fca.org.uk/publications/policy-statements/ps26-2-operational-incident-third-party-reporting) introduces unified operational incident and material third-party reporting from March 18, 2027, for firms within its scope. Those firms must also maintain an annual third-party register. This is not a universal AI law, and an AI supplier is not automatically material under the rule. It does mean that an in-scope financial firm needs evidence about suppliers and incidents when the relationship meets the regulator's threshold.

Contract review matters alongside regulation. The product status page is not the service order. The order form, service-level terms, support tier, data terms, and limitation provisions determine what the supplier promised. A procurement team should ask which product and component the commitment covers, how downtime is calculated, what exclusions apply, what notice arrives during an incident, and whether the remedy is a credit or operational help. A credit can offset a bill while doing nothing to clear Monday's work queue.

The decision artifact below turns those questions into an AI work-continuity ledger. Create one row per business workflow, not one row per vendor.

| Ledger field | Decision or evidence to record |
| --- | --- |
| Business service | Outcome and recipient, stated without a product name |
| Deadline and arrival rate | When work must finish and how quickly the queue grows |
| Primary route | Interface, application, provider, model, account, region where known, and tools |
| Shared dependencies | Identity, gateway, cloud, retrieval, source systems, network, logging, and secrets |
| Impact tolerance | Maximum acceptable interruption before the route changes |
| Detection and authority | Health signal, incident owner, and person allowed to switch or stop |
| Queue rule | What is accepted, rejected, expired, deduplicated, and prioritized |
| Technical alternate | Tested model, region, provider, degraded service, or explicit decision to wait |
| Alternate evidence | Last test date, test set, quality result, tool result, audit result, and switch time |
| Human route | Named role, access, authority, current procedure, throughput, and fatigue limit |
| State portability | Inputs, context, pending actions, and records that can move safely |
| Restoration rule | Health check, replay boundary, reconciliation owner, and correction path |
| Communication | Employee, customer, supplier, legal, and regulator notice thresholds |
| Incident record | Provider ID, internal timestamps, error classes, decisions, residual queue, and later cause update |

A completed row might describe candidate interview scheduling. The outcome is a confirmed slot and one accurate message to each participant. The workflow receives 40 requests an hour and can wait two hours before interviews start to slip. The primary AI route interprets replies and drafts messages, but the calendar system makes the booking. A degraded route can send a form without generation. Two coordinators can manually process 25 cases an hour. Every request carries an idempotency identifier, and uncertain sends are checked before replay. Recovery occurs when the queue is reconciled, not when the model endpoint returns HTTP 200.

That example may lead the company to reject a second model. A form plus manual prioritization could meet the tolerance at lower cost and with less data movement. Another workflow, such as incident diagnosis during a product outage, may justify a separately tested provider because every delayed minute extends customer impact. The ledger makes the difference visible.

The tested recovery has to change a business outcome enough to justify its cost. Workflows that do not clear that bar can queue, degrade, or stop under an explicit rule.

## Monday's first task needs an offline route

The first continuity exercise should be small enough to run next week and real enough to fail.

Choose one workflow that has a deadline and already uses an external model in normal work. Do not begin with the most regulated or dangerous process. Do not choose an optional writing aid that can disappear without creating a queue. Pick a middle case where interruption matters and recovery can be observed, such as support triage, internal incident summaries, candidate scheduling, sales-call follow-up, or contract search.

Then run the following sequence:

1. Name the business outcome, recipient, deadline, and impact tolerance.
2. Draw the route from interface through provider and tools to the final system of record. Mark every dependency shared by the primary and alternate.
3. Export or reconstruct ten to thirty representative tasks, including edge cases and at least one transaction with an external side effect.
4. Disable the primary model route in a test environment. Do not merely ask the team to imagine it is unavailable.
5. Start the clock when the workflow owner receives the alert. Record the switch decision, queue behavior, first alternate output, manual throughput, and any permission request.
6. Restore the primary route while some work remains pending. Reconcile uncertain, manual, alternate, and primary transactions.
7. Compare the result with the stated tolerance. Record defects, labor used, displaced work, customer communication, and the next owner and date.

The test has failed if the alternate returns text but cannot complete the business action. It has also failed if people complete the queue but need unrecorded access, ignore a policy, work beyond planned capacity, or cannot tell which messages already went out. Those failures are useful before an actual incident.

Management should expect one uncomfortable result: the lowest-cost recovery may be to stop. A business can pause a feature, display a clear delay, and preserve the queue instead of routing consequential work through an untested model. Availability pressure should not turn an incident into a privacy, quality, or duplicate-action problem.

The result also informs purchasing. If a tested same-provider model switch meets the tolerance, another vendor may add little. If every model route shares a failing internal gateway, fix the gateway. If manual work clears the queue but exhausts the only trained operator, add training and capacity. If no route preserves the deadline, leadership must either fund recovery or change the promise made to customers and employees.

Charlie Dai, a Forrester vice president and principal analyst, told [ITPro](https://www.itpro.com/security/ai-is-increasingly-becoming-operational-infrastructure-rather-than-a-productivity-add-on-yesterdays-triple-ai-outage-should-be-a-wake-up-call-for-enterprises) that enterprises should consider multi-model strategies, fallback workflows, continuity planning, and supplier transparency. That is analyst advice, not proof that a particular architecture survived September 3. The exercise converts the advice into local evidence.

Ana Paula Assis, an IBM senior vice president, framed AI resilience as a matter of control over models, data, infrastructure, and skills when IBM released its survey. IBM has a commercial interest in that framing. A company does not need to adopt the vendor's whole control model to test a narrower fact: can this team finish this work when this route is unavailable?

September 3 supplies the drill scenario, not the answer. At 13:26 UTC one provider began investigating. At 13:30 another provider's outage clock started. At 14:43 a third provider's routing error began. The incidents overlapped. Their shared cause remains unconfirmed.

On Monday, take one production-like queue offline for 45 minutes. Let the workflow owner make the switch. Require the alternate or manual route to finish the action, preserve the evidence, and reconcile after restoration. Put the elapsed time, error list, labor record, and reconciled queue into the first completed row of the work-continuity ledger. Those local results will tell the company more than three provider uptime percentages.

### Related Reading

- [AI Safety Tests Reached Systems They Did Not Own](/2026/09/01/ai-safety-tests-live-systems/)
- [Legal AI Splits Between Owned Models and Frontier APIs](/2026/08/25/legal-ai-owned-models-frontier-apis/)
- [When HR Agents Ship, Someone Has to Carry the Pager](/2026/06/22/hr-agents-operating-pager/)
- [AI Credits Distort the Startup Headcount Plan](/2026/07/07/ai-credits-startup-headcount-plan/)
