Enterprise AI procurement should begin with an accepted business outcome and a test set, not a vendor category or a spending forecast. A buyer needs to know which work is changing, what evidence will count as success, which data and authority the system receives, who owns failure, and how the organization exits. Price comparison comes after scope is stable.

There is no universal percentage of AI pilots that fail, no dependable industry discount, and no default return on investment. Those figures combine different products, stages, and definitions. A reusable process instead moves through five gates: problem fit, evidence, risk, contract, and production acceptance.

Define the decision before contacting vendors

Write one sentence that describes the current process, affected user, required outcome, and constraint. “Buy an enterprise agent” is not a problem statement. “Reduce the time analysts spend locating current policy clauses while preserving source citation and existing access rules” is testable.

The UK’s AI procurement guidelines advise buyers to focus on the challenge rather than prescribe a specific AI solution. They also recommend a data assessment and early risk work. The guidance targets public procurement, but the sequence transfers well to private buyers.

Document the non-AI alternative. A search upgrade, rules engine, workflow redesign, or added training may solve the problem with less uncertainty.

Classify the workflow and consequence

Not every use case deserves the same diligence. A drafting assistant that cannot publish has a different failure cost from a system that changes benefits, ranks applicants, moves money, or controls industrial equipment.

FactorLower consequenceHigher consequence
OutputSuggestion reviewed before useDecision or external action
DataPublic or low sensitivityPersonal, regulated, confidential, or safety data
AuthorityRead onlyWrite, send, approve, delete, or purchase
ReversibilityEasy correctionIrreversible or costly recovery
ScaleFew known usersAutomated use across many people or records
ContestabilityError is visible and appealableAffected person may never see it

The highest applicable consequence should drive the control level. An average risk score can hide one unacceptable capability.

Build an evidence matrix before the demo

Prepare questions and required artifacts before vendors shape the evaluation around their strengths.

ClaimMinimum evidenceBuyer verification
Task qualityMethod, dataset, model and product versionRun a representative blind test
ReliabilityError distribution and known failure modesAdd edge cases and repeated runs
SecurityArchitecture, access controls, incident processThreat model and technical testing
PrivacyData flow, retention, training use, subprocessorsContract mapping and deletion test
FairnessPopulation, groups, metric and limitationsReview relevance to actual users
IntegrationSupported APIs, limits, field mappingTest failure and reconciliation
EconomicsDefined unit and included servicesModel full workflow cost
PortabilityExport format, timing and scopePerform a sample export

A vendor document can answer what the vendor claims. It cannot replace a buyer test using the intended workflow and configuration.

Evaluate the full system, not the base model

Enterprise performance depends on prompts, retrieval, tools, data, permissions, interface, and human review. A benchmark for the underlying model does not validate the purchased application.

NIST’s AI Risk Management Framework organizes risk work into Govern, Map, Measure, and Manage. NIST says the framework is voluntary, non-sector-specific, and use-case agnostic. It does not certify a product or prescribe one score.

Use the intended production configuration in testing. Record source quality, unsupported claims, refusals, latency, cost, permission errors, tool errors, and whether a person accepted the result. Repeat non-deterministic tasks to see variation.

Build versus buy is a control decision

Buying can accelerate access to models, infrastructure, security features, and support. Building can preserve control over data flows, evaluations, interfaces, and switching. Most enterprises combine both: they purchase model or platform services and build the workflow layer.

Choose based on the scarce capability. If the workflow is common and differentiation is low, a configured product may be sensible. If proprietary process knowledge, regulatory obligations, or integration depth determine value, more of the application may need to remain under buyer control.

Do not confuse building with training a frontier model. An enterprise can own prompts, retrieval, evaluation, and orchestration while using external models. It can also buy a packaged application and still retain responsibility for configuration and use.

Total cost follows the accepted outcome

Token price or seat price is only one input. Include implementation, data preparation, connectors, security review, evaluation, human review, support, monitoring, incident response, model changes, and exit work.

Use a cost equation tied to the workflow:

cost per accepted outcome = all operating and review costs / outputs accepted for intended use

Separate fixed implementation cost from variable usage. Model low, expected, and high volume using contract units. If the vendor prices by action, define an action and determine whether failed calls, retries, background processing, or test traffic count.

No generic discount percentage can replace a quote. Ask for the rate card, minimum commitment, overage treatment, renewal method, and a sample invoice based on your expected pattern.

Data terms need operational detail

“Your data is private” is not a contract specification. Identify each data type, where it travels, how long it remains, who can access it, whether it is used for training or service improvement, and what happens in logs and backups.

Require a current subprocessor list and notice process. Define deletion timing and exceptions. Map residency and cross-border transfer requirements. Determine whether customer prompts, retrieved records, outputs, feedback, and usage metadata receive different treatment.

Test access inheritance. A retrieval system should not expose a document merely because the model can find it. The user’s existing authorization should constrain the context supplied to the model.

Security review must cover agent authority

Conventional application review remains necessary: identity, secrets, encryption, vulnerabilities, tenancy, logging, and incident response. Agent systems add tool choice, generated arguments, untrusted context, and side effects.

CISA’s Secure by Demand guide recommends placing product-security requirements into procurement and contract language. For an AI agent, include least privilege, separate read and write tools, approval for consequential actions, idempotency, rate limits, and a kill switch.

Ask the vendor to demonstrate a failed permission check, revoked credential, prompt injection attempt, partial tool outage, and uncertain transaction result. Security claims become useful when the failure behavior is visible.

Contract for change, not one model snapshot

AI services can change model versions, system prompts, safety behavior, limits, subprocessors, and features during a contract. Define what change notice is required and which changes trigger revalidation or a termination right.

The contract should address:

  1. The named service, features, regions, and model policy.
  2. Permitted data uses, retention, deletion, and subprocessors.
  3. Security controls, incident notification, and investigation support.
  4. Evaluation access, logs, audit cooperation, and documentation.
  5. Service levels for the complete workflow, not only API uptime.
  6. Intellectual-property rights and claim handling for inputs and outputs.
  7. Usage units, minimums, overages, renewal, and price changes.
  8. Export formats, transition support, deletion proof, and exit timing.

Allocate responsibility to the party able to control the risk. A vendor may control model operations; the customer may control job criteria, user access, and final use.

Portability must be tested before signature

The U.S. Government Accountability Office’s 2026 report on federal AI acquisitions says agencies encountered difficulty understanding AI costs and accessing technical expertise. Its review of acquisition guidance notes data and model portability, clear licensing, pricing transparency, testing, and vendor-lock-in terms.

Although the report covers federal agencies, the failure modes apply broadly. A contractual export right is weak if the export omits prompts, evaluations, source links, configuration, or audit history. Run a small export and import before the contract is final.

Identify replacement time and dual-running cost. A system that can export data but not reproduce decisions may still be difficult to leave.

Use a staged acceptance plan

Procurement award is not production acceptance. Divide rollout into stages with evidence and stop conditions.

StageAllowed behaviorExit evidence
Offline testNo production data or actionTask and risk baseline met
Controlled pilotLimited users and approved dataQuality and workflow fit confirmed
Shadow modeObserve real cases without actingComparison with current process
Approved actionHuman confirms consequential stepsBounded incidents and reliable recovery
Expanded productionWider scope within defined limitsSustained accepted outcomes and cost

Attach material acceptance criteria to payment, expansion, or renewal. Keep the old path until recovery and rollback have been tested.

Governance needs named owners

A steering committee is not useful if no one can stop the system. Name an executive outcome owner, technical owner, data owner, security owner, legal or compliance owner, and operational reviewer. Define who accepts residual risk and who handles a user complaint.

OMB’s current federal acquisition memorandum, M-25-22, emphasizes clear requirements, testing, interoperability, transparency, and avoiding vendor lock-in in government AI purchasing. Private companies are not bound by that memorandum, but its questions are useful prompts.

Review ownership after a model or use-case change. The team that approved a drafting pilot may not be qualified to approve autonomous external action.

Monitor value and risk together

Production monitoring should connect technical and business outcomes. Track task completion, accepted outputs, review effort, latency, spend, incidents, complaints, override rates, access failures, and subgroup effects where lawful and relevant.

A lower unit cost can accompany worse service. Faster processing can create more rework. High use can reflect a mandatory tool rather than value. Compare metrics with the baseline and preserve context about other process changes.

Set explicit pause conditions. Examples include unauthorized data access, a severe harmful outcome, an unexplained quality drop after a model change, loss of required logs, or spend outside the approved bound.

Bottom line

AI procurement is a continuing evidence process, not a one-time vendor selection. Define the workflow, classify consequence, test the complete system, model the full cost, contract for change and exit, and expand only when production evidence supports it.

This approach does not promise a universal ROI or eliminate uncertainty. It makes uncertainty visible and assigns it to a decision owner. That is the foundation for buying AI without confusing a compelling demo with a dependable system.