Matt Garman’s mandate at Amazon Web Services is to keep AWS useful across the full AI stack: data centers and networking, custom silicon, model access, and the services needed to run agents in production. This is less a contest over a single model than a test of whether AWS can turn large capital commitments into reliable, economical customer workloads.

This profile was checked against public sources on September 14, 2026. Product performance and adoption figures below are identified as Amazon claims unless they come from a regulatory filing.

His role and operating brief

Amazon’s 2026 proxy statement says Garman has served as AWS CEO since June 2024. It also records a long progression through EC2, compute services, sales, and product leadership. That history matters because his job is not limited to AI marketing. He runs a cloud business whose existing customers expect continuity, security, capacity, and cost control while AI changes the workload mix. The proxy statement is the most authoritative public source for his title and career chronology.

AWS under Garman offers three connected layers. Trainium and Inferentia are Amazon-designed accelerators. Amazon Bedrock supplies managed access to models from multiple providers. AgentCore and related services address runtime concerns such as identity, observability, memory, and tool access. The strategic idea is that customers should be able to choose models while AWS retains the underlying compute and operational workload.

Custom silicon changes the bargaining position

Amazon has strong incentives to reduce its dependence on any one accelerator supplier. In its 2025 shareholder letter, the company said Trainium2 had about 30 percent better price-performance than comparable GPUs, that capacity was largely sold out, and that most Bedrock inference ran on Trainium. Those are Amazon’s own performance and utilization claims, not independent benchmark results.

The strategic value is broader than a benchmark. A credible in-house chip gives AWS another capacity source, a way to tune hardware and software together, and leverage in supplier negotiations. It also introduces switching costs. Customers must evaluate compiler maturity, supported operators, debugging tools, portability, and the availability of people who can operate the stack. A lower quoted compute price does not guarantee a lower total workload cost.

At re:Invent 2025, Amazon presented Trainium3 UltraServers, Bedrock model expansion, Nova models, and AI Factories. The announcement describes the intended integrated stack, but its speed and cost comparisons are vendor-reported. Buyers should reproduce them with their own model, context length, batching pattern, reliability target, and regional capacity requirements.

Model choice is AWS’s distribution strategy

Bedrock is designed as a managed catalog rather than a bet on one frontier-model provider. That lets AWS serve customers that want Anthropic, OpenAI, Amazon, or open-weight models without moving the surrounding application platform. The approach can reduce integration work, but only if model switching is real in practice. Prompts, tool schemas, safety controls, latency behavior, and evaluation baselines often differ between models.

The supplier relationships are therefore important. Amazon and Anthropic described AWS as Anthropic’s primary cloud provider and said future models would use Trainium and Inferentia. Amazon also announced an AWS and OpenAI compute agreement in November 2025. The OpenAI workload announcement confirms the relationship and identifies Bedrock availability, while leaving workload economics and realized utilization undisclosed.

Agents make operations the product

Garman has argued that AI value will move from generating content to completing tasks. Amazon’s account of his re:Invent 2025 remarks connects that view to Bedrock AgentCore. The company interview is useful evidence of management’s direction, not proof that agents deliver positive returns for every customer.

Production agents raise questions that model demos avoid. Teams need to know which identity an agent uses, which tools it can call, how approvals work, where traces are retained, how failures are contained, and what a rollback actually reverses. NIST’s AI Risk Management Framework provides a vendor-neutral structure for governing, mapping, measuring, and managing those risks.

A practical buyer test

AWS customers should evaluate Garman’s strategy with workload evidence rather than platform breadth alone. A useful comparison includes end-to-end latency, successful task completion, human-review time, failure recovery, data-egress exposure, reserved-capacity commitments, and the cost of changing models or accelerators.

The strongest case for AWS is integrated choice: several model families, Amazon and third-party chips, and mature cloud controls in one operating environment. The main risk is that integration becomes dependence across several layers at once. Garman’s performance should therefore be judged by repeatable customer economics and operational reliability, not by the number of AI announcements.

Evidence boundaries

Amazon discloses AWS segment results, but it does not separately publish audited revenue or margin for Bedrock, Trainium, or AgentCore. Public materials also do not establish a universal cost advantage over GPUs or competing clouds. Customer-specific results remain sensitive to architecture and contract terms.

As of the check date, the public record supports three conclusions: Garman is AWS CEO, AWS is investing across chips, models, and agent operations, and major model companies have announced infrastructure relationships with AWS. Claims about market leadership, superior economics, or customer outcomes require workload-level validation.