Konstantine Buhler’s public case for an agent economy is narrower, and more useful, than the sensational version often attached to it. The Sequoia Capital partner is not on record promising a particular multitrillion-dollar market or a fixed timetable for autonomous businesses. His documented thesis is that software is moving from answering questions toward completing work, and that this change requires new infrastructure, management practices, and controls.

That distinction matters. A venture thesis is a map of where an investor expects value to form. It is not a market forecast, an independent product test, or proof that a portfolio company can perform a regulated workflow. Readers should evaluate Buhler’s framework against deployment evidence, not repeat it as an established economic outcome.

This analysis separates three layers: what Buhler and Sequoia have published, what independent evidence currently supports, and what remains a judgment about the future.

Who Konstantine Buhler is

Sequoia’s official biography identifies Buhler as an investment partner who works with seed and early-stage companies. It says he studied artificial intelligence and looks for businesses that combine data cycles with intelligent algorithms. The page also lists investments, but a portfolio list reflects the firm’s selection decisions. It does not establish product quality or investment performance.

His public work is unusually focused on operating systems around AI, rather than a single model vendor. Sequoia’s AI Ascent 2025 recap says Buhler discussed how an agent economy could change work and which technologies would be needed. The recap confirms the subject and the firm’s framing. It does not publish a market-sizing model, independent benchmarks, or the economics of the companies mentioned.

The responsible reading is therefore biographical and analytical. Buhler is an investor articulating a thesis through a company publication. Sequoia is the interested party, and its articles should be labeled as company disclosures.

What Sequoia means by an agent economy

In Buhler’s Always-on Economy essay, the central idea is that AI agents may operate continuously, communicate with people and other software, and handle parts of economic coordination. He places the broad transition on a five-to-seven-year horizon. That is his forecast, not a measured adoption curve.

The useful part of the concept is not a claim that agents will form a separate economy. It is a checklist of capabilities that ordinary business software often lacks: persistent identity, permissions, memory, transaction authority, auditability, and a way to resolve exceptions. A system that can draft a purchase order but cannot identify the approving person is a tool. A system that can place the order, document the authority used, and stop when policy is ambiguous is closer to an operational agent.

This changes the unit of analysis. Buyers should stop asking only whether a model can produce a convincing response. They should ask whether the complete system can deliver an accepted result under real constraints.

Work shifts from answers to completed actions

Sequoia’s 2025 event recap records Buhler’s view that AI is moving from chat toward completed enterprise workflows. That direction is visible in products that search records, draft an action, call another service, and request approval. Yet the number of connected tools is not evidence of autonomy.

A completed action has at least four requirements. The goal must be clear enough to test. The agent must receive only the permissions it needs. The output must pass an acceptance check. The organization must know who handles failure. Without those conditions, automation can move an error faster and make its origin harder to trace.

The difference is important for search readers evaluating the agent economy thesis. A polished demo usually shows the happy path. Production value depends on the frequency and cost of every other path.

A practical stack for evaluating agent businesses

The thesis becomes decision-useful when it is translated into layers that can be tested.

LayerNecessary capabilityEvidence to request
ModelProduce a usable next stepTask-specific evaluation set and error distribution
ContextRetrieve current, authorized informationSource trace, freshness policy, and access tests
ActionCall tools within defined limitsPermission map, sandbox results, and transaction logs
ControlStop, escalate, and recoverApproval rules, rollback path, and incident history
EconomicsBeat the existing processCost per accepted outcome, including review and rework

No layer can compensate indefinitely for a missing one. Better model output does not repair stale permissions. A detailed audit log does not make an unprofitable workflow valuable. Investors and buyers need evidence across the stack.

Probabilistic output changes management

Buhler’s essay on a stochastic mindset argues that probabilistic systems require organizations to rethink how work is assigned and reviewed. This is a valuable management observation. Conventional software is expected to execute the same defined operation repeatedly. Generative systems can produce different outputs from similar inputs, and the errors may look plausible.

The operational response is not to demand perfect determinism from every model. It is to match controls to consequence. A research assistant can surface several candidate documents with lightweight review. A system that changes payroll, deletes production data, or sends a contractual commitment needs explicit authority boundaries and a reliable human escalation path.

This is where the thesis connects to real organizational design. The likely adoption unit is a bounded workflow with known owners, not an abstract digital employee.

Reliability limits the length of delegated work

Independent evidence counsels restraint about long, unattended workflows. METR’s time-horizon methodology estimates the task length at which models can complete certain software tasks with a given success rate. METR also states important limits: results depend on the task set, scoring method, model, and scaffold. The measure does not imply that a model can reliably perform every real-world job of the same duration.

Stanford’s 2026 AI Index economy chapter presents a similar tension at the organizational level. AI use is widespread, while agent deployment remains in the single digits across nearly all measured business functions. Adoption of AI tools and delegation to agents are different states.

These findings do not refute Buhler’s forecast. They identify the gap it must cross. Longer task chains multiply opportunities for a bad assumption, outdated record, failed tool call, or missed exception. Progress should be measured by reliable delegation length under realistic conditions, not by the maximum number of steps shown in a demo.

Portfolio examples are not market proof

Sequoia articles often illustrate a thesis with portfolio companies. That is normal venture communication, but it introduces selection and financial interests. A company named as an example may have a capable product. The mention itself is still not independent validation.

For each example, readers should seek evidence that is specific to the claimed workflow: retained production use, an externally defined success measure, failure rates, review time, and the cost of operation. Customer logos show a relationship, not scope or value. Funding shows investor demand, not user demand. A benchmark created by the vendor may be informative, but its task design and scoring need inspection.

This source discipline prevents a circular argument in which Sequoia’s investments prove Sequoia’s thesis and the thesis validates the investments.

Safety and control are product requirements

The U.S. National Institute of Standards and Technology describes human oversight as an operating responsibility, not a slogan. Its AI Risk Management Framework guidance on human and AI interaction calls for organizations to define roles, responsibilities, and how people interact with systems. NIST’s generative AI profile adds risk considerations specific to generative systems.

An agent product therefore needs more than an accuracy claim. It needs scoped credentials, durable logs, data handling rules, approval thresholds, and a tested recovery process. The system should expose uncertainty when it affects a decision. It should not quietly substitute an inferred objective for an authorized one.

These controls are not merely compliance overhead. They determine whether a company can trust the product with more valuable work.

Economics should be measured per accepted outcome

Seat-based software metrics can obscure agent economics. If an agent performs work, a more useful denominator is the accepted outcome. The calculation should include model inference, orchestration, external services, human review, failed attempts, rework, and incident cost.

Suppose a system drafts 1,000 support resolutions. Eight hundred are accepted after quick review, 150 require substantial editing, and 50 create escalations. The cost per draft will look attractive. The cost per accepted resolution may not. The same distinction applies to sales research, security remediation, recruiting operations, and finance workflows.

This also clarifies pricing. Outcome pricing can align vendor and customer incentives only when the outcome is defined, attributable, and resistant to gaming. Otherwise it transfers a measurement dispute into the contract.

A decision test for founders and buyers

Founders applying Buhler’s thesis should be able to name the narrow workflow they own. They should know who authorizes each action, what constitutes acceptance, how the system detects a boundary, and how performance changes when inputs are messy. A roadmap toward greater autonomy is credible only if each stage has observable proof.

Buyers should run a controlled comparison against the existing process. Measure completion, review minutes, correction rate, elapsed time, and total cost. Include difficult and adversarial cases. Check whether performance persists after the vendor’s implementation team leaves. Confirm that logs, export, deletion, and permission revocation work in practice.

Investors need one further test: whether the product accumulates a defensible advantage from deployment. Proprietary data is not automatically a moat, particularly when customers restrict its reuse. Durable advantages may instead come from workflow integration, evaluation knowledge, distribution, trust, or a lower cost of resolving exceptions.

Bottom line

Buhler offers a coherent way to think about software that completes work rather than only supplying information. His strongest contribution is the recognition that agents need an operating layer around the model: identity, memory, permissions, transactions, controls, and a management system suited to probabilistic output.

The evidence does not justify a precise market size, a guaranteed timetable, or a claim that autonomous agents have already transformed most business functions. Sequoia’s publications are company disclosures and investment arguments. METR, Stanford, and NIST provide useful independent checks on reliability, deployment, and governance.

The agent economy thesis should be treated as a testable direction. Its progress will be visible when bounded workflows deliver more accepted outcomes at lower total cost, with failures that organizations can detect, contain, and correct.