The Handoff Boundary: Why Enterprise AI Programs Stall Where Decisions Change
Value capture in enterprise AI tracks the depth of decision-right transfer — not adoption breadth, model performance, or spend.
Abstract
Enterprise artificial intelligence programs are conventionally evaluated as deployment exercises: seats licensed, workflows instrumented, pilots completed. Value outcomes, however, have dissociated from those counts. This paper argues that the binding constraint is neither model capability nor organizational change resistance, but governance design. Drawing on transaction cost economics (Williamson, 1985) and dynamic capabilities theory (Teece, 2007), the analysis identifies a failure point the literature has largely overlooked: the handoff boundary, at which a working pilot must convert into a changed decision right. The paper advances a falsifiable claim — that realized value per unit of spend tracks the depth of decision-right transfer once model quality and adoption breadth are held constant — and derives a three-archetype diagnostic (augment; delegate with appeal; delegate and log) through which boards can classify their own portfolio. The practical implication is a sequencing claim: authority must be reallocated before spend is scaled.
Keywords: decision rights; governance design; transaction cost economics; dynamic capabilities; enterprise AI; organizational reconfiguration
1. Introduction
The dominant framing treats enterprise AI transformation as a deployment problem, and measures it accordingly. Licences are counted, workflows are instrumented, adoption dashboards are green. Value outcomes, however, do not follow those counts at the rate the counts imply [insert industry benchmark, year]. The dissociation is not randomly distributed. It concentrates in firms where deployment succeeded in every measurable sense — the technology is in place, the users are active, the pilot closed on schedule — and the decision the technology was commissioned to influence did not move.
This paper argues that the unexamined variable is neither model accuracy nor user adoption, but the allocation of decision authority. Every enterprise AI use case implicitly determines who is permitted to decide, and what becomes of human accountability when the answer is wrong. That determination is a governance choice. It is frequently made by default, and it is rarely recorded anywhere. The consequence is a category of failure that current metrics cannot see, because the metrics measure activity at exactly the point where value would be lost.
The analysis proceeds as follows. Section 2 establishes the theoretical foundations. Section 3 locates the failure point in the capability sequence. Section 4 derives the decision-right test and its three archetypes. Section 5 states the implications for boards. Section 6 states the limits of the argument, including the observation that would disconfirm it.
2. Theoretical Foundations
2.1 Transaction cost economics: deployment as a governance decision
Transaction cost economics holds that the central analytic choice in organizing an activity is governance rather than technology: whether a task is performed by hierarchy, by market, or by a hybrid arrangement, and at what cost (Williamson, 1985). Applied to artificial intelligence, the implication is direct. Every use case embeds a governance structure, and that structure allocates a decision right between a human principal and a non-human agent. It carries a measurable cost structure in supervision, error recovery, and — critically for boards — reversibility.
The predictive failure of most programs is legible under this lens. Firms default to what may be termed hierarchy without authority: the system is built inside the existing control architecture, reports upward, recommends, and leaves the decider precisely where it was. Governance is nominally preserved, and value is correspondingly absent. The firm has purchased advice rather than delegated judgment — and has typically paid a premium for the distinction, since advice scales linearly while authority transfers do not.
2.2 Dynamic capabilities: where the sequence breaks
Dynamic capabilities theory specifies a sequence that firms must continually renew to remain competitive: sensing, seizing, and reconfiguring (Teece, 2007). Enterprise AI programs characteristically complete the first two. Sensing is a market scan. Seizing is a funded pilot. Reconfiguration — the harder work of dismantling a prior arrangement and standing up a different one — is where the handoff occurs, and it is the step least often resourced.
Absorptive capacity, the capacity to internalize what a pilot taught, mediates this effect. It should be understood as a mechanism rather than an independent explanation: firms that cannot learn from a completed pilot will not benefit from commissioning more of them.
3. The Handoff Failure Point
Failure at the handoff boundary presents in three recognizable signatures. The first is shadow deployment, in which the system advises and a human signs — the accountability surface is preserved while the decision surface is left unchanged. The second is automation of tasks that never carried decision authority, which produces clean efficiency gains and no change in what the firm actually decides. The third is repeated re-piloting of the same use case under new labels, which registers activity across successive reporting periods without altering the underlying allocation of rights. Comparative work on stalled enterprise programs documents the same pattern [insert comparative case study, sector]; the pattern is sufficiently consistent that it functions as a diagnostic signature rather than a coincidence.
4. The Decision-Right Test
The test is applied per use case and asks three questions. First, name the decision the system is intended to make. Second, name the decider before deployment and the decider after. Third, state the reversal cost — the time, money, and organizational trust required to restore the prior arrangement. The answers sort each use case into one of three archetypes.
Augment.
The human decides; the system improves the input. The decision right is unchanged, and returns are correspondingly bounded by the proficiency of the operators using it. This archetype is legitimate and often correct. It should, however, be labelled as what it is: a productivity instrument, not a transformation.
Delegate with appeal.
The system decides; a human may reverse on request. Authority has formally moved, but the firm retains de facto control, and the record of decisions is not systematically built. This is the most common architecture and the most frequently misread as delegation. It generates the appearance of autonomy without the compounding benefit of it.
Delegate and log.
The system decides, the decision is recorded with its rationale, and reversal requires a positive act. Only this archetype compounds, because each decision trains the record on which the next one is made. The first two plateau at operator proficiency. This distinction should be the board’s central question, because it is the one that distinguishes a portfolio that is learning from a portfolio that is merely busy.
5. Implications for the Board
Three directives follow, each of which runs counter to prevailing practice.
Reallocate authority before scaling spend. A use case without a named change in decider is a cost centre with a demonstration attached, and it will absorb additional budget and additional executive attention rather than less.
Budget for reversibility design rather than training. The binding constraint on safe delegation is the capacity to reverse — the time, money, and trust required to restore the prior arrangement — not the fluency of the workforce. Budgets weighted toward enablement address a constraint that is rarely the operative one.
Treat authority transfer as a governance act requiring explicit mandate. A delegation is an executive-level allocation of accountability, not a technology initiative that inherits accountability it was never granted. The mandate should be named, dated, and revocable. Scaling spend ahead of this allocation is the specific mechanism by which programs accumulate evidence of activity without evidence of value.
6. Limitations and Disconfirming Evidence
The argument rests on a single falsifiable claim: that realized value per unit of spend tracks decision-right depth once model quality and adoption breadth are held constant [insert empirical test, sector]. It is a demanding prediction, and it is stated here so that it can be defeated. A single firm demonstrating high value at unchanged authority would substantially disconfirm it. Boards should treat such a case as evidence, not as an anomaly to be explained away.
Two boundaries warrant statement. The account underweights cases of genuine technical unreliability, where the binding constraint is model performance rather than authority; not every stalled program is a governance failure. And reversibility is not always recoverable: once accountability has been publicly reallocated, the trust cost of restoring the prior arrangement may exceed the benefit of doing so. That asymmetry should temper enthusiasm for authority transfer, and is the strongest argument for sequencing it deliberately rather than opportunistically.
References
Teece, D. J. (2007). “Explicating dynamic capabilities: The nature and microfoundations of (sustainable) enterprise performance.” Strategic Management Journal, 28(13), 1319–1350. https://doi.org/10.1002/smj.640
Williamson, O. E. (1985). The Economic Institutions of Capitalism: Firms, Markets, Relational Contracting. New York: Free Press.

