Two failure modes appear again and again. In the first, every AI action requires review, so the organization gains no capacity and quietly stops using the system. In the second, actions flow without oversight until something with real consequence goes wrong, and the business discovers it accepted a risk nobody explicitly agreed to.
The line between them is not a matter of temperament. It can be decided deliberately, written down, and revisited as trust is earned.
Four tests for any action
For each action the system can take, work through four questions. Any single strong answer is enough to require a named human.
- Legal exposure — could this create, alter, or waive an obligation? Contracts, commitments, and regulated communications always hold.
- Financial consequence — does money move, or does a figure enter a record that money will later move against? Set a threshold rather than a feeling.
- Relationship risk — would a customer, employee, or vendor reasonably expect a person behind this? Apologies, terminations, escalations, and pricing conversations qualify.
- Reversibility — if wrong, can it be undone quietly within an hour? Irreversible or publicly visible actions hold.
Approval is not a measure of how much you distrust the system. It is a statement about which consequences the business insists on owning.
Three postures, not two
Treating oversight as a binary is what forces the choice between paralysis and exposure. There are three useful postures, and most actions belong in the middle one.
Proceed
The action runs and is logged. Appropriate for internally visible, reversible, low-consequence work: classifying a document, retrieving context, updating an internal status, preparing a draft.
Prepare and hold
The system does all the work and stops at the last step. A person reviews something that is already complete and either releases it or corrects it. This is where most customer-facing and financial work belongs, because it captures nearly all the time savings while keeping the decision human.
Escalate
The system recognizes it is outside its defined boundary and routes to a named person with the context it has gathered. Exceptions, ambiguity, and anything the design did not anticipate belong here — and the rate of escalation is a useful health metric in itself.
Name people, not roles
An approval assigned to a department is an approval nobody performs. Each hold point should map to a named person with a named backup, a stated expectation for turnaround, and a defined behavior when that time passes — escalate, or continue holding. Silence is not a decision the architecture should permit.
Let the boundary move with evidence
The right approval line at launch is not the right line a year later. Review held actions periodically: where reviewers approve without changes at a very high rate over a meaningful volume, the hold may be relaxed — deliberately, with a date, and with sampling in place afterward. Where corrections are frequent, the problem is usually the design of the action rather than the reviewer.
Governance done this way is not a brake on the system. It is what allows the business to widen the system's authority over time without ever guessing about the consequences.
If this is a live question in your business, schedule an AI strategy conversation.