One handoff flag was answering two different questions
We had one handoff flag doing two jobs: deciding whether the AI could speak and whether a human still owed the shopper a reply. Splitting AI Paused from Needs Human turned a fallback into a real lifecycle.
We thought human handoff needed a state that meant the AI had stepped aside. That state existed. It was called aiPaused. The problem was that we were using it to answer two different questions: may the AI speak right now, and does a person still owe this shopper a reply? Those questions sound close enough to share a flag, but in practice they belong to different parts of the conversation. Treating them as one state made it impossible to change one without changing the other.
That was the point where human handoff stopped looking like a fallback button and started looking like a lifecycle.
Pausing the AI is not the same as completing the handoff
A shopper asking for a person creates at least two responsibilities. The assistant has to know whether it should continue answering, while the merchant side has to know whether someone still needs to deal with the request. Our single flag answered both. If aiPaused was true, the AI stayed quiet and the conversation appeared to need a person. If it became false, the AI could speak again and the human-work signal disappeared with it.
That coupling was convenient until we needed those two facts to move independently. A conversation can reach a point where allowing the assistant to speak again is useful while the merchant still owes the shopper a human response. Resuming automation does not prove that the request for a person was handled, but with one flag the system had no way to say both things at once. Turning the AI back on also meant telling the merchant queue that the work was gone.
That was not a model problem. It was a state problem.
We split "AI Paused" from "Needs Human"
The fix began with vocabulary. We kept AI Paused for the question it actually answers: may the assistant speak? We gave Needs Human the other job: does this conversation still require a person? That distinction was written into the decision record and into the product vocabulary so it would not survive merely as an implementation detail somebody had to rediscover later.
The important part was not renaming a variable. It was allowing two facts about the same conversation to be true independently. An AI system can be allowed to answer while a human obligation remains open. An AI system can also be paused without that pause being proof that a person has arrived, read the thread or resolved anything. Once those states are separate, the product can stop pretending that silence means service.
A handoff button hides the difficult part
From the shopper's side, human handoff can look simple: there is a request to talk to a person, and there is an outcome where a person helps. Between those two points are state transitions the interface has to represent honestly. Somebody has to decide whether automation continues, whether the conversation belongs in a human queue, and what event actually clears that obligation.
A single "Talk to a human" button does not answer any of those questions. It only starts them. The same is true of a fallback sentence such as "I'll pass this to the team." The sentence describes an intention. It does not prove that somebody has taken responsibility for the conversation.
That difference matters in an AI shop assistant because automation is still present after escalation. It can potentially resume. The shopper can continue sending messages. The merchant can enter and leave the thread. The conversation does not stop existing simply because one system decided not to answer the next message. The handoff therefore needs state that survives those changes.
Silence is a behaviour. Human attention is an obligation.
Separating the two states gave us a useful way to think about the system. AI Paused controls behaviour, while Needs Human records an obligation. The first tells the runtime what it may do. The second tells the merchant side what still needs doing.
When one field tries to represent both, every transition becomes overloaded. Resuming the AI starts to mean "nobody needs to respond anymore". Pausing it starts to mean "a person is now handling this". Neither conclusion necessarily follows from the action itself. That is how a handoff feature can look correct in code and still lie in the interface.
A boolean is attractive because it makes the transition easy to describe: AI on, AI off. The conversation is not binary. A shopper can still be waiting for a person even if the assistant is allowed to speak again, and an assistant can be paused before any person has actually arrived. Those combinations are not edge cases. They are exactly the states a real handoff creates.
The state machine starts with asking separate questions
We did not need a large diagram to discover this. We needed to stop asking one field two questions: can the assistant speak, and does a person still owe the shopper a reply? Once those questions have separate answers, later behaviour has somewhere honest to attach. The runtime can decide whether to answer without silently clearing merchant work. The merchant queue can keep a conversation visible without requiring the AI to remain silent forever.
That is the difference between an escape hatch and a handoff lifecycle. An escape hatch says what the AI should stop doing. A lifecycle also records what still has to happen afterwards. For us, the important change was not adding more automation around escalation. It was recognising that "AI stopped talking" and "a human handled it" had never been the same event.