Skip to content
Back to the workshop

A Timeout Is Not a Cancellation

Promise.race stops waiting. It does not stop the work.

E
EugeneBuilding Cleo
3 min read

Every tool Cleo runs has a time budget. Making an image, sending an email, publishing a post: if a tool takes too long, the system stops waiting and reports a failure, so one slow call cannot hang a whole conversation.

The budget was implemented the way most JavaScript timeouts are. The tool's promise raced a timer, and whichever settled first decided the outcome.

That pattern has a property that is easy to forget. Promise.race abandons the losing promise. It does not stop it.

What the timer actually did

When the timer won, the AI was told the tool had failed. Meanwhile the tool kept running, because nothing had told it to stop. The image finished rendering. The energy for it was charged. In the worst case, the email went out.

So the system's account of what had happened was wrong in the least helpful direction. It reported failure where there had been success, and an assistant told that an action failed will reasonably offer to try again. A retry after a false failure is how one email becomes two.

Give the work a way to hear "stop"

The fix had two halves: actually cancelling where possible, and being honest where it is not.

Each tool execution now owns an AbortController. Its signal is passed into the tool through the abort-signal option the AI SDK already hands to tool functions, combined with the SDK's own signal so that a client disconnecting still cancels the work as before. When the time budget runs out, the controller aborts. A tool that passes the signal on to its network calls, and checks it between steps, genuinely stops. Tools that ignore the signal behave exactly as before, which made the change safe to ship before every tool had been taught to listen.

Three outcomes, not two

The more important change was in what the system reports.

If the tool acknowledges the abort, the work stopped. Reporting a failure is now true, and the existing failure response is returned.

If the tool does not acknowledge the abort, nobody knows what happened. The work may have finished a moment after the timer, or got halfway, or never started. Calling that a failure is a guess presented as a fact. So there is now a third outcome: unknown. It is marked as not retryable, and it carries an instruction to check the actual state with a read before doing anything else. Did the email send? Look. Was the image made? Look. Only then decide whether anything needs doing again.

It sounds like a small change to a set of labels. In practice it changed what happens after a timeout, from offering a second send to checking whether the first one landed.

Where this applies

Whenever you put a timeout around an operation with side effects, there are three possible states, not two. It succeeded. It definitely did not. Or you stopped watching. Folding the third into the second is comfortable, because it keeps error handling simple, and it is wrong in exactly the cases where being wrong costs most: payments, messages, anything public.

The honest options are to make the operation cancellable, to make it idempotent so a retry is harmless, or to report uncertainty as uncertainty and verify before acting. Ideally all three. What you cannot do is let a timer decide what happened in a system it was never connected to.

E

Written by Eugene

Building Cleo, an AI marketing operating system. These posts cover the architecture decisions, technical challenges, and lessons learned along the way.

More from the workshop