Skip to content
Back to the workshop

Two Workers, One Frame

Check-then-act is a race, and money makes it expensive

E
EugeneBuilding Cleo
4 min read

Cleo can build a short video ad as a storyboard: a row of frames, each rendered as a still image before the clips are made. Rendering a frame costs real money, so it costs the user energy, the unit Cleo uses to price work. Each frame carries a status, and the render path followed a pattern you will find in a great many codebases.

Read the frames. Pick the ones that are idle. Mark them as rendering. Render them.

It works in every test that runs one request at a time. It fails when two requests arrive together.

The window between the read and the write

Two render calls for the same storyboard do happen. A double press, a retry from an unreliable connection, the AI and the user both deciding the frames need doing at the same moment. Each call reads the frames and sees the same idle ones, because neither has written yet. Each marks them as rendering. Each renders them. The user pays twice for the same picture, and one of the two results is simply discarded.

This is the classic time-of-check to time-of-use race. The check (is this frame idle) and the action (mark it as mine) are separate operations, and the world is allowed to change between them. Nothing in the code was wrong line by line. The bug lived in the gap between two correct lines.

Make the write carry the check

The fix is to stop treating the status write as a record of a decision already made and to make it the decision itself.

The rendering mark is now a conditional update. It says: set this frame to rendering, but only if its status is still the one I read. In SQL terms, the assumption from the read moves into the WHERE clause of the write. The database applies these one at a time, so exactly one of two racing calls finds the frame still idle and wins. The other's update matches no rows.

The losing call then has to do something sensible with that result. It drops the frame from its own list of targets, so it does not render it. And because energy was reserved for the whole batch up front, the share for the frame it lost is refunded immediately rather than at the end of the run.

That last part needed more care than the race itself. A batch can be interrupted partway. Some frames rendered, some failed, some were lost to another caller. The refund arithmetic has to keep money only for renders that actually succeeded, return it for everything else, and never refund a lost claim twice: once when the claim was lost and again when the interrupted batch is reconciled. I wrote the tests for that path first, because it is the kind of arithmetic that looks right and is not.

Status columns are claims

The way I think about it now is that a status column on a piece of work is a claim, and claims have to be taken atomically. If a row says rendering, some process owns that work, and the only safe way to become that owner is a write that fails if somebody else got there first.

The pattern reaches well beyond video. Any job that is picked up by reading a status and then setting it has this window. Sometimes the consequence is harmless duplicate work. When the work costs money, sends something to a customer, or posts something in public, the duplicate is the bug.

There are several ways to close the window: row locks, advisory locks, a queue with exactly-once delivery, a unique constraint that refuses the second insert. For a status transition on a row that already exists, the conditional update is the lightest of them. No lock held across a network call, no new infrastructure, one extra predicate on a write that was already there. The database had always been willing to arbitrate. I had simply not asked it to.

E

Written by Eugene

Building Cleo, an AI marketing operating system. These posts cover the architecture decisions, technical challenges, and lessons learned along the way.

More from the workshop