Skip to content
Back to the workshop

An Unordered Limit Is a Lottery

And the same tickets lose every draw

E
EugeneBuilding Cleo
4 min read

Cleo runs a number of scheduled jobs. One of them summarises long conversations, so that when a conversation grows past what fits comfortably in a request, the older part travels as a summary rather than being dropped. It ran every thirty minutes. It reported success every thirty minutes.

Over seven days it wrote one summary. During that week, forty-five active conversations were long enough to need one.

Green on every run

The query that chose candidates fetched conversations with a limit of fifty and no ordering. The table held around seven hundred rows.

A query with a LIMIT and no ORDER BY returns whichever rows the database finds convenient. That set is not random in any useful sense. In practice it tends to be the same rows, run after run, decided by physical layout and the plan the database chooses. So every half hour the job fetched the same arbitrary fifty conversations, found each one either already summarised or too short, skipped them all, and exited successfully.

The conversations that needed a summary were never among the fifty. No error was raised, because nothing failed. No alert fired, because the job finished. It is the silent-success form of starvation: a loop that spends all its effort on items that need none.

The same shape, nine more times

Once I knew what to look for, I audited every scheduled job. Nine others had the same pattern: a bounded fetch with no ordering.

Two were already past their limit. The jobs that refresh each workspace's summary and send the weekly summary email both fetched organisations, unordered, a hundred at a time, from a table that had grown beyond that. Whichever workspaces fell outside the first hundred would never have been refreshed or received a weekly email. Three others only looked at paying workspaces and were well inside their limit, so they were safe for now. I recorded those rather than changing them, because changing working code on a hunch is how the next incident begins.

Order by what matters

The fixes are narrow.

The summariser now takes recently active conversations, busiest first, within a forty-eight-hour activity window. A dormant conversation with a backlog is picked up within one cycle of its next message, because that message makes it recent again. Before relying on that, I checked that no active conversation was missing the timestamp the window depends on.

The two organisation jobs now order by activity. There is still a limit, and something is still deferred when it is reached, but what gets deferred is the most dormant workspace rather than an arbitrary one, and it moves up the queue as soon as it does anything.

All three now log a warning when the fetch comes back full. A saturated window is the early sign of this whole class of problem, and it should announce itself rather than hide behind a successful response.

An assumption that was true of the code

There was a knock-on effect worth recording. A day earlier I had reduced how much recent conversation history is sent with each request, on the grounds that the summariser covered anything older. That reasoning was correct about the code and wrong about production. The summariser existed. It simply was not reaching the conversations in question. The ordering fix is what made the earlier change safe in practice.

The rule

A LIMIT without an ORDER BY says you do not care which rows you get. That is occasionally true, for a sample or a quick look. For a work queue it is almost never true. The question to ask of any bounded fetch is what happens to the row that is not selected, and whether it ever will be. If the answer depends on luck, it is the same luck every time.

E

Written by Eugene

Building Cleo, an AI marketing operating system. These posts cover the architecture decisions, technical challenges, and lessons learned along the way.

More from the workshop