People share images with Cleo all the time. A product photo, a screenshot of an advert that is underperforming, a picture of a shopfront. The chat inlines those images into the request it sends to the model, so the model can actually look at them.
The vision API I use documents two limits that matter here. There is a limit in bytes, and there is a limit on dimensions: an image more than 8,000 pixels on a side is rejected outright. Separately, the model downsamples anything larger than about 1,568 pixels on its long edge before reading it, so pixels beyond that point cost bandwidth and buy nothing.
The chat already checked bytes. It did not check dimensions.
Bytes are the wrong fence
It is easy to assume that a size limit in bytes stands in for a size limit in pixels. For photographs it roughly does. For flat graphics it does not. A wide PNG with large areas of solid colour compresses extremely well. The image that surfaced this was about 8,500 pixels wide and a little over 700 kilobytes, comfortably inside the byte limit and comfortably outside the dimension limit.
On its own, that would have meant one failed request. What made it serious is how conversations work.
A chat with a language model is stateless on the model's side. Every turn sends the conversation history again, and if earlier messages contained images, those images are sent again too. So an image the API refuses is not refused once. It is refused on the turn it was shared, and on the next turn, and on every turn after that, because it is now part of the history every request carries. One oversized upload turns a conversation into a guaranteed failure for as long as anyone keeps using it.
Normalise before inlining
The fix has two parts.
First, every image's dimensions are now checked before it is inlined, by parsing the image header rather than decoding the whole file. Header parsing is cheap, so this runs unconditionally. Anything larger than the model will usefully read is downscaled to 1,568 pixels on its long edge. Nothing is lost: the model would have downsampled it anyway, and the request is smaller and faster.
Second, the rare image that is both beyond the hard limit and impossible to decode for resizing is not inlined at all. The turn carries a short note in its place, saying an image was shared and could not be read, with its address. That is a degraded turn, but a recoverable one. The alternative was a conversation that could never send another message.
While doing this I moved the normalisation into a shared utility. The same logic already existed in the service that describes images in the media library, written against the same API and the same limits. Two copies of a limit drift apart. One copy cannot.
Inputs that persist need stricter gates
One smaller detail along the way: the chat now detects an image's type from its bytes rather than trusting the content type reported by the storage server. Servers are sometimes wrong about this, and the vision API refuses a request where the declared type and the actual bytes disagree. That is another input which, once in the history, would fail every turn.
The general lesson is about where validation belongs. Anything that becomes persistent state, and is then replayed, needs to be checked against every limit of every consumer before it is stored, not only the limit that was most obvious. A transient input that fails costs one request. A persistent input that fails costs every request that ever touches it, and the person who shared it has no way of knowing that the photo they sent an hour ago is the reason nothing works now.