Blog · September 28, 2026

Which part of your AI call is actually slow?

The reflex, when a feature that calls a model feels sluggish, is to go shopping: try the smaller model, try a different provider, read a benchmark table. That is a guess dressed up as engineering. An LLM call has two clocks running inside it, they have completely different causes, and until you know which one is eating your seconds, every fix you try is a coin flip.

The two clocks

The first clock is time to first token: request sent, nothing back, and then the first piece of the answer arrives. The second clock covers everything after that — the answer arriving one token at a time until the model stops.

They are separate because the server is doing two different jobs. Producing the first token means reading your entire prompt: system instructions, examples, chat history, any image you attached. That work happens across the whole prompt at once, so it is roughly proportional to how much you sent, and it is over before a single word of the reply exists. After that the model generates sequentially — each token depends on the one before it, so they cannot be produced in parallel. That second phase is proportional to how long the answer is, and to essentially nothing else.

This is the same asymmetry that shows up on your invoice: providers charge several times more for output tokens because output tokens cost several times more to produce. Latency and cost are two readings of one physical fact.

Measure them, it takes four lines

You do not need a tracing stack for this. Stream the response and record three timestamps — when you sent the request, when the first chunk arrived, when the stream closed — then pull the input and output token counts out of the usage field every provider returns.

t0    request sent
t1    first chunk           -> t1 - t0  = prefill, scales with input
t2    stream complete       -> t2 - t1  = generation, scales with output
usage prompt_tokens: 2,140  completion_tokens: 610

Log that for a week of real calls and the diagnosis is usually sitting there in plain sight. If the gap to the first token is the bulk of it, your prompt is the problem. If the tail is the bulk of it — and for most features it is — you are waiting on text you asked the model to write. Those are not the same bug and they do not have the same fix.

Fix one: stop asking for words you throw away

Generation time is close to linear in output length. If your stream produces tokens at some steady rate, doubling the length of the answer roughly doubles the wait, regardless of how clever the prompt was. So the single cheapest latency win available to most apps is to ask for less back.

Look at what you do with the response. A feature that needs five headlines and then parses five headlines out of the reply does not need the paragraph of friendly framing in front of them, or the explanation of why each one works, or the closing offer to try again. All of that is generated at full sequential cost and then discarded by your parser. Say return only the JSON array, no prose, give a one-shot example of the exact shape you want, and set a hard max_tokens that fits the real answer with headroom. Trimming 400 tokens of preamble is a larger, more reliable win than any model swap you were considering.

The same logic kills "explain your reasoning" in production paths — a fantastic debugging tool and an expensive shipping default. Put it behind a flag you turn on when something breaks.

Fix two: know whether your model thinks before it answers

Reasoning models produce a block of internal tokens before the reply you see. Those tokens are generated one at a time exactly like visible output, billed like output, and timed like output — the user just never reads them. A feature that "suddenly got slow" after a model upgrade is very often a feature that started thinking.

Whether that is on by default depends on the exact model, and the defaults differ within a single family — Google's Flash-Lite tier ships with thinking off while its larger siblings reason by default, and both accept an explicit budget. Do not assume: look it up for the model string you actually shipped, and set the value explicitly so a future default change can't quietly cost you four seconds.

There is a failure mode here beyond slowness, and we hit it. Every Gemini call ShotCanvas makes passes a thinking budget of zero, and not only for speed: with a modest output cap and thinking left on, a model can spend the entire budget reasoning and return a truncated answer, or nothing at all. The symptom looks like a parsing bug. The cause is a token budget spent somewhere you weren't looking.

Fix three: attack the first clock with structure, not brevity

If your measurements point at prefill instead, the instinct is to hack words out of the system prompt. That helps a little and costs you quality. The bigger lever is prompt caching: send the stable part of your prompt — instructions, examples, schema — byte-identical and first, every time, so the provider can reuse the work it already did on it. Caching pays you in both currencies here, cutting the bill and the wait on the same tokens. It has one structural requirement: the shared prefix has to be genuinely identical, so anything that varies (timestamps, user names, a session id) belongs after the stable block, never woven through it. We wrote up the mechanics in the 90% discount most indie devs never use.

Images are the other prefill story: an attached screenshot is worth a large, fixed number of input tokens before you have written a single instruction, and sending it at 4K when the model downscales it anyway buys you nothing.

Fix four: stop queueing calls that don't depend on each other

Multi-step AI features accumulate latency by accident. Generate a headline, then a subtitle, then a description, each awaiting the last, and you have paid three prefills and three full generations end to end — when nothing in step three needed step two's answer. Two rules cover it. If the calls are independent, fire them concurrently and wait on the set. If they share a large prompt and produce small outputs, fold them into one call returning one structured object: you pay the prefill once, and the combined output is shorter than three separate replies with three separate wrappers.

What streaming does and doesn't buy

Streaming does not make anything faster. It makes the first clock the one the user experiences and hides the second behind text they are already reading — close to free, and worth doing wherever the output is prose a human reads as it lands.

But if your feature parses the response — JSON you validate before rendering — the user waits for the last token regardless, and streaming buys you nothing except a progress signal. That is exactly the case where output length matters double, and exactly the case where a well-designed wait state earns its keep: making a ten-second AI wait feel fast covers what to put on screen while the tokens land.

The order to work in: instrument the two clocks and read real numbers before changing anything. Cut discarded output — preamble, explanations, restated inputs. Set the thinking budget explicitly for the model string you shipped. Cap max_tokens to the real answer plus headroom. Restructure the prompt so the stable prefix is cacheable. Parallelise or merge independent calls. Then, if it is still slow, go shopping for a smaller model — with a before number to compare against.

The version of this you can skip

Most of what is above is work you do once per feature and then forget. In ShotCanvas it is already done: the AI that writes your headlines, descriptions and keyword sets runs with the budgets capped, the thinking turned off, the shared prompt prefix arranged for caching, and the independent calls running side by side — so generating a full metadata set for a listing takes about as long as generating one field.

Generate your listing copy → Why a hardcoded model name breaks