Never let an image model write your headline.
The pitch is hard to resist: describe the whole screenshot (the scene, the glow, the headline in big white type) and get a finished frame back from one call. It even works now. Google describes its current Flash image model as having “reliable text rendering”, and for a short English line it usually earns that. Usually is the problem. Once a word is painted into an image it stops being text. You can’t spell-check it, edit it, restyle it or translate it. You can only regenerate it and hope.
Why spelling was ever hard
A text-to-image model never sees your headline as letters. The prompt goes through a text
encoder first, and most encoders read subword tokens, chunks like track or
habit, the same way a chat model does (we measured how unevenly that chopping
falls across languages in an earlier post). To paint
a word, the model has to recall what that chunk looks like as a row of glyphs, something it
was never directly told.
A 2022 study tested exactly this. Image models built on encoders that read raw characters spelled better across the board, with gains of more than 30 points on rare words. Newer models have closed much of the gap, but the weakness still has the same shape: common words come out fine and rare ones wobble. The rarest word on your listing is usually the most important one, too: your app’s name, which you probably invented.
“Mostly right” compounds badly
Nobody publishes a spelling accuracy for your headline in your style on your background, and I haven’t measured one either. So pick any rate you believe and run it through the arithmetic. A listing isn’t one image. It’s a set of six, say, repeated in every locale you ship, and every one of them has to be right.
| Accuracy per image | One set of 6 clean | 6 frames × 10 locales clean |
|---|---|---|
| 99% | 94% | 55% |
| 95% | 74% | 5% |
| 90% | 53% | 0.2% |
Look at the top-right cell. Even at 99% per image, far better than “usually”, a ten-locale listing has close to even odds of shipping at least one typo. At that volume a misspelled headline isn’t bad luck. It’s the expected outcome.
Catching them is harder than it sounds. A typo in live text is a string you can diff, spell-check or search. A typo in pixels needs a pair of eyes, or an OCR pass, which is another model with its own error rate. For the locales you can’t read yourself, you are shipping words you have no practical way to check.
Pixels can be redone, not edited
The cheapest loop in store optimisation is copy: change a headline, ship it, see what moves. Both stores will even run that test for free. Bake the headline into a generated image and every copy change becomes a new generation. Image models sample just like text models do, so a fresh generation brings a fresh background with it. The glow you approved moves, the colour drifts, and the frame no longer matches the other five.
Editing has improved, and you can hand the image back and ask for one word to change. What comes back is still a whole new image, though, and nothing in the API promises that the pixels you didn’t mention survive untouched. You end up re-checking the entire frame to change one word.
The type stops being yours
A headline in a design system is a specification: a family, a weight, a size, tracking, a line break, a colour with a known contrast ratio. A painted headline only has a description of one, such as “bold white sans-serif”, and the model fills in the rest. Across six frames it fills them in slightly differently each time. The a in frame two isn’t the a in frame five, the weight breathes and the baseline wanders. Each frame passes on its own, but side by side the set looks homemade.
You also lose the small decisions that make type good. You can’t choose where the line breaks. You can’t lighten the weight to compensate for white on dark. And your layout can’t measure a headline it didn’t set, so nothing downstream knows where the words end and the device frame can begin.
Backgrounds forgive upscaling. Letters don’t.
Gemini’s current image models default to 1K output, which for a tall 9:16 frame is 768 × 1376 pixels. An iPhone 6.9″ screenshot is 1320 × 2868. Filling that height means stretching the generated image by about 2.1×.
A gradient, a cloud of light or a softly focused scene survives that stretch without anyone noticing, because there were no hard edges to lose. Letterforms are made of nothing but hard edges. Scale a glyph up twofold and every crisp edge turns into a short ramp of grey (the same smearing that turns a 1px line grey). The headline, the one element that has to be sharp, ends up the softest thing in the frame.
You can request 2K instead, which gives 1536 × 2752, close enough to native. On Google’s price list as I write this, that’s $0.101 an image instead of $0.067, half again as much, and you pay it on every regeneration. Live text costs nothing at any size, because the renderer draws it at the export’s native resolution for whichever slot it’s filling.
The pattern: paint the plate, set the type
None of this means you should stop using image models for backgrounds. They’re excellent at atmosphere. Split the job along the line where each tool is reliable: the model paints a text-free plate, and your renderer sets the words on top. Four rules make the split work.
Ban text in the prompt, and say why. ShotCanvas’s backdrop prompt lists “no text or lettering” among its hard constraints, next to logos and phones. It also names the space the words will occupy: “A large headline will be overlaid onto the TOP ~30% afterwards — it is NOT part of this image, never paint it.” Then it asks for that region to be the quiet part of the scene itself (sky, haze, distance) rather than an empty block. Our reasoning is that a calm band across the top of a tall poster is exactly where posters put words, so we would rather tell the model the words are coming than hope it leaves the space alone.
Set the headline live, at export resolution. Use your font, your weight and your line break, and measure the text before layout so the device frame knows where the words stop. One plate then serves every locale. Ten locales become six generations plus sixty strings you can proofread, instead of sixty generations you can’t.
Review the composite, not the plate. A plate can look beautiful on its own and still put its brightest highlight directly behind the headline. In the ShotCanvas app, a vision model checks the composed frame (backdrop, device and headline together), and if it finds a real flaw the backdrop is repainted once with that feedback before you ever see it. For the readability side of this, including scrims, shadows and placement, see how apps make text readable over any image.
Ban incidental lettering too. Text inside the scene, like a shop sign or a label on a jar, rarely says anything real, and it competes with your actual headline for the first glance. A model told “no text” should hear that as no text anywhere, not just none at the top.
Where ShotCanvas draws the line
ShotCanvas generates bespoke backdrops on exactly this split. The image model paints the atmosphere, and your headline is set live on top in the template’s type, so you can rewrite it, translate it or swap the font without touching the art. Every slot exports at its native size with the words drawn crisp.