Every image stack build I have written up so far starts from nothing. A listing, a review pool, a keyword surface, a tactic library, and the model has to produce a creative argument from scratch under a set of rules about what it may not say.
This one starts from a finished answer. The client already had a stack that worked. Three secondary images on a Barn Red fence stain, approved, locked, performing. What they needed was the same three images on Cumberland Brown, Sable Brown, Sierra, and Black, with the colour changed and nothing else broken. Twelve images, four SKUs, one afternoon.
That sounds like the easy version. It is the version where the model has the least room to be creative and the most room to be subtly wrong, because "the same, but different" is a brief that lets an error hide inside sameness. The skill is ~/.codex/skills/amazon-image-stack-rollout, 1,194 words in the SKILL.md plus a 215-word QA checklist, dated May 19. The run I am writing up is Wood Defender's Semi-Transparent colour family on June 4.
The problem, and what it costs by hand
A rollout is the most common creative request in a marketplace shop and the one least likely to get a strategist's attention. The strategy happened on the hero SKU. The rollout is production.
Done by hand, a designer takes the three approved layouts, swaps the stained-fence photography for the new colour, re-shoots or re-composites the product pail with the new label colour, changes the colour name in the type, and checks that the swatch actually looks like the swatch. Per colour, per image. In our shop that is roughly a day of designer time for a four-SKU batch, and I am deliberately not going to turn that into an annual savings figure because I don't have run logs across a team and a made-up efficiency number is exactly what I would call out in somebody else's post.
The real cost is not the day. It is that the day never gets scheduled. The hero SKU gets the strategist, the siblings get "we'll roll it out when there's time," and there is never time, so four colours of the same product run three different stacks for a year and the brand's own catalogue reads as four brands in the search grid.
The build
No API. No scraping. No compositor. The rollout skill runs inside Codex with the built-in image generator, on the same subscription seat as everything else. The run folder records no per-image token cost because there is no per-image line; the marginal cost of a batch is the seat and the strategist's attention, and the strategist's attention is the expensive one.
The skill takes two required inputs: the final PNGs from the approved stack, and a batch list of target SKUs as Amazon URLs, ASINs, or source packets. The first rule in the workflow is the one the whole thing rests on:
The source of truth for the visual system is the final PNGs from the approved stack, not memory or a vague description.
That line exists because the first version of this workflow took the original prompts as the source, and prompts describe what somebody asked for, not what got approved. The approved images are downstream of three rounds of client notes. The prompts are not.
From the PNGs the skill extracts a brand-image-stack-system.md before touching any target SKU. On the Wood Defender run that file came out at eight sections: overall impression, palette and contrast behaviour, typography feel, layout patterns per image role, product staging, copy tone and claim discipline, SKU-specific colour direction, and reusable versus SKU-specific elements. The layout section reads like a spec:
Stacked headline left:
CONTRACTOR'S/CHOICE/SEMI-TRANSPARENT/[COLOR NAME]/FENCE STAIN & SEAL
The bracketed colour name is the entire rollout in one token. Everything outside the brackets is the brand. The thing inside the brackets is the SKU.
Each target then gets its own creative brief and its own prompt file. Here is the image-03 prompt for Cumberland Brown, verbatim from the run:
Create a square 2000x2000 upload-ready Amazon secondary image for Wood Defender Cumberland Brown color swatch. Full-bleed design, no screenshot frame, no rounded border, no white outer buffer, no drop-shadow card. Background is a close-up of three vertical rough cedar pickets stained Cumberland Brown: warm amber-brown, golden brown undertones, visible wood grain and knots. Top header is bright Wood Defender yellow with the Wood Defender-style red/white logo at top left and a slanted black/red stripe under it. Large stacked condensed uppercase text: "CUMBERLAND BROWN" in white, "COLOR SWATCH" in yellow, blue bar "SEMI-TRANSPARENT FENCE STAIN". Add a yellow outlined callout near bottom left reading "WOOD GRAIN VISIBLE" with a small woodgrain icon. Add a blue slanted wedge at the lower right. No extra claims.
Read the second sentence again. "No screenshot frame, no rounded border, no white outer buffer, no drop-shadow card." That sentence is in all twelve prompts, and it was not in the skill. It is the fossil from what broke.
What broke: the brand element that was a screenshot
The skill says the final PNGs are the source of truth. The client sent screenshots of the final PNGs.
Not maliciously. They opened the approved images in a viewer, took screenshots, and attached them, which is how most people send an image they already have. Each screenshot carried the viewer's chrome: a white outer margin, a rounded light-grey card, a soft drop shadow. And the extraction step did exactly what it was told. It treated the supplied file as the approved artwork and wrote this into the brand system document, under Palette and Contrast Behaviour:
The outer Amazon frame is white with a light gray rounded border and a subtle shadow.
There is no outer Amazon frame. There is a macOS preview window. The model had faithfully catalogued a screenshot border as a brand colour decision, and if the run had continued from there, twelve upload-ready images would have shipped with a fake rounded card baked into a 2000-pixel canvas, which on Amazon renders as a product image inside a smaller product image with a grey margin around it.
The client caught it on the first proof. The fix has three layers, and only one of them is in the skill.
The run-level fix is the "client correction" paragraph now sitting at the top of both the brand-system file and the batch brief, and the no-frame sentence in every prompt. The cheap fix, and it worked.
The skill-level fix is the one I still owe the file: an input gate that asks whether the supplied source images are the exported artwork or a capture of it, and refuses to extract a system from a capture. The QA checklist already says "final image is true upload-ready artwork, not a screenshot/mockup." That check runs at the end. The failure happened at the beginning, and a rule that catches an error after twelve generations is a rule that costs twelve generations.
The general lesson is the one I keep re-learning in different costumes: the model cannot tell the difference between the artefact and a picture of the artefact, and it will never ask. Source-of-truth rules only work if the thing you hand over is actually the source.
What broke: the smaller things
The generator returned undersized squares. The built-in tool came back with square images below 2000 pixels. The skill anticipates this and permits exactly one kind of post-processing: "deterministic resize/crop/pad is allowed only to enforce exact dimensions after a generation." So the run resized and logged it. The rule matters because the alternative is a model that decides to "fix" a small image by regenerating it, which changes the artwork, or by compositing it onto a canvas, which reintroduces the border problem from the other direction.
Server errors mid-batch. The generation-status file notes the tool "resumed successfully after the earlier server errors." Twelve images across four SKUs, and the run stalled partway. What made that survivable is that every SKU has its own folder, brief, and prompt file, so a resume is a folder check rather than a memory exercise. A batch job with no per-unit artefacts has to start over when it dies. This one picked up at the next missing PNG.
One reroll. Black is the hardest colour in a system built on a dark left-side gradient, because the colour name and the gradient are the same colour. The first Black Real Results render put "BLACK" into the shadow and it disappeared. Rerolled once for legibility. The QA line that caught it, "color name is not lost against the background," is the kind of sentence that looks obvious until you realise it only exists because a specific image failed it.
Two claim conflicts, resolved by silence. Wood Defender's website supports a three-year warranty, an up-to-six-year life expectancy with conditions, and a coverage range. The Amazon listings show a different coverage range. The brand-system file resolved this the only safe way: "Do not claim exact coverage because Amazon and Wood Defender pages show different coverage ranges." When two sources you are told to trust disagree, the correct output is not the average and not the more flattering one. It is nothing, on that claim, in that image. Same logic killed "deck" from every prompt: it is a fence stain, the website says fence, and a lifestyle render of a stained deck is a return waiting for an address.
The tension the skill doesn't name
Three weeks ago I wrote that a brand's own catalogue is where camouflage bites hardest: siblings that share every design decision become indistinguishable in a grid, and the shopper picks the wrong one or none. A rollout's entire purpose is to make siblings share every design decision. So the rollout skill is, on its face, a machine for manufacturing the problem I spent August warning about.
It got away with it on this run for one reason, and it is a merchandising reason, not an engineering one. The variable that differs between these four SKUs is colour, and the system puts the colour in two places a shopper reads first: the full-frame stained-wood background, and the colour name as the largest block of type on the image. The sameness is structural. The difference is the loudest thing in the frame.
Run the same skill on a family that differs by size, or count, or strength, and the same structure produces four stacks where the difference lives in a small word inside a headline. That is intra-brand camouflage, delivered efficiently.
The skill's only defence is the word "loosely" in its own description, and one workflow line: "do not force exact layouts when the product shape, shopper question, or category requires a better adaptation." That is a mood, not a rule. The rule I am adding this week: the batch brief must name the single variable that distinguishes each sibling from the source SKU, and that variable must be the largest type element or the dominant visual field in at least two of the rolled-out images. If the brief cannot name the variable, the batch stops, because a rollout of products that differ by nothing is a rollout of the wrong products.
The other skill that got built that morning
The folder timestamps tell a story I had half-forgotten. The rollout run started at 14:05. Between 09:48 and 12:16 the same day, a separate skill was written: wood-defender-image-stack, 3,830 words, brand-specific, with its own brand-system reference, product-line reference, and QA file.
The two skills answer two different questions. The rollout skill asks "how do I copy an approved stack onto siblings." The brand skill asks "how do I build a new Wood Defender stack from scratch without inventing a fence." Its required inputs include real client lifestyle photography, and it refuses to generate finals without it. That requirement came from the same afternoon: the rollout's real-results frames were built from supplied jobsite photos as colour truth, and the difference between a rendered fence and a photographed one is visible to anyone who has stained a fence, which is the entire customer base.
A generic rollout skill got a brand out of a four-SKU backlog in an afternoon. A brand skill turned the lessons from that afternoon into a permanent constraint. Most of the value in the second one is a sentence that says "do not proceed without real photographs."
Time and cost, honestly
The run folder timestamps put the first source file at 14:05 and the last final PNG at 16:10, with a colour reference sheet built at 16:34 to check the four colours side by side. Call it two and a half hours for twelve upload-ready 2000-pixel images, including the server stall, one reroll, and one round of client correction that would have cost a full re-brief on a manual job.
No per-image cost, because the generator is the built-in tool on a subscription seat. No cron, no daemon, no unattended write path. A human invoked it, a human read the proofs, a human uploaded the files. The judgment about whether Sable Brown was reading too orange did not get cheaper, and it is the part the client is paying for.
What you can replicate
- Hand the extractor the artefact, not a picture of it. Export the approved images and attach the exports. If you cannot tell whether a file is the artwork or a capture, neither can the model, and it will catalogue the capture.
- Extract the system before touching a target. One document that separates brand from SKU, written once, reused across the batch. If your first prompt for a sibling is written before that document exists, you are re-deriving the brand per image and it will drift.
- Put the sibling variable in brackets. Literally.
[COLOR NAME],[SIZE],[COUNT]. Everything outside the brackets is locked. If you cannot name what goes in the brackets, stop. - Make the bracket the biggest thing in the frame. A rollout that changes a small word is a camouflage generator. The differentiator gets the largest type or the dominant field, or the stack is not doing its one job.
- One folder per SKU, one prompt file per image. Batches die mid-run. A job with per-unit artefacts resumes at the next missing file. A job without them resumes at the beginning.
- When two trusted sources disagree, claim nothing. Not the average. Not the better one. The absence of the claim is the only output that survives a compliance read and a customer with a tape measure.
The run produced twelve images that shipped. The most useful thing it produced was a sentence in a brand-system file describing a border that did not exist, because it is the clearest example I have of a model doing exactly what it was told with exactly the wrong input, and neither of us noticing until somebody who knew what the real images looked like opened the proof.