I Audited My Own Amazon Creative Skill Library: 11 Automations, 831 Words, and the Word 'Only' in 9 of Them
📢
← Back to Blog

I Audited My Own Amazon Creative Skill Library: 11 Automations, 831 Words, and the Word 'Only' in 9 of Them

John Aspinall · · 12 min read

I run a creative shop. Eleven of the jobs my team does on client accounts now start as a slash command — image stacks, premium one-pagers, Sponsored Brands Video packages, Q&A banks, attribute audits, AI shopping visibility audits.

Last week I opened the folder those commands live in and read all of them end to end for the first time since I wrote them. Eleven files. 831 words total. The longest is 107 words. The shortest is 48.

Then I counted something that stopped me: the word "only" appears in nine of the eleven.

That's the build log. Not the workflows — those were an evening's work and they're boring. The thing I actually built, without noticing I was building it, is a constraint layer. Every one of those files is mostly a list of things the model is not allowed to do. And every restriction in there is scar tissue from an output that would have gone onto a live listing.

The problem: this automation writes to a $200K/mo asset

Most AI workflow posts are about saving time on something reversible. Draft an email, summarize a call, clean a spreadsheet. If it's wrong you notice, you fix it, you move on.

Amazon creative isn't that. The output of these skills is a 2000x2000 PNG or a block of listing copy that gets uploaded to a detail page doing real revenue. Once a fabricated fact is baked into a pixel and that pixel is live:

  • It's not a design revision, it's a claim on a product page.
  • "Third-party tested" on a supplement image you can't substantiate is a suppression risk, not a typo.
  • A dimension that's 4mm off in a callout is a return, and returns feed the frequently-returned badge, and that badge costs more than the image ever made.
  • Nobody catches it in QA, because it looks great. That's the whole problem. Fabricated facts arrive in beautiful typography.

Doing this work manually is slow in the way that makes brands skip it. Building a secondary stack brief for one SKU — reading the reviews, pulling the return drivers, deciding which objection gets which slot, writing the copy map — is a strategist afternoon before a designer touches it. Across a 25-SKU catalog that's a project, which is why most brands do it once for the hero and never again.

So the automation is genuinely worth having. The question was never whether to build it. It was what happens on the run where the model is confidently wrong and everything looks fine.

The build: a 48-word file that isn't a prompt

Here's the entire image stack skill, verbatim. This is the real file:

Use $amazon-image-stack with the user's Amazon listing URL, ASIN,
brand/product website URL, product images, or non-live source packet.

Create a 5-7 image secondary Amazon image stack. Exclude the hero/main
image. Use listing, packaging, source-packet, and brand-site facts only.
Generate final `2000x2000` PNGs with built-in Codex image generation only.

Forty-eight words. Look at what they're doing:

  • Sentence 1 demands inputs. Not "help me with images" — a named list of acceptable evidence sources.
  • Sentence 2 specifies the deliverable and immediately carves out the hero.
  • Sentence 3 is a fact boundary: those four sources, nothing else.
  • Sentence 4 pins the generator.

There is no "you are an expert Amazon creative strategist." No tone guidance. No step-by-step workflow. The actual craft — slot sequencing, objection ranking, copy hierarchy — lives in the $amazon-image-stack skill these commands delegate to. The command file's only job is to constrain the invocation.

The stack is unglamorous: Claude Code and Codex, native image generation, browser/computer-use for evidence capture, and a folder of markdown files synced through iCloud so every machine I work from loads the same contract. That last part matters more than it sounds. The constraint isn't in my head or in a doc somebody has to remember to read. It fires automatically on every invocation, including the ones I run at 11pm when I'm tired and inclined to trust the output.

What the audit found

Eleven marketplace skills. Here's the real inventory:

Skill Words
image-stack 48
image-stack-rollout 56
premium-social-ad-pack 61
premium (one-pager) 65
qa-generator 67
attribute-fill-rate 76
sbv-veo 82
ai-shopping-audit 84
wood-defender-image-stack 92
machine-legible-infographic 93
listing-pipeline 107

Total: 831 words. Markers I counted across those eleven files:

  • "only" — 9 files. Fact sources only, this generator only, this mode only.
  • "first" — 6 files. Almost always "capture evidence first."
  • Explicit "Do not" — 3 files.
  • "Exclude the hero/main image" — 3 files.

I went through clause by clause and classified them. In most of these files, more than half the words are telling the model what not to do, what to leave out, or what to require before it starts. That's my read of my own writing, so take the classification as mine — but the counts above are just there in the files.

Then there's main-image-lab, the one long skill at 452 words, which has an explicit section at the bottom titled Guardrails. When a skill gets big enough, the constraints stop being embedded in the instructions and get their own heading. That happened by itself.

The six things I ended up constraining

Reading them together, the eleven files converge on the same six boundaries. Each one is worth more than the workflow it wraps.

1. Fact boundaries. The most repeated idea in the library. From attribute-fill-rate:

Draft missing attributes only when they are source-backed. Do not invent dimensions, materials, compatibility, certifications, ingredients, medical claims, or compliance attributes.

That's a specific list because those are specifically the fields that get listings suppressed. main-image-lab carries its own version — "Do not invent certifications, awards, ingredients, variants, origin claims, pack sizes, or claims" — and adds a second-order rule I like a lot: "Do not choose tactics that depend on unsupported facts." Not just don't state the unsupported thing. Don't pick a creative direction whose whole premise needs it. A model that can't say "clinically tested" will happily build a concept around the idea of clinical credibility if you let it.

2. Evidence before analysis. Five skills carry this sentence, word for word:

If an Amazon URL/ASIN is provided, autonomously capture visible listing evidence first using available browser/computer-use tooling when direct access is blocked or incomplete.

Capture, then reason. Without it you get analysis of a listing the model never actually looked at, which is the most convincing wrong output in the whole category.

3. The hero is carved out of every generation skill. Three skills say "exclude the hero/main image," and it's deliberate. The hero is the highest-stakes asset on the page, it has hard compliance requirements, and it's the one image where a generated approximation of the product is unacceptable rather than merely risky. So the automation isn't allowed near it. I've optimized 14,000+ hero images; that decision isn't about capability, it's about which mistakes I'm willing to be one review cycle away from.

4. Don't claim access you don't have. From ai-shopping-audit:

Do not claim direct access to Rufus, Alexa for Shopping, Amazon ranking systems, or proprietary Amazon models.

This one isn't about product facts, it's about epistemic honesty in the deliverable. An AI-shopping visibility audit is inference from public surfaces. The instant a report implies it queried Rufus directly, I'm handing a client a document that overstates its own basis — and they'll make budget decisions on it. The constraint keeps the audit honest about what it is.

5. An approval stop, written into the prompt. From premium-social-ad-pack:

Create static social ad concept previews first. Generate distinct 1080x1080 concept options... then stop and ask which direction to roll out, scale, revise, or resize. Do not generate platform-size finals until the user approves a concept direction.

The gate is in the file, not in my intentions. And main-image-lab adds the useful default: treat generated images as concept work unless product identity, label accuracy, and branding are production-safe. Concept until proven otherwise.

6. Pin the generator. Seven of these skills say some version of "built-in Codex image generation only." That word "only" is a pin. I wrote about the silent model swap problem — subscription defaults move underneath you and nothing errors. Same logic here: if a skill can reach for whatever image tool happens to be available, then the visual consistency of a client's stack depends on which tool was exposed the day it ran. One path, named, so output drift has a cause I can find.

And the client-specific version of all this, from the Wood Defender skill: "require real Wood Defender lifestyle photography and exact selected-color swatch references before generation." Stain colors are the product. A generated approximation of a color is a returned bucket of stain. So that skill can't generate until real assets are in hand.

What broke

Here's where I'll be straight with you, because the honest version is more useful than a dramatic one.

I have exactly one failure from this library that I've written up with a date and an ASIN: building the attribute fill-rate audit, v1 confidently invented attributes — "third-party tested" on a product with no such documentation, a country of origin nobody had given it. That's the incident that produced the source-backed claim rules, and you can see the fix frozen into the invocation line today: "Audit public/visible product facts separately from uploaded backend attributes" plus the do-not-invent list.

The other constraints are scar tissue from the same category of near-miss — outputs I caught in review, which is exactly why I can't hand you a dated incident report for each one. I'm not going to dress up "I read a draft, went 'no,' and added a sentence" as a war story with a timestamp. What I can tell you is the shape of every one of them, because it never varies:

The model was not wrong in a way that looked wrong. It was fluent, on-brand, well-composed, and confidently specific about something it had no source for. Every constraint in that folder exists because a plausible output is more dangerous than a broken one. A broken output announces itself. A plausible one gets approved.

The one structural failure worth naming: early on, my skills didn't distinguish "I can't see it" from "it isn't there." A visible-only pass would report a gap that might just be a field it couldn't read. That distinction is now the first sentence of the attribute skill, and it's the reason qa-generator and ai-shopping-audit both capture evidence before scoring anything.

qa-generator has my favorite constraint in the library, because it's the one that isn't about accuracy at all:

Keep answers source-backed, on-brand, concise, and clearly drafted as seller-provided content rather than fake customer submissions.

A Q&A generator will happily write you a customer asking a question and a customer answering it. Both fabricated, both reading as organic social proof on your detail page. That's not an accuracy problem, it's a manipulation problem, and it's the kind of thing an automation does enthusiastically unless you tell it not to.

What it cost, and what it saved

The workflows were an evening each, and some were twenty minutes. The constraint sentences took longer than the workflows. Deciding that the hero is off-limits, that fact sources are enumerated, that finals stop for approval — those are merchandising decisions written in prompt form, and they required thinking about which failures I could tolerate.

Runtime cost is a few dollars per SKU on image-heavy runs, less on the audit and copy skills. There's no cron, no daemon, no unattended write path. A human invokes it and a human approves the output.

On savings, I'll give you what I can defend and skip what I can't. The compression is real on the front half of the work: getting from raw evidence to a structured stack brief, a Q&A bank, or an attribute report is the part that used to eat a strategist afternoon per SKU, and it now comes back in one pass for me to correct. What I'm not going to publish is a tidy hours-saved-per-month figure across a 9-person team, because I'd be making it up, and a made-up efficiency number is exactly the thing I'd call out in someone else's post.

The number that actually justifies the guardrail work is on the other side of the ledger: one fabricated compliance claim on one live hero SKU costs more than every hour this library has ever saved me. That's not a hypothetical — it's a suppression event, or a return-rate move on a page doing five figures a month. The constraint layer isn't overhead on the automation. It's the reason the automation is allowed to exist.

What an operator could replicate this week

You don't need my skills. You need this shape:

  1. Write the constraint before the workflow. If you can't name what the output must never contain, you're not ready to automate the task. Start with the do-not list; the steps are the easy part.
  2. Demand your inputs in the invocation line. Name the acceptable evidence sources. A model with no sources doesn't refuse — it substitutes.
  3. Exclude your highest-stakes asset. Whatever your equivalent of the hero image is, carve it out. Automate the ninety percent that's recoverable.
  4. Pin your generator with the word "only." Then upgrades are a decision you make on a Tuesday instead of a change you discover in a client's stack.
  5. Put the approval stop in the file, not in your head. "Stop and ask before finals" survives a tired evening. Your intention to check doesn't.
  6. Separate "I couldn't see it" from "it's missing." Every audit automation needs this line or it will hand you a confident wrong diagnosis, which is worse than no diagnosis.

Eleven skills, 831 words, and the word "only" nine times. I set out to build automations and what I actually built was a set of boundaries with workflows attached. If you're doing this on a catalog that pays your salary, that's the right ratio.

Put AI to work inside the business you already run.

The Aspi OS Bootcamp is a 4-week live build: second brain, Claude Code workflows, Codex execution — on your real business. Starts Mon, Aug 3 · $1,500 · 12 seats.

Explore the bootcamp →

Not ready? Get the free newsletter — the AI workflows I actually ship, when they're worth your inbox.