The question you have been asking your agency about AI is out of date as of this week.
"Do you use AI" returned no information in 2026 โ everyone says yes. "Which model do you pin" was better, and I've been telling people to ask it since July. As of Wednesday there is a sharper one, and it has a factual answer for the first time: can you produce the session record for the work your agents did on my account, and does it name a human?
That question got teeth this week because Anthropic shipped both halves of the thing in a nine-day window, and the industry is only going to write about one of them.
What happened
On August 19โ20, 2026, Anthropic moved computer use to general availability as computer_toolset_20260801, launched a new browser use tool (browser_toolset_20260801) for driving hosted browsers, and took the Files API and Agent Skills / Skills API to GA on the Claude Developer Platform (Anthropic release notes).
In the same release: Enterprise Admin API user-management endpoints reached GA, Managed Agents gained allowed_domains / blocked_domains restrictions on their web search and fetch tools, and the Console session viewer was rebuilt with an Inspector panel carrying session cost, raw events and per-tool statistics. Eight days earlier, on August 11, the Compliance API was extended to return transcripts of Cowork and Claude Code sessions โ a single server-hosted record per session containing prompts, responses, tool call content, a verified user ID and email address, and timestamps.
Why most brand owners will read this wrong
The dumb take is the exciting one: agents can drive browsers now, so agents are going to run Seller Central.
They aren't, not this quarter, and not in any way that touches your P&L before January. Hosted-browser agents are slow, expensive per task, and terrible at exactly the thing marketplace work demands, which is knowing when a page didn't load versus when a field was genuinely empty. I've built enough of these to say plainly that the interesting failures are still the ones that come back beautifully formatted and confidently wrong.
The real signal is the other half of the same release, and it isn't a capability at all. The week the agent got hands, the platform shipped the paperwork. Domain allowlists so an agent can only reach sites you named. Hard spend caps per session, added August 7. Per-tool statistics and per-session cost. And a compliance endpoint that returns a transcript of what the agent actually did, attached to a verified human identity, with timestamps.
For two years the honest answer to "who touched my catalog" was a shrug wearing a process diagram. It is now a supported API call.
Except โ and this is the part that decides whether it matters to you โ almost none of that exists on an individual subscription seat. The Admin API is Enterprise. The Compliance API is Enterprise. The session viewer, the domain controls, the spend caps, the tool statistics all live on the Developer Platform. If the AI work on your account is being done by a person on a personal Pro or Max plan โ and across the agencies I know, that is where the overwhelming majority of it happens โ none of this shipped for them. Nothing about their work became more auditable on Wednesday.
That is not an accusation. It's an architecture fact, and it's the whole reason the RFP question is worth asking.
What actually changes for a brand doing $200K a month
1. The agency question moves from capability to evidence. Every agency in this market will have "agentic" on a slide by October. That slide has never cost anyone anything to produce. Asking whether they can produce a session record is not a gotcha โ plenty of good shops will say no โ it's a fast way to find out whether anyone there has thought about the write path into your account. A shop that answers with a plan ("our automation runs on the API under a service identity, here's who reviews it") has thought about it. A shop that answers with a paragraph about responsible AI has not, and that reply is itself the answer.
2. You are the party carrying the obligation, and until this week nobody in your vendor chain had a product that could produce the evidence. Amazon's Business Solutions Agreement added Section 19 covering Agents effective March 4, 2026 โ automated tools accessing Amazon Services must identify themselves as automated. That obligation sits on the seller account. Your repricer, your PPC automation, your listing tool, and now potentially a hosted browser your agency is driving are all in scope, and the account-health consequence lands on you, not on the vendor who built the thing. As of this week the vendor side finally has a way to show you what it ran. Ask for it.
3. AI cost is now attributable per client, and it is going to disappoint the people hoping otherwise. The Inspector panel gives session-level cost and token use. For the first time an agency can genuinely answer "what did AI cost on my account this month" instead of estimating. I'd expect that number to land where I've said it lands all year: inside a $6,000 retainer, the AI spend attributable to your specific account is low hundreds of dollars a month, sometimes double digits. Three separate repricing events in eleven weeks โ a tokenizer change in May, a competitor cutting a tier 80% in July, a scheduled increase quietly cancelled this month โ and not one of them moved a rate card in this industry. The cost in agency work is the person who checks the output, and that person did not get cheaper this week either.
4. The allowed_domains control is the most useful thing in the release and nobody will cover it. Being able to say an agent may reach these domains and no others is the difference between a research tool and a research tool that wandered onto a forum and came back with a competitor's blog post dressed as a finding. If you run any of your own automation, that's a one-line configuration change with a real quality benefit, independent of any governance argument.
What I'd do this week
Add one line to the vendor question set. You should already be asking two: do you pin your model version, and does your delivery process preserve file metadata. The third is now: what identity does your automation run under on my account, and can you produce a record of what it did. One email. The reply resolves it.
Write down which of your tools touch a write path. One page. Tool, what it can change, whose credentials it uses, which named human is responsible. Most operators cannot produce this list, and it takes an afternoon. Add a column for whether it identifies itself as automated under Section 19, because that is the one item on the page with an Amazon enforcement date behind it rather than a vendor's pricing page.
If you run your own agent work, set domain restrictions and session spend caps now. Both shipped this month, both are configuration, neither requires a migration. A hard spend cap on a session is the cheapest insurance available against a loop you didn't anticipate.
Do not migrate anything to a hosted browser agent before January. Ten weeks from peak is the wrong moment to put a new class of tool on a live catalog write path. If it's interesting, test it against a fixed set of twenty of your own real SKUs in Q1, when a wrong answer costs you a Tuesday instead of a week of Q4.
Declare an integration freeze. Mid-October to mid-January, no new write-access tools. Write it down and tell your agency. This costs nothing and removes the single most common way a good quarter gets broken by somebody being helpful.
What I'd ignore
The computer-use benchmark discourse. There is no published evaluation that measures whether an agent writes a bullet point that won't get your listing suppressed. That is still the only benchmark that bills you when it's wrong, and you still have to build it yourself out of your own SKUs.
Anything sold as an "agentic readiness audit." It's a one-page tool inventory and three emails. I've now published the whole thing twice for free.
The "agents will replace agencies" cycle, and its mirror image. Both are content. The thing that actually moved this week was the ability to see what an agent did after it did it, which is neither a threat nor a defence โ it's a filing cabinet, and filing cabinets are how these arguments eventually get settled.
The urge to switch providers over a GA announcement. Migration is a config string. Re-validating your guardrails against a golden set is the expensive part, and if you don't have that set, a GA release is a bad week to discover it.
Four times this year I've written about something moving underneath an operator who did nothing wrong: a model swapped silently on a Friday, a price that expired with no memo, an increase cancelled in an undated note, a shutdown date buried in a changelog. Every time, the only durable fix was the same shape โ find the artifact, write it down, put a date on it.
This week is the first one where the artifact got easier to obtain rather than harder. The capability half of that release will get all the coverage, and the paperwork half is the part that changes a conversation you're actually going to have. Ask your agency for the session record. Not because you'll read it. Because the answer tells you whether anyone over there knows what their software is allowed to touch on your account.