Last month a client forwarded me an email that one of my agents had drafted on my behalf. "This doesn't sound like you," she said. "It sounds like every other AI email I get." She was right. The sentence structure was correct. The information was accurate. The tone was that flat, aggressively helpful, slightly over-enthusiastic register that every untrained AI defaults to. It read like a chatbot wearing a John Aspinall nametag.
I'd spent weeks building that agent's workflow โ the data it pulls, the decisions it makes, the format it outputs. I'd spent zero minutes teaching it how I actually communicate. The agent was technically excellent and stylistically invisible. Every operator I know has this problem. You build an agent that does the work, but the output doesn't sound like you. So you edit every email, rewrite every report, and tweak every draft until the time savings evaporate. You've automated the thinking but not the voice.
Training AI to write in your voice is the single highest-leverage thing most operators skip. Not because it's hard โ it takes about two hours to set up properly โ but because most people don't realize it's a discrete, solvable problem. They assume "AI just sounds like AI" and accept it. It doesn't have to.
What Is AI Voice Training?
AI voice training is the practice of providing an AI agent with enough information about your specific communication style โ your vocabulary, sentence patterns, opinions, tone, and personality โ that its output reads as authentically yours without manual editing. It's not fine-tuning a model. It's not building a custom GPT. It's structured context: a set of documents and examples that travel with every agent you build, defining not just what to say but how you say it.
The difference between a voice-trained agent and an untrained one is the same difference between hiring someone who's read your company handbook and hiring someone who's worked beside you for two years. The handbook person gets the facts right. The two-year person gets the feel right. Voice training bridges that gap without requiring two years of exposure.
Why All Untrained AI Sounds the Same
Every foundation model โ Claude, GPT, Gemini โ defaults to a statistical average of everything it was trained on. That average sounds polished, professional, and completely generic. It hedges with "it's important to note." It uses filler phrases like "absolutely" and "great question." It writes in a smooth, conflict-free tone that belongs to no one and impresses no one.
This isn't a flaw. It's a feature of the training process. The model learned to be broadly acceptable. Your job is to narrow it โ to make it specifically yours.
Most operators try to fix this with a single line in their prompt: "Write in a conversational tone" or "Sound professional but approachable." This does almost nothing. It's like telling a new hire to "be themselves" and expecting them to match your brand. They need specifics. So does the model.
The operators I know who produce genuinely on-brand AI output all do the same thing: they give the model a voice document. Not a vague style guide. A dense, specific, example-rich reference that defines exactly how they communicate.
The Voice Stack: Four Layers That Control How Your Agent Writes
I use a four-layer system to train every agent on my voice. Each layer adds specificity. Most operators only need the first two layers to see a dramatic improvement. Layers three and four are for teams and scaled operations where consistency across agents matters.
Layer 1: The Voice Identity Document
This is a single document โ I keep mine at about 600 words โ that defines who you are as a communicator. Not who your brand is. Who YOU are. It covers five dimensions:
Vocabulary range. Which words you use and which you never use. I never use "leverage," "utilize," "synergy," "holistic," or "game-changer." I do use "build," "ship," "break," "stack," and "compound." I say "costs $47/month" not "represents an affordable investment." List ten words you always use and ten you never use. That alone moves the needle significantly.
Sentence rhythm. Do you write in long, complex sentences or short punchy ones? My writing mixes both โ a short declarative sentence followed by a longer one that unpacks the idea. I rarely use three long sentences in a row. I almost never use semicolons. Write five sentences that capture your rhythm and include them verbatim.
Tone boundaries. Where do you sit on the spectrum between casual and formal? Between optimistic and skeptical? Between prescriptive and exploratory? I'm direct and opinionated โ I tell people what to do rather than listing options. I'm skeptical of hype but optimistic about building. I use "I" constantly because I'm writing from my own experience, not a corporate "we."
Opinion patterns. What do you believe that shapes how you write? I believe operators should build their own tools. I believe most AI advice is written by people who don't run businesses. I believe specific numbers beat vague claims. These beliefs color everything I write. Tell your agent what you believe and it'll stop hedging.
What you'd never say. This is more important than what you would say. I'd never write "In today's rapidly evolving landscape." I'd never open with a question I'm about to answer. I'd never use a bulleted list of synonyms when one specific word would do. Your "never" list is the fastest way to eliminate the generic AI register from your output.
Here's a condensed example of what my voice identity document looks like in practice:
## Voice: John Aspinall
Direct, opinionated, first-person practitioner. Not a thought leader โ
a builder who writes about what he's built.
ALWAYS: specific numbers, real tool names, first person, short
sentences mixed with longer ones, opinions stated as facts,
imperative verbs (build, ship, cut, test, measure)
NEVER: "leverage," "utilize," "in today's landscape," "game-changer,"
"it's important to note," rhetorical questions as openers, hedge
words (might, perhaps, could potentially), exclamation marks,
emoji in professional output
TONE: skeptical of hype, confident from experience not authority,
treats the reader as a peer operator who builds things,
assumes intelligence, doesn't over-explain
BELIEFS: build > buy, specifics > generalities, systems > one-off
hacks, compounding > shortcuts, operators > commentators
That's 150 words. It takes twenty minutes to write. And it transforms agent output from "generic AI" to "recognizably mine."
Layer 2: Do/Don't Rules
These are specific, mechanical rules about formatting, structure, and language. They're more granular than the voice identity document and easier for the model to follow because they're binary โ do this, don't do that.
My do/don't list includes about forty rules. Here are ten that make the biggest difference:
- Don't start any paragraph with "It's worth noting" or "Importantly"
- Don't use transition phrases between sections โ just start the next idea
- Do include at least one specific dollar amount, percentage, or count in every substantive paragraph
- Don't use the word "simply" โ nothing in business is simple
- Do use contractions โ "don't" not "do not," "I've" not "I have"
- Don't end paragraphs with questions unless the next paragraph answers them immediately
- Do name tools, platforms, and specific versions โ "Claude Code" not "AI coding tools"
- Don't use passive voice for anything I did โ "I built" not "it was built"
- Do lead with the conclusion, then explain the reasoning
- Don't write more than three sentences without something concrete โ a number, a name, a date, a tool
These rules are easy to test. You can read an agent's output and check each one mechanically. When I find a rule the agent keeps breaking, I move it higher in the list and add an example of the violation alongside the correction.
Layer 3: Reference Examples
This is where most voice-training guides start โ and where most operators fail. They dump ten blog posts into the context and say "write like this." The model produces a pastiche that sounds vaguely similar but misses the texture.
The fix: curate small, representative samples, not large volumes. I keep a reference file with exactly twelve paragraphs that represent my voice across different contexts:
- Two paragraphs from client emails (professional, direct)
- Two from blog posts (teaching, opinionated)
- Two from internal memos (concise, action-oriented)
- Two from social posts (punchy, conversational)
- Two from product descriptions (benefit-led, specific)
- Two from critical feedback (honest, constructive)
Twelve paragraphs. About 1,200 words total. That's enough for the model to pattern-match without overfitting. More than twenty examples and the model starts averaging across them, which ironically makes the output more generic.
For each example, I include a one-line annotation: "Email to a client who asked about pricing โ notice the directness and the specific next step." The annotation tells the model what to learn from the sample, not just what to copy.
Layer 4: Output Validation
This is the feedback loop that keeps the voice consistent over time. Every time an agent produces output I edit before sending, I ask myself: what was wrong with the voice, and which layer should have caught it?
If the agent used "leverage" โ that's a Layer 1 failure. Add it to the vocabulary list. If the agent wrote three paragraphs without a concrete number โ that's a Layer 2 failure. The do/don't rule exists but the model ignored it. Move it higher in the list or add an example. If the agent adopted a tone that doesn't match any reference โ that's a Layer 3 failure. You might need a reference example for that context.
I review voice failures weekly, the same way I review agent output quality. About once a month, I update the voice documents based on patterns I've noticed. The documents are living files, not set-and-forget assets.
Deploying Voice Across Multiple Agents
Here's where it gets interesting for operators who run more than a few agents. You can't paste a 600-word voice document into every prompt. It eats context, it's redundant, and when you update it, you have to update it everywhere.
The solution is to centralize your voice documents and reference them. In Claude Code, I keep my voice files in a project's .claude/ directory and reference them from skill files. The voice identity document, the do/don't rules, and the reference examples each live in their own file. Any skill that produces written output includes a line like:
Before writing, read and follow the voice guidelines in
voice-identity.md, voice-rules.md, and voice-examples.md.
Three files. One canonical source. When I update a rule, every agent picks it up on its next run.
For agents that serve different functions โ a client email agent vs. a blog drafting agent vs. a report generator โ I use the same voice identity document but swap the reference examples. The core voice stays constant. The register shifts for context. This mirrors how you actually communicate: you sound like yourself in an email and in a presentation, but the formality level is different.
The Three Mistakes That Kill Voice Consistency
After training voice across 30+ agents over two years, I've watched operators (including myself) make the same three mistakes repeatedly.
Mistake 1: Training on aspirational voice instead of actual voice. You write a voice document describing how you wish you sounded โ more polished, more authoritative, more articulate. The model produces that aspirational voice and now your output sounds like someone else entirely. Your clients notice. Train on how you actually communicate, not how you wish you did. Pull real samples from your sent folder, not from a branding exercise.
Mistake 2: Over-specifying and under-specifying in the same document. You write forty rules about formatting (indent lists, use em-dashes, capitalize acronyms) and zero rules about tone. The output is mechanically perfect and emotionally flat. Voice is primarily about feel, not formatting. If you have to choose, write ten tone rules and five formatting rules, not the reverse.
Mistake 3: Never updating after the initial setup. Your voice evolves. Six months ago you used a lot of sports metaphors. Now you don't. Your voice documents still reference the old style. Output sounds subtly off and you can't pin down why. Schedule a quarterly voice document review โ fifteen minutes, same cadence as your prompt library review.
FAQ
How long does it take to see results from AI voice training? About two hours for the initial setup โ one hour to write the voice identity document and do/don't rules, one hour to curate reference examples. The output improvement is immediate on the first run. The refinement from Layer 4 validation takes two to four weeks to dial in, and the voice gets tighter with every update.
Does AI voice training work across different AI models? Yes. The voice documents work with Claude, GPT, Gemini, and any model that processes system-level instructions. The specifics (vocabulary rules, do/don't lists, reference examples) are model-agnostic because they describe your voice, not the model's behavior. I've moved my voice stack between models without rewriting anything.
How is this different from fine-tuning? Fine-tuning modifies the model's weights โ it changes what the model knows. Voice training modifies the model's context โ it changes what the model considers when generating output. Fine-tuning is expensive, requires training data, and locks you to one provider. Voice training is free, requires twelve paragraphs, and works everywhere. For operators, voice training wins on every dimension.
Can I train AI to match different team members' voices? Yes, and this is the single biggest unlock for agency operators. Each team member gets their own voice stack โ identity document, rules, and examples. When an agent drafts on behalf of a specific person, it pulls that person's voice files. I've done this for three team members and the result is that clients can't tell the difference between a human-written and agent-written email โ because the voice is genuinely matched.
What if my writing isn't consistent enough to train on? That's more common than you'd think, and voice training actually helps. Writing your voice identity document forces you to articulate patterns you follow unconsciously. Most operators discover they do have a consistent voice โ they just never documented it. The document becomes both a training asset for the AI and a style guide for yourself.
Three Actions to Take This Week
-
Write your voice identity document. Open your sent email folder, read your last twenty messages, and note the patterns. Write 150-300 words covering vocabulary, tone, rhythm, opinions, and what you'd never say. Store it somewhere every agent can reference.
-
Curate twelve reference paragraphs. Pick six different contexts you write in regularly (emails, reports, posts, proposals, feedback, instructions). Pull two representative paragraphs from each. Annotate each one with what makes it representative of your voice.
-
Apply to one agent and test. Pick the agent whose output you edit most. Add your voice identity document and reference examples to its instructions. Run it on a real task and compare the output to last week's. The difference will tell you exactly how much voice consistency you've been leaving on the table.
The goal isn't perfection on the first try. It's eliminating the 70% of editing you currently do because the output "doesn't sound right." That editing time is the tax you pay for not training your agents on your voice. Two hours of setup buys it back permanently.