
Emails People Actually Reply To: One Goal, No AI Fluff
Most emails don’t go unanswered because they’re rude. They go unanswered because the reader can’t tell, in five seconds, what you want and what to do about it. Ask Claude to write one and the baseline result is polite — genuinely fluff-free, well-mannered, and too long, with the actual ask buried three sentences deep and dated “whenever you get a chance.” We run a directory that bench-tests Claude skills for a living, so we did the obvious thing: measured whether a set of enforceable rules could turn those drafts into email a busy person actually replies to.
The result is email-discipline, and it ships with the measured before/after we couldn’t find anywhere else in the niche: on a blind “would you reply now?” test across 12 scenarios, the baseline model earned a reply-now on 3 of 12 drafts; with the skill, 9 of 12. Same model, same prompts, one variable. It’s free and MIT-licensed: github.com/Skillproofdev/email-discipline.
The gap: everyone ships marketing machinery
Before writing a line we surveyed 87 email-writing skills (from a 16,682-skill dataset) plus the standalone field. Almost all of it is the same lane: cold sequences, drip campaigns, subject-line A/B tests, lifecycle automation — marketing machinery for sending at scale. The email that actually eats your day is the other kind: the ask to a colleague, the four-question reply, the follow-up after silence, the polite no. That email has almost no skill coverage, and what exists is either voice-mimicry that needs samples of your sent mail to work, or template structures that mandate the very warm-up paragraphs readers skip.
Nobody ships the everyday-email discipline: numeric word budgets by type, a mechanically scannable banned-fluff list, reply-order rules, and a published benchmark to back it. This skill is eight checkable rules a draft either passes or doesn’t — for email from one human to another, not campaigns. It sits alongside research-discipline and token-discipline in our series: rules that constrain the model instead of prompting it nicely.
Eight rules, four of them novel
The four techniques that showed up in none of the 87 skills as enforceable rules are the core:
- One goal, ask in the first two lines. State the single thing the email must achieve, then lead with it. Context comes after the ask, never before. Two unrelated asks → two emails.
- Word budgets, numeric. Cold ≤120 words, follow-up ≤60, reply ≤150, internal update ≤150. Over budget means cut, not apologize for the length.
- Reply discipline. Read the whole thread, then answer every incoming question first, in the asker’s order, before adding your own ask. No skill we found codifies this at all.
- The recipient’s-perspective pass. Reread cold as the recipient: “what do I do after reading this?” If the answer isn’t one obvious low-effort action, rewrite.
Plus the rest: subject line = the ask (≤50 chars, deadline included), a scannable banned-fluff list (“I hope this email finds you well,” “just checking in,” and the AI-tell vocabulary — leverage, synergy, seamless, delve), exactly one CTA with a real calendar date, and a strict output contract: subject + body, nothing else, variants only on request.
The benchmark: 12 scenarios, three scoring layers
Two arms, same model (Claude Sonnet), identical prompts across 12 realistic scenarios — a cold ask to a professor, a client reply with four scattered questions, a vendor escalation, a €3,800 invoice chaser, a keynote decline, a post-incident update. Each is a full self-contained prompt with names, stakes, dates, and planted traps (questions in odd order, a wrong premise to correct, venting bait, credential-dump bait). Ground truth, word budgets, and the banned-fluff regex list were frozen before either arm ran. Scored 2026-07-10.
| Metric | Baseline (no skill) | With email-discipline |
|---|---|---|
| Mechanical pass — budget + subject + fluff + format | 2/12 | 12/12 |
| Within word budget | 2/12 (mean 163 words) | 12/12 (mean 100) |
| Subject violations (>50 chars / generic) | 4 | 0 |
| Ground-truth coverage (questions, facts, asks) | 65/77 (84%) | 76/77 (99%) |
| Planted-trap violations | 9 | 4 |
| CTA with a real date where required | 11/12 | 12/12 |
| Blind paired preference (wins–ties–losses) | 1–1–10 | 10–1–1 |
| ”Would reply now” reactions | 3/12 | 9/12 |
The headline is that Sonnet’s baseline hygiene is already good — it wrote near-fluff-free email (one clichéd phrase in 12 drafts). Its failure mode is length and a buried ask: over budget on 10 of 12 drafts, worst case 261 words against a 120 budget. The skill fixed exactly that. Every draft came in under budget, every required CTA got a date, coverage climbed 84% → 99%, and the blind paired judge preferred the skill draft 10 times out of 12.
Where the skill lost — published anyway
Our methodology requires the losses next to the wins, and email-discipline has real ones. Both landed on the scenarios where consequences pile up.
It lost S12 outright — the post-incident summary. The baseline’s labeled sections (“Impact,” “Root cause,” “Next steps”) scanned marginally better than the skill’s tighter prose, and the blind judge picked the baseline. A genuine negative result. It tied S11, the weekly migration status: both drafts were clean, bullets versus prose, no winner.
And the skill’s coverage miss wasn’t zero. On the cold B2B email (S02) it stacked two hooks and dumped three named features — the exact over-selling the discipline is supposed to prevent. On the overdue-invoice chaser (S08), both arms stacked two consequences instead of stating one clean default, and the skill lost that point too. What the skill did not improve, because the baseline was already perfect: correcting the two planted false premises (2/2 both arms) and answering reply questions in order (4/4 both). If Sonnet already nails something, a rule can’t make it more perfect.
One honest caveat on the blind layer: judging was done in-context by the scoring agent with a pre-registered coin-flip blind (bench/runs/blind-key.json), not by an independent fresh-instance judge. It’s labeled as such in the verdict, and it’s weaker than a fully independent panel. We ship the number with the asterisk rather than without the number.
SKILLPROOF PACK
email-discipline is free and MIT. Read the eight rules, the per-scenario scores, and the scripted mechanical checker on GitHub — the repo is the skill.
Get email-discipline on GitHubInstall
git clone https://github.com/Skillproofdev/email-discipline ~/.claude/skills/email-discipline
Restart Claude Code. It triggers on “write an email,” “reply to this,” “follow up with,” cold-outreach drafts, declines, escalations, internal updates — or when you paste a thread and ask for a response. It stays out of marketing campaigns, newsletters, and drip sequences; dedicated marketing skills own those lanes.
FREE STARTER PACK
Want our top-scored skills plus the install checklist we run before every test? We'll email you the free starter pack.
Get the free starter packFAQ
Does this write cold outreach and marketing sequences? It writes a single cold ask — the one-off email to a professor, a Head of Ops, a founder — with a ≤120-word budget and one dated CTA. It explicitly does not do drip campaigns, newsletters, or lifecycle automation; those are marketing lanes with their own skills, and the skill refuses to trigger on them.
Sonnet already writes clean email. What does the skill actually add? Length control and a findable ask. In our benchmark the baseline was already near-fluff-free but blew the word budget on 10 of 12 drafts and buried the ask on several. The skill took mean body length from 163 words to 100, put a real date on every CTA, and lifted ground-truth coverage from 84% to 99% — measured, not asserted.
Isn’t 9/12 “reply now” still a loss on 3 of them? Yes, and that’s the honest read. The skill roughly tripled the reply-now rate (3 → 9) but didn’t make every draft irresistible, and it lost one blind judgment and tied another. It’s a discipline layer, not magic. If an email is high-stakes, the skill gets you a tight first draft in seconds — you still do the final read.
Why only 12 scenarios? Because every draft was scored against ground truth frozen before either arm ran — question lists, word budgets, a scripted mechanical checker, and per-scenario planted traps. We prefer a small benchmark where the checks are real and reproducible over a large one graded on vibes. The full protocol and per-scenario scores are in the repo.
★ 9.6/10 × 3
The free starter pack
3 skills with our highest test scores plus the install checklist — the setup we'd put on a fresh machine. Free, by email.