It's Never the Copy

Akhil Agrawal · June 5, 2026

Across 7.5 million strictly cold emails last year, the average reply rate was 0.45%. One reply per 222 sends. Gong's larger dataset says the average rep now sends 344 cold emails to book a single meeting.

And when a sequence dies, the post-mortem almost always starts in the same place: the subject line.

Here's the claim this whole post defends: reply rate is a targeting metric before it's a writing metric. A cold message is a bet on four variables at once, the right person, the right pain, the right moment, the right ask. The words are the last mile. When the bet fails, copy is usually the least broken variable, and it's the one everyone fixes first because it's the one you can fiddle with for free at 11pm.

There's a correct debugging order. It runs deliverability, then targeting, then timing, then offer, then copy. Most people run it in reverse.

The five-layer debug ladder for dead outbound

Layer 1: did it even arrive?

In GlockApps seed testing for Q4 2025, Gmail put 57% of test emails in the inbox. Outlook managed 45%. Even fully legitimate, opted-in marketing email now averages 83% inbox placement, with 10.5% landing in spam and another 6% simply vanishing. Cold email does worse.

Sit with that before you touch a word of copy. If placement is against you, a meaningful chunk of your "ignored" outreach was never seen by a human at all. You're A/B testing subject lines on an audience of spam filters.

The rules also hardened recently. Google and Yahoo's bulk-sender requirements (February 2024) put anyone sending 5,000+ emails a day under a 0.3% spam-complaint ceiling with mandatory SPF, DKIM and DMARC. Microsoft followed in May 2025: fail authentication at volume and you go to junk, then get rejected outright. This is the machinery that decides whether layers 2 through 5 get to exist.

One more trap: you can't diagnose this layer with open rates anymore. Apple's Mail Privacy Protection auto-fires tracking pixels, so roughly half of tracked "opens" are machines, not people. Belkins, an agency sending millions of cold emails a year, turned open tracking off entirely in 2025. Diagnose with seed tests and placement tools, not opens.

Layer 2: was it ever the right person?

The dirty secret of list building: contact data rots while you're using it. ZeroBounce's 2026 report, built on 11 billion checked addresses, found 23% of email lists went bad during 2025, and only 62% of all addresses checked were safe to send to. (You'll see "data decays 30% a year" quoted everywhere; that number traces back to 2000s folklore. The measured 2026 figure is 23%. Close, but now you know where it comes from.)

Then there's precision. In the same Belkins dataset that produced the 0.45% average, emails to companies with fewer than ten people replied at 0.72% while emails to enterprises with 10,000+ replied at 0.22%. Same channel, same effort: a 3.3x difference purely from who you pointed it at. Woodpecker's 20-million-email dataset shows the same shape from another angle: within their platform, deeply personalised campaigns (which really means deeply researched targeting) pulled roughly double the reply rate of generic ones.

The list is upstream of every sentence you'll ever write. An ICP is a list of nameable humans, and the sharpness of that list sets the ceiling on everything below this line.

Layer 3: the right moment, not the right hour

Timing is the most misunderstood layer, because people hear "timing" and think send-time optimisation. The evidence there is genuinely weak: a review of nine send-time studies found them contradicting each other (five said Tuesday; the rest split), and Belkins' own measured "morning advantage" was 0.54% versus 0.45%. Basis points. Schedule reasonably and move on.

The timing that matters is the trigger: did something just happen that makes your message legible right now? A funding round, a new VP, a stack change, a public complaint. The published numbers here are all vendor-published, so treat them as direction rather than gospel, but the direction is consistent: UserGems reports job-change leads replying around 20% against a ~10% baseline, and Common Room's signal-based plays land in the same 2x territory. A mediocre email that arrives the week the problem became urgent beats a beautiful one that arrives at a random Tuesday, 9:04am.

Layer 4: what did you actually ask for?

Gong's analysis of 28 million sales emails found that pitching your product in a cold email cuts reply rates by as much as 57%. Their earlier study of 300,000+ emails found interest-based asks ("worth a look?") beat calendar asks in cold outreach; the specific-time meeting request only starts winning once a conversation already exists.

The size of the yes determines the reply rate. "Do you want 30 minutes with a stranger" is an enormous ask. "Is this a problem for you" is a small one. And persistence on the ask is quantified too: in Woodpecker's data, adding a single follow-up lifted total reply rates from 4.1% to 8.3%. Most founders rewrite the first email five times and never send the second one.

Layer 5: now, and only now, the words

Copy matters. Within large datasets, the difference between weak and strong copy is real: shorter emails, question asks, no pitch, one idea, all show relative lifts in the tens of percent. That's worth having.

But hold the magnitudes side by side. Copy improvements: tens of percent, relative. Targeting precision: 3.3x. A fresh trigger: roughly 2x. Inbox placement: binary, it arrived or it didn't. The layers aren't equal, and copy is the thinnest one at the top of the stack. Polishing it while the foundation is cracked is how you end up with a beautifully written email in a spam folder, sent to someone who changed jobs eight months ago, asking a stranger for half an hour.

Why everyone debugs backwards

Because copy is visible and the other layers aren't. You can see your own words; you can't see the spam folder from your side, you can't feel the 23% of your list that silently went stale, and no dashboard tells you the moment passed. Copy also flatters the ego: fixing it feels like craft, whereas fixing deliverability feels like plumbing. So the plumbing stays broken and the poetry gets a sixth revision.

Reversing the order is mostly a discipline, and it fits in one paragraph:

Before you rewrite anything, run the ladder: 1) Placement test this week, pass or fail? 2) Would this exact person, checked this month, recognise the pain in line one? 3) What happened in their world in the last 30 days that makes now the moment? 4) Is the ask smaller than a meeting? 5) Fine. Now the words.

That's the whole chapter, honestly. The rest is you resisting the urge to open the copy doc first.

Where does your outbound actually fail? If you've run placement tests or trigger-based sequences and your numbers disagree with the ones above, reply and tell me. This post is a chapter of a book I'm writing in public, and disagreement now is cheaper than a correction printed later.