← sunnyray.com Book 15 minutes

Tools / Cold Outreach Tests

Free tool · No signup

How to run a cold outreach test

Most outbound tests measure nothing, because the instrument was broken before the first send.

1 variable per arm 500+ sends before you read it 3 rounds, in order

Before anything: prove the instrument works

Take your last dozen campaigns and put two columns next to each other, bounce rate and reply rate. If the campaign with the lowest bounce is also your best on replies, and the one with the worst bounce sits at or near zero, you have your answer and it is not about copy.

Same offer, same writer, same list quality. The only thing that changed is which mailboxes the campaign sent from. That is not a copy result, that is a delivery result.

Why this ends the discussion about copy. Run a copy experiment across a pool like that and every arm returns near zero. You will conclude the ideas do not work, when in fact nothing was ever tested. Worse, you will have burned the list, so the real test can never be run on those contacts again. Fix delivery, then test.

Thresholds worth holding to: under three percent bounce is healthy, three to five percent is a warning, and anything above five percent means pause the campaign the same day. If you are averaging seven percent with a worst case near seventeen, you are not running experiments, you are running an outage.

One honest caveat about this diagnostic. If open tracking is switched off, and it usually should be, bounce is the only delivery proxy you have. The correlation between bounce and replies is strong and consistent across campaigns, but it remains a correlation. Treat it as the thing to rule out first rather than as proof of causation.

The rules that make a result mean something

Bind every arm to one clean pool

Every arm sends from the same authenticated, warmed, low-bounce mailboxes, explicitly selected rather than inherited from campaign defaults. If arm A sends from healthy domains and arm B inherits the broken ones, you have tested infrastructure and labelled it copy.

Change exactly one thing

If the subject line is the variable, the body is the control copy word for word. Two changes in one arm produce a number you cannot act on, and you will still argue about which change caused it a month later.

Make it a big swing

At realistic reply rates, detecting a twenty percent improvement needs sample sizes you do not have. Test changes big enough to move the number by a multiple. Tweaks are for teams sending millions.

Read it late

Two positive replies against four is noise. Five hundred or more sends per arm before any conclusion, and judge on contacts per positive response rather than reply rate, which counts every out-of-office.

Same days, same hours

Tuesday to Thursday for every arm. If one arm ran over a Friday and a long weekend, day of week is a second variable you did not intend to test.

Write it down before you send

Name the variable, the hypothesis and the number that would change your mind, in writing, before launch. Otherwise the result gets reinterpreted to match whatever you already believed.

Round one: four arms worth running

These four isolate the changes most likely to produce a multiple rather than a percentage. Run them together against comparable slices of the same list, five hundred to seven hundred contacts each.

Arm A. The subject gap

Variable: subject line only

The body is your existing control copy, unchanged, word for word. Only the subject differs, so whatever moves is the subject. Replace generic title-cased subjects with short, mismatched-case lines carrying the first name.

Arm B. Two emails, and a real second angle

Variable: sequence shape

Collapse a five or six step sequence to two. The second email carries new information rather than a reminder, and a PS adds a different offer angle. If your current step three opens with a line about following up in case the last one got buried, this arm is the highest-value one on the page.

Arm C. No pretence

Variable: the opener

Strip the fake-research opening completely and say so. No claim to know their business, no generated observation. This is the arm most people refuse to believe in, and it consistently beats generated personalisation.

Arm D. Value first, no pitch anywhere

Variable: the offer itself

The opener contains no pitch at all. Offer something genuinely useful that costs you a few minutes per reply and nothing per send, ask for nothing in return, and let the commercial conversation happen on the reply if it happens at all. Demote any invitation to the PS.

The catch with this arm is honest: it costs real time per reply. Budget for that before launching it, because failing to deliver on a free offer is worse than never making it.

Round two and round three

Round two: length

There are two large, credible and directly contradictory claims about cold email length. One dataset says 55 to 65 words and that shorter always wins. Another, drawn from more than a thousand campaigns and tens of thousands of positive responses, says 70 to 90 is the band and that both longer and shorter underperform, because under 70 words you lose the credibility line entirely.

Two large claimed datasets, opposite conclusions. That is the cleanest A/B available to anyone doing outbound, and it is exactly why it should not be confounded into round one. Hold every round-one arm at 70 to 90 words so the variable being tested stays readable, then run length on whichever arm wins.

Round three: where the contacts come from

This is the highest-value round and the one almost nobody runs, because it is a data project rather than a copy change. Everyone pulls from the same handful of databases, so the same decision maker receives thirty to fifty near-identical emails a week generated from the same record.

Sourcing contacts from places nobody is scraping changes who receives the email rather than what it says, and that is why it moves the number by more than rounds one and two combined. Public posts from people describing the exact problem you solve. Job postings that signal the gap. Attendee and speaker lists. Communities where your buyer actually talks. It is slower, and it is worth it.

Grading the outbound advice you get sent

Someone forwards you a post with seven tactics and a screenshot. Most of it is real, some of it is wrong, and one item is actively dangerous. A four-bucket grading pass takes twenty minutes and stops you from launching the bad one.

BucketCriterionWhat to do
AdoptAn independent source with different data reached the same conclusionBuild it into the next round as a test arm.
Test laterCredible but contradicts a source you already trustIsolate it in a later round. Never confound a contradiction into round one.
Not a campaignTrue, but it is an infrastructure or procurement decisionRoute it to whoever owns the sending stack. It does not belong in a copy test.
Hard noFilter evasion, guaranteed outcomes, anything you would not want quoted backDo not build it. Write down why, so it does not get proposed again next quarter.

Two examples of the hard-no bucket

Character substitution to slip past filters. Swapping letters for lookalike characters, or seeding deliberate typos the brain autocorrects, on the theory that filters are pattern matchers rather than readers. It inverts. Modern filters normalise those substitutions and score the obfuscation itself as evasion, so you buy the exact outcome you were trying to avoid, on domains you cannot easily replace. It is also deliberate filter evasion by a business with its name on the envelope, and the downside case is not a bad campaign, it is the story about the company that obfuscates characters to get past spam filters.

Outcome guarantees in the copy. Ten qualified meetings in thirty days or your money back, and its many cousins. Unenforceable, a liability in any regulated or capital-adjacent context, and it reads to an experienced buyer as a lottery ticket rather than a service. Take the mechanics from a post like that, leave the posture.

Run a compliance pass on your own copy before it ships

The dangerous sentences are not the obvious ones. They are the ones where a true historical statement quietly turns into a promise, and they are easy to write without noticing.

The rule to check against: never write a sentence where we, or the engine, or the platform is the subject, and investors, customers, meetings or capital is the object of find, reach, book, land, get, source or introduce. You provide the technology and the service. The client owns every relationship.

Breaks it

After listing outcomes other clients reported: that is the actual offer. One sentence converts a permitted historical account into a promised outcome.

Fixed

I cannot promise you any of that. What I can promise is the work, and that you keep it.

Breaks it

That is the part I would want for your company. Softer shape, same problem: it implies the outcome transfers.

Fixed

Name what is built rather than what results. Infrastructure, list, copy, review cadence. Those are real and they are yours.

The rewrite is better copy, not merely safer copy. Saying plainly that you cannot promise the good part is more disarming than implying that you can, and it is the opposite of what everything else in that inbox is doing.

Three more things to check before any send, none of which are copy problems but all of which are yours: no non-ASCII characters in subjects or bodies, no links in a first email, and sender identity plus a postal address plus a working unsubscribe present at the sending layer. Cold business-to-business mail is not a blanket exemption under CAN-SPAM, CASL or GDPR.

The test log

One row per arm, filled in before launch and closed after. If a row cannot be filled in, the test is not ready.

FieldExample entry
ArmA, subject gap
VariableSubject line only. Body is control, unchanged.
HypothesisShort mismatched-case subjects with the first name beat generic title case.
Sending poolThe named clean pool. Same for every arm.
Contacts600, drawn at random from the same segment as the other arms.
WindowTuesday to Thursday, recipient timezone.
Read at500 sends. Not before.
MetricContacts per positive response. Bounce logged separately.
Kill criterionBounce above 5 percent pauses the arm the same day, whatever the replies say.
ResultFilled in after, alongside what you now believe and what you will test next.

Want the tests designed and run for you?

We build the sending layer, the list and the arms, run them Tuesday to Thursday on a clean pool, and read the result with you. Three month engagement, flat fee. You keep everything that gets built.

Built with help from AI. We use AI tools to research, draft, and assemble pages like this one. A human reviews everything, but if something looks off, tell us and we will fix it fast.