What is a postcard A/B test?
It is a controlled comparison between two versions of the same mailing. Version A is usually your current design and version B changes one planned element. Everything else stays the same: audience type, quantity, mailing date, card size, paper and measurement rules. Because only one thing differs, a clear difference in results can be linked to that change.
Changing several elements at once still produces a comparison, but only between two complete concepts. You learn which card did better, not why, and you cannot reuse the lesson on the next campaign.
What should you test first on a postcard?
Test the element most likely to change the reader's decision. For most small businesses that is the offer or the headline, followed by the call to action, then the main image. Layout details and color usually move results less and need larger tests to detect.
| Element | Example of a single change | Good first test? |
|---|---|---|
| Offer | 10 percent off versus a free add on | Yes |
| Headline | Benefit headline versus question headline | Yes |
| Call to action | Scan to book versus call to book | Yes |
| Main image | Product photograph versus illustration | Sometimes |
| Card size | 4 x 6 versus 6 x 9 inches | Only with cost tracked |
| Paper or finish | Matte versus glossy | Rarely first |
Write the question before designing, for example: does a free add on produce more bookings than 10 percent off for past customers? If you need help sharpening the offer itself, see the guide to postcard offer copy.
How do you split the audience fairly?
With a mailing list, assign each address to A or B at random, for example by sorting the list randomly and alternating. Random assignment spreads differences in neighborhood, customer history and household type evenly across both groups.
Route based mailing is harder. USPS Every Door Direct Mail delivers to every address on a carrier route, with a minimum of 200 pieces per mailing, so you cannot randomize individual homes. Instead, split by routes that are as similar as possible in size and household profile, mail both on the same day, and write the limitation into your report. Direct mail postcards bundled for EDDM make this practical for local businesses.
How many postcards do you need for each version?
More than most people expect. Response to mail is often a small percentage, so small differences are easily produced by chance. The NIST Engineering Statistics Handbook describes two standard ways to compare two proportions: a normal approximation for reasonably large samples and the Fisher exact test when samples are small.
Worked example with illustrative numbers: 500 cards per version, 10 responses for A (2 percent) and 15 for B (3 percent). A two proportion test gives a z value of about 1.0, well short of the 1.96 usually used for 95 percent confidence. B might be better, or the gap might be ordinary variation. To detect a one point difference at these response levels you would need several thousand cards per version.
If the business decision is small, a modest test can still guide you, but report it as a direction, not a proven winner.
How do you track each version separately?
Give each version its own tracking route. Use a separate tagged link and short address for A and B, distinguished by utm_content, and a separate offer code if people can respond by phone or in person. Our guide to postcard QR tracking covers the setup. Decide the outcome that counts, such as completed bookings, and the observation window, such as four weeks from the in home date, before any card goes out.
Do not stop the test the moment one version pulls ahead. Early leads often shrink as more responses arrive.
How should you read the results?
Report raw counts, the number mailed, response rates and costs for both versions. The insert experiment planner calculates observed response rates and the difference between versions in percentage points, which suits postcards as well as package inserts. If versions have different costs, as with a larger card, compare cost per outcome, and use the campaign break even calculator to check whether the better version also pays for itself.
Keep the result attached to its context. A test that finds an illustration beat a photograph for one audience and one offer tells you about that campaign, not about all photographs.
What should your test plan include?
- One written test question and one changed element.
- Random or closely matched groups of equal size.
- The same mailing date, size, paper and quantity for both versions.
- A separate tracking link and offer code for each version.
- A primary outcome and observation window set before launch.
- A planned sample size, or an honest note that the test is directional.
- A report with raw counts, costs and limitations.
What do you do after the test?
If one version clearly wins, make it the new control and test the next element against it. If the result is unclear, either repeat the test with a larger mailing or keep the cheaper version and move on. Keep a simple log of every test, its question, sample size and outcome, so later campaigns build on evidence instead of starting over.
Frequently asked questions
How many postcards do I need for an A/B test?
It depends on your expected response rate and the size of difference you want to detect. When responses run at a few percent, detecting a one point difference usually takes several thousand cards per version. With 500 per version, a result of 10 against 15 responses is not statistically reliable, so treat small tests as directional rather than proof.
Can I A/B test an Every Door Direct Mail campaign?
Yes, but not by randomizing individual homes, because EDDM delivers to every address on a carrier route. Split the mailing by routes that are similar in size and household profile, mail both versions on the same day and track each with its own link and code. Record that the groups were matched, not randomized.
What should I test first on a postcard?
Start with the element most likely to change the reader's decision, usually the offer or the headline, then the call to action. Image, color and layout details tend to produce smaller differences that need larger tests. Write one clear question before designing, and change only that element between the two versions.
When can I call a winner in a postcard test?
Only after the observation window you set before launch has ended, and only if the difference is large enough relative to the sample size. The NIST handbook describes standard two proportion tests for this comparison. Stopping as soon as one version leads is a common way to declare a winner that later disappears.
Sources & specification notes
The references below support the relevant technical or product details in this guide. Examples and checklists are not claims of practical testing or universal supplier requirements. Confirm the current specifications for your chosen product before production.
- NIST/SEMATECH e-Handbook of Statistical Methods: comparing two proportions ↗
Describes the normal approximation for large samples and the Fisher exact test for small samples when comparing two proportions.
- USPS: Every Door Direct Mail ↗
Route based delivery with a minimum of 200 pieces per mailing, which is why EDDM tests split by route.
- Google Analytics Help: URL builders and UTM parameters ↗
Lists utm_content for differentiating creative, useful for tracking A and B separately.




