Skip to main content
OmerH
Partner - Gold
2026 Champion
August 18, 2026

What I split test before BFCM instead of subject lines

  • August 18, 2026
  • 1 reply
  • 68 views
Omer Hazer, Atlas Studios

Here are the pre-BFCM experiments I run (and why), how to run them as holdouts, the metrics that decide whether a cohort goes live on the big send, and how I leverage Klaviyo features and Composer.

Most pre-BFCM advice covers subject lines, send times, and offer copy. None of that tells you what to actually do on Black Friday.

Not because testing the above doesn't matter, but because leading into a compressed window like Black Friday Cyber Monday (BFCM), every test costs a send you can't get back, and those 3 don't change what you do on Black Friday.

So I apply one filter to everything: does this test inform the actions I need to take in BFCM?

Why is deliverability the real BFCM testing filter?

Optimize a subject line, and you improve one email.

Improve inbox placement, and you improve every email you send for the rest of the year, including the key communication that happens during BFCM.

Run the arithmetic on your own account. Move placement from 85% to 92% on a 500,000-person send, and that's 35,000 more people who actually see the email, at the moment your revenue per recipient peaks. No creative work. No extra discount. Nothing else you can test produces a multiplier like that.

And deliverability is the one thing you cannot fix in November. Sender reputation is built on months of engagement history. Arrive at peak week with a bloated list and poor deliverability and the damage is already done. What's worse is that I've seen merchants cause most of this damage during their testing.

The question stops being "what could I test?" and becomes: what test changes how much of my list I can safely reach on the biggest send of the year?

Test 1: how deep into the list can you safely send?

On Black Friday you'll want to reach more people than usual.

Most people answer this by expanding sequentially. Send to 30-day engaged, then add 30–60 next time, then 60–90, and eyeball the results. Don't. That confounds the cohort with the week you sent it, so a segment that looks weak may have just caught a bad send.

Run it as a holdout instead. Take the cohort you're considering, say 60–90 days engaged, split it randomly, send to one arm, suppress the other. Then measure the delta between arms on incremental revenue per thousand and on complaint and bounce rate. Same week, same campaign, same conditions. The only variable is whether they got the email.

There's usually a cohort where incremental revenue goes near-flat but complaint rate doesn't. That's your ceiling, and now you know it from a controlled read rather than a gut call at 11 p.m. the night before.

Test 2: will your newest subscribers still be sendable in November?

Your welcome series is a deliverability test. Almost nobody treats it as one.

Every subscriber you acquire in September and October has to arrive at Black Friday warm. If they don't, they're not just neutral. They're cold, sitting in your list, suppressing engagement rates on the send where placement matters most.

Acquisition volume also peaks in this window, so the flow is processing more profiles now than at any point in the year, which gives you the volume for a clean test.

Two variants worth splitting:

  • Sequence length. 3 messages versus 5. More touches means more engagement signal, but only to the point where unsubscribe and complaint rates inflect. Find that in October, not on November 28.
  • First-purchase versus first-engagement framing. A welcome series optimized purely for immediate conversion often produces a worse-behaved list than one optimized for opens and clicks. For BFCM, an engaged non-buyer is an asset. A subscriber who ignored 5 emails is a liability.

Klaviyo supports A/B testing individual flow messages, and you can split flow branches with a conditional split on a random sample for sequence-level tests. Test one variable at a time, whether that's your subject line in the welcome flows, opening copy, or send time.

Test 3: average order value (AOV) mechanics, and what you really can't test beforehand

Be honest about the constraint first, because most articles skip it. You cannot test BFCM discount depth before BFCM. Running 30% off in October to see whether it beats 40% teaches your list the sale is already here and pulls demand forward. The test corrupts the environment it's testing.

So I test the mechanic, using promotions that don't cut headline price: free shipping thresholds, spend-and-save tiers, bundle pricing, and gift with purchase.

All 4 run cleanly in September and October, none train the list to wait for a discount, and all 4 are AOV levers, which is what you want to understand before a period where high margin basket size does more for you than conversion rate.

Judge them on AOV and revenue per recipient, never conversion rate alone. A flat incentive usually beats a threshold mechanic on conversion and loses on contribution margin, because you've given something away on every order including the ones that were coming anyway. A test that wins on conversion and loses on margin is a loss. It just doesn't look like one on the dashboard.

What transfers isn't the winning offer. It's whether this audience responds to threshold mechanics at all, and where the threshold sits relative to current AOV. Both hold when you layer a discount on top in November.

What I don't test

Subject lines. Contaminated metric, context-specific result, and the winner doesn't transfer to a week where every brand your customer has ever bought from is shouting simultaneously. I'll still A/B test these during BFCM, but it doesn't produce directional guidance pre-BFCM.

Send time. Send time optimization assumes the inbox environment on your test day resembles the one you're optimizing for. On Black Friday it doesn't, remotely. Your carefully validated 10:14 a.m. is landing somewhere completely different.

The result that changed my approach

A few years ago, a high eight-figure health and wellness brand. Record BFCM, their strongest year-on-year revenue yet. We sent an ungodly number of emails and tested subject lines and content relentlessly. We went into that period with placement in the 60s. We came out of it in the 40s. Email revenue the following Q1, Q2, and Q3 was flat, and below prior year in some months. We'd borrowed a record November from the 9 months after it, and the interest rate was brutal.

The next year we did it differently. Non-discounted mechanics tested with holdout groups, incrementality measured instead of attributed revenue, complaint rate delta read by cohort before committing to a sending depth. Then a much tighter promotional calendar off the back of it. Fewer emails, fewer people per send.

23% more revenue year on year, off an already record BFCM. 18% more margin, with less discounting. And the number I actually care about: the following Q1 to Q3 up 69%, holding 90+ placement throughout.

You live and die by two things. The offer, and deliverability.

How does list size change what you should test?

Smaller lists can't always reach significance fast, and pretending otherwise is how people act on noise. 3 adjustments:

Test the thing with the biggest effect size. Subject lines move open rate a few percent. AOV mechanics move basket size by double digits. If you can only detect large effects, only test things capable of producing one.

Pool across sends, not within one. A single send won't clear significance on a small list. The same directional result across 4 consecutive sends is real signal.

Separate the metric you test on from the metric you decide on. Klaviyo can call a winner on open or click rate and roll it out automatically. Useful for mechanics. But the decision gets made on revenue per recipient and complaint rate, read manually, once order data settles. Automatic rollout on a proxy metric is how you confidently ship the wrong variant.

How I have actually gotten value out of Composer for testing

The bottleneck there has always been the manual pull: cohort by cohort, send by send, in a spreadsheet, usually on a Sunday. This is where I've been getting real use out of Composer. Its audit side analyzes your existing campaigns, flows, and segments, so instead of assembling the cohort-level picture by hand, you can ask for it and interrogate what comes back. When you're pooling results across campaign sends under time pressure, cutting the gathering from hours to minutes decides whether you act on the data at all.

What carries forward and what doesn't

One test: does this tell me about my audience, or about that specific week?

Carries forward: engagement ceilings, and AOV mechanics. Both describe how your audience relates to your brand, and that doesn't shift between October and November.

Doesn't: anything measured mainly on opens or clicks. BFCM inbox context is unlike any other week of the year.

A test I got wrong

Early in my retention career, I tested send times properly. Months of it, across campaign types, on a brand with the volume for clean reads. Early evenings won almost every time. About as unambiguous as A/B testing gets. So we rolled it into BFCM, and landed in inboxes that had already taken 10 Black Friday emails that day.

The data wasn't wrong. It was answering a question about a normal Tuesday, and I applied it to the one week the inbox works differently. During BFCM, you're competing for share of wallet harder than at any other point in the year, and it is genuinely early bird gets the worm.

Some things you test and act on. Some things you just do. Being early is one you just do.

How I use this in my own workflow

When I start: 8 to 10 weeks out. The holdout tests need 4 or more sends for a poolable read, and flow tests need time for people to get through the flow. Any later and you're guessing with extra steps.

What makes me change a plan: complaint rate delta, above everything. Meaningful incremental revenue and complaint rate holds, the cohort's in. Marginal revenue and complaint rate moves, it's out, and I don't agonize, because the cost of that one shows up in March, not November. Revenue results I note. Deliverability results change the calendar.

One test only: the cohort holdout. The only one where getting it wrong damages every send after it.

Repeating this year: cohort holdouts, deliverability metrics (unsubscribe, bounce, spam complaint, and engagement), and non-discounted mechanics. The holdout is the only thing that tells you whether you generated revenue or just took credit for it.

Skipping: send time. Permanently, for BFCM. I'll optimize that in March.

If a test won't tell you what to do differently during BFCM, don't run it. That's the whole filter.

I spent years reading BFCM testing advice that was all about subject lines or copy. The stuff that actually cost me money was never in it.

Share your experience

What's the one test you always run before BFCM, no matter how tight the timeline gets? And, has anyone actually run a holdout during peak week? It took a bad year for me to start.

If you've done it and the number surprised you, I want to know.


Founder & Chief Executive, Atlas Studios | Shopify Plus & Klaviyo Master Gold Partner

Growth strategy, retention marketing, and smart Shopify development for mid-market brands.

www.atlasstudios.agency

1 reply

Contributor I
August 19, 2026

nice post