The skills library is open: 86 files, one email

Incrementality Testing for B2B: Geo Holdouts, Audience Holdouts and the Power Problem

How to run a holdout that shows whether a channel caused pipeline, which design fits a B2B business, and how to tell when you are too small to test at all.

Mert · Founder7 min read
Post Share

Switch LinkedIn off in Austria for ten weeks and count what happens to pipeline. That sentence makes marketing leaders physically uncomfortable, which is why it almost never happens, and it is still the only measurement in the building that owes nothing to a model.

Closed-loop attribution mentions holdouts in a paragraph and moves on. This is the method, including the part where you find your company is too small to use it.

Incrementality asks what would have happened anyway

Attribution asks which touches appeared in the path of closed deals. Incrementality asks a counterfactual: of the pipeline that arrived, how much would have arrived anyway.

The gap is not academic. Branded search converts beautifully, every model credits it, and many of the people clicking it were going to type your name into a browser anyway, because a colleague recommended you last Thursday. Retargeting has the same shape. Both look like your best channels and both can be substantially non-incremental. An experiment settles it because you construct the counterfactual instead of inferring it, so nothing depends on a credit rule or a click identifier surviving a CRM conversion.

Three designs, and what each needs to be valid

Geo holdouts split the addressable market by geography. They need cells that behaved alike beforehand, established with a pre-period of equal length and checked rather than assumed, and they need separation, because a German-language campaign in Bavaria reaches Austrian buyers. Treating DACH as three clean cells is usually wrong; splitting Germany by Bundesland is cleaner.

Audience holdouts hold back a matched share of target accounts using the platform's suppression lists. They need an account-addressable channel, ruling out broad search and display, and accounts matched on what predicts conversion rather than picked alphabetically. They also need honesty about contamination, because a held-out account still gets your newsletter and your SDR sequence. That makes the result the channel's marginal effect on top of everything else, worth writing down so nobody reads it as the total.

Switchback tests alternate the channel on and off in time blocks across the whole market, so every cell is its own control. They need the effect to appear and disappear quickly relative to the block length, and with a three-month cycle, spend from block one still produces meetings in block three. They suit fast-acting changes read on a fast event: a chat widget, a pricing page variant, a retargeting burst. They are rarely right for a media channel in B2B, and saying so early saves a quarter.

The power problem is where B2B tests die

A holdout is readable when the difference between cells is large relative to the noise, and noise on a count is roughly its square root. A cell producing twenty opportunities carries noise of about four and a half, so a gap of four is not a finding. It is Tuesday.

Run it the other way. Detecting a 20 per cent lift with any confidence needs counts in the high hundreds per cell; detecting a doubling needs a few dozen. Most B2B companies close between forty and a hundred and twenty deals a year. Split that in half, run for a quarter, and each cell holds five to fifteen deals. There is no experiment there. There is a coin. The same thinness rules out mix modelling.

The practical answer is a denser event, stated as an assumption

Move up the funnel until the count is large enough to read: qualified opportunities, then meetings, then demo and pricing submissions, then high-intent sessions. Each step buys statistical power and spends relevance. The assumption belongs in the test document before anything is switched off: the treatment does not change conversion from that event onward. You are asserting that a demo request produced by LinkedIn converts like any other.

That assumption fails in a knowable direction. A channel good at cheap form fills can lift demo requests 40 per cent and produce no pipeline at all. So pair the test with a follow-up: read the dense event, then track the treated cohort's downstream conversion for two more quarters. If it converts materially worse, the lift was volume, not demand.

The window is set by the sales cycle, which is what makes it expensive

Two lags decide duration. The first is the gap between exposure and the measured event. If the median time from first touch to opportunity creation is six weeks, a ten-week test measures only four weeks of treated traffic, so size the window as the lag plus the measurement period you actually want.

The second is carryover in the holdout cell, which keeps converting pipeline generated by spend that ran the month before. Those opening weeks are contaminated by your own history, and including them understates the effect. Discard them, and fix the length in advance.

A realistic B2B geo holdout is therefore a quarter of committed calendar, planned around a stretch with no launch or price change.

Size the cost instead of calling it burning money

The reason nobody runs the test is the feeling of leaving money on the table. That feeling is a number, so calculate it. If the holdout is a fifth of the market and the channel costs 40k a month, you withhold roughly 24k across the quarter, and the pipeline forgone is that spend times your believed return. That is an upper bound, since it assumes the channel is fully incremental, which is the claim under test. Against it: if the channel is 40 per cent non-incremental and you never find out, you spend 480k a year on a belief. Put both numbers in one sentence to the CFO and the conversation changes.

Reading the result without fooling yourself

Fix the cells, the metric, the window and the threshold in writing first. Criteria set afterwards are not a test, they are a search. And do not peek and stop: checking weekly and halting when the gap looks good guarantees a false positive eventually.

Report an interval, not a point. "No significant difference, interval from minus 15 per cent to plus 35 per cent" is the truth; the first half alone is not. A noisy null is not a proven zero, and treating it as one is how a working channel gets defunded by a badly sized experiment.

And give the result a shelf life: response is concave, so the answer at 80k a month is not the answer at 40k.

When your company is too small to run one, say so plainly

If the densest event you can honestly connect to the channel yields fewer than about thirty occurrences per cell in a quarter, no design rescues it. Do not run the test. You will get a number, it will be noise, and someone will present it. The same holds for a company with one meaningful channel, where the only holdout is switching off all its pipeline.

At that scale the honest instruments are cheaper: a self-reported attribution field on the demo form, a clean signal ledger, and an audit of whether the data flows. Escalate to holdouts when one channel's annual spend crosses the point where being wrong costs six figures. That is the threshold, not headcount and not revenue.

Frequently asked questions

What is incrementality testing?

Incrementality testing measures how much of an outcome a channel actually caused, by comparing an exposed group against a comparable unexposed one. Unlike attribution, which credits touches preceding deals, it builds a real counterfactual, so the result depends on no credit rule and no identifier surviving the journey. It separates demand a channel created from demand it intercepted.

How long should a B2B holdout test run?

Long enough to cover the lag between exposure and the measured event, the measurement period itself, and a burn-in you discard because the holdout keeps converting pipeline generated before the switch-off. With a six-week lag to opportunity creation: an eight-week pre-period, three to four discarded weeks and eight to twelve readable ones. A quarter of committed calendar.

What if there are not enough conversions to read the result?

Move the measurement to a denser event higher in the funnel, such as demo requests or qualified meetings, and state the assumption: that the treatment does not change the conversion rate from that event onward. Verify it later by tracking the treated cohort for two more quarters. If the densest honest event still yields under thirty per cell in a quarter, do not run the test.

Does a null result mean the channel does nothing?

No, and this is the most common misreading. A test finding no significant difference usually means it lacked power, not that the effect is zero. Report the interval: one from minus 15 per cent to plus 35 per cent holds both a harmful channel and an excellent one, so the honest conclusion is that you learned nothing. Only a tight interval near zero is evidence of absence.

see where you stand

Twelve questions. Then your build order.

The diagnostic returns your operating stage, the three widest gaps in your motion and what to build first. Two minutes, no sales sequence, one human reply.

Keep reading

All articles →