The Key Insight
"We tried it, it did not work" is a conclusion, and a conclusion is only as good as the test that produced it. Run a dead test through five diagnostic questions: if any finds a fault, the channel verdict is unsafe and a re-test can be justified; if none do, the kill stands and you can stop wondering.
Somewhere in your reporting there is a dead channel with a one-line epitaph: "we tried it, it did not work." That sentence closes the conversation, usually for a year or more, and nobody enjoys reopening it.
But the conclusion is only as good as the test that produced it, and in many failed expansions the test design is the more likely culprit than the channel. A dead test can usually be diagnosed with five questions. If any of them finds a fault, the channel verdict is unsafe. If none do, the kill stands, and you can stop wondering.
The five questions map to the parts of a proper test design, the one laid out in our guide to testing a new paid media channel before scaling. Run your dead test through them in order.
Diagnosis 1: Was It Judged Against the Wrong Benchmark?
A common fault. The new channel was compared to the incumbent channel's mature CPA, often within the first weeks. The incumbent had years of optimisation, warm audiences, retargeting pools, and accumulated creative learnings; the new channel had none of them. On that comparison, many channels look weak before they have produced a fair answer.
The tell: nobody can say what the pass mark was, only that the numbers "compared badly" to the main channel. And if the comparison was made on platform-reported ROAS, the problem gets worse, because attributed ROAS can understate a cold channel while a mature retargeting-heavy channel flatters itself, for the reasons covered in platform ROAS vs incremental ROAS.
The fix for next time: a written pass mark in unit economics, agreed before launch, judged on its own terms.
Diagnosis 2: Was There a Real Hypothesis?
Could anyone on the team have stated, before launch, what success looked like in numbers? Which segment the channel was meant to reach, at what cost per new customer, within what window, on what budget?
If not, the test had no clean way to fail or succeed. It could only end. A test without a hypothesis produces a feeling, not a finding, and feelings usually default to the status quo.
The fix: the bracket format from the pillar: Channel X can reach [segment] at no worse than [pass mark] within [duration] on [cap], measured by [blended method]. Every bracket filled before a pound moves.
Diagnosis 3: Could the Tracking Read the Result?
Many tests are measured mainly inside the platform being tested, which means the channel is too close to its own scorecard. Cold attribution windows, missing conversion imports, and no blended view can make a working channel look dead, and occasionally a dead one look alive.
The questions to ask: was blended new-customer data available during the test? Were platform claims reconciled against real revenue? Where volume allowed, was there a controlled read, of the kind described in what is a geo holdout test? If the answer to all three is no, the test produced platform reporting, not evidence, and the incrementality ladder shows what evidence would have looked like.
The fix: tracking readiness before launch, and a measurement method named in the hypothesis.
Diagnosis 4: Were the Audiences Polluted?
Check what the test campaigns actually served on. If existing customers, active retargeting pools, or brand searchers leaked in, the results are harder to trust in either direction: harvest can make a weak channel look passable, or make a genuinely promising channel look unnecessary because its conversions were claimed by audiences you already owned.
The tell in the data: test-period conversions concentrated in returning visitors or existing customer segments rather than new-to-business buyers.
The fix: customer and retargeting exclusions on prospecting tests, so the channel has a clear job and the result describes it.
Diagnosis 5: Was the Sample Too Small to Answer?
An invented illustration: a test spends £900 over three weeks against an £80 target cost per new customer. Even at target performance, that is around eleven customers, a number that normal variance can produce or erase on its own. The test can look scientific while giving a noisy answer, the same problem that undersized geo tests can have.
Underpowered tests are common because caution feels safe. But the honest conclusion of an underpowered test is "unknown", not "failed". Unknown should not close the conversation for a year.
The fix: budget derived from the evidence required, as the pillar's sizing section shows: enough expected conversions to read against the pass mark, capped and time-boxed.
What to Do With the Diagnosis
If one or more of the five found a fault, the channel verdict is unsafe. The learning from the dead test is still useful, creative that flopped, audiences that engaged, landing paths that leaked, but the conclusion "this channel does not work for us" is not supported. A re-test can be justified, with the pillar's full design and one named fix per re-test, under the same kill, iterate, or scale discipline. What a diagnosis does not justify is an enthusiastic relaunch with the same design; that risks buying the same unreadable answer twice.
If none of the five apply, the test was sound and the kill stands. That is not a disappointment; that is the method working. Write the result down, revisit only when something material changes, and spend the energy on the next hypothesis. Before any re-test, clean the base: audit the waste in the account that will fund it.
Running this post-mortem, and designing the re-test where one is justified, is part of how we run new channel testing engagements: it is usually cheaper to diagnose the failed test than to abandon the channel without knowing what went wrong.