Why Your A/B Testing Programme Plateaus After a Year

The early gains were repairs, not discoveries. Write a falsifiable claim about why customers behave as they do, run three or four tests against it, then judge the claim rather than the variants.

KISSmetrics Editorial

|11 min read

A testing programme without a thesis produces wins that do not compound. Each test answers whether one variant beat another and teaches you nothing transferable, so after forty tests you have forty facts and no model of your customer. The programmes that compound are testing a belief, not a button.

This is about the difference between a backlog of ideas and a thesis, why the first plateaus at around a year, and how to run the second without turning it into a research project.

The newsletter

Join our KISS newsletter

One short read a week on what actually moves revenue, in a free email. Read by 10,000+ operators and founders.

No spam. Unsubscribe in one click.

I.Why a programme of wins adds up to nothing

Each isolated test is locally valid and produces no knowledge that survives it.

A.Forty facts and no model

The standard programme: a backlog of ideas, ranked by some effort-and-impact score, run in order. Green button beat blue. Shorter headline beat longer. Three fields beat five.

Every result is real and none of them transfers. Nothing about the button tells you what to do about the pricing page, because you never learned why it won, only that it did.

After a year you have a list of settled arguments and no better idea what your customer is thinking than when you started.

B.The plateau, which is the diagnostic

The pattern is consistent. A new programme produces good results for two or three quarters, then flattens, and the team concludes the site is optimised.

It is not. What has happened is that the obvious tactics are exhausted. Those first wins came from fixing things that were plainly wrong, which any competent person would have spotted, and testing merely confirmed it. Once the plainly wrong is gone, generating the next idea requires a model of why people behave as they do, and a tactical programme never built one.

What the plateau looks like

Metrics view
Q1: obvious fixes
8.1% lift per quarter
Q2: remaining obvious fixes
5.4% lift per quarter
Q3: running out
1.9% lift per quarter
Q4: noise
0.4% lift per quarter
Q5: noise
0.7% lift per quarter
Illustrative, not measured. The shape recurs because the early wins were never really tests: they were repairs, confirmed by a test.

II.What a thesis is, and what it changes

A falsifiable claim about why your customers behave the way they do, specific enough to be wrong.

A.Three examples, and one non-example

“Our buyers stall because they cannot estimate what they will pay.” Falsifiable, specific, and it generates a dozen tests: a calculator, typical-usage examples, a simplified unit, a starting price on the enterprise tier.

“People who arrive from comparison searches need proof we are different, not proof we are good.” Generates a different dozen, all about the arrival experience for that segment.

“The blocker is a security review nobody tells us about.” Generates tests about surfacing compliance information earlier.

The non-example: “we want to increase trial signups by 20%”. That is a target. It generates nothing, because it contains no claim about why people currently do not sign up.

B.Where a thesis comes from

Not from a workshop. From behaviour you can observe plus conversations you have had.

The behavioural half is where the sequence lives: which people hesitated, what they looked at before they stopped, what the ones who converted did differently. A drop-off rate cannot produce a thesis because it has no shape. A path can.

The conversational half is unavoidable. Behaviour tells you where people hesitate and never tells you what they were thinking, so somewhere in the formation of any thesis there are five conversations with people who did the thing.

The same programme, run both ways

Comparison view
TacticalThesis-led
Unit of workOne testA few tests against one belief
A loss meansTry the next ideaEvidence against the belief
After 12 monthsA list of settled argumentsA model that generates tests
Failure modePlateauAttachment to a wrong thesis
Who can run itAnyone with a backlogSomeone who talks to customers
The thesis-led failure mode is real and different: a team can defend a belief past the point the evidence supports it, which a tactical programme cannot do because it believes nothing.

III.Running it, including the tests you should lose

A few tests per thesis, then judge the thesis. And expect to be wrong about a third of the time, or you are not testing beliefs.

A.Test the belief, not the variant

Take the thesis and design three or four tests that would each be expected to win if the thesis holds. Run them. Then make a decision about the thesis, not about the variants.

If three of four moved the metric, the belief is probably right and there is a whole direction to exploit. If none did, the belief is wrong and the useful output is eliminating it, which is worth more than a green button because it redirects everything that comes after.

B.Why losses only pay under a thesis

In a tactical programme a losing test is a dead end: the idea did not work, cross it off, take the next one. Nothing was learned that helps with the next idea.

Under a thesis, the same loss is evidence. Four tests all aimed at making the price estimable, all flat, tells you the stall is not about estimation, and that eliminates a whole region of the backlog.

This is also the honest defence of testing at a company where most tests lose, which is most companies. A programme where everything wins is testing things that were already known.

C.The two ways this goes wrong

Attachment. A team that has built a thesis will defend it, reinterpreting flat results as poor execution. Write down in advance what result would make you abandon it, and hold to that.

Underpowering. Testing a belief across several tests multiplies your sample size requirement, and running four underpowered tests to judge one thesis is worse than running one properly. See sample size and multiple comparisons.

Kissmetrics records the variant as a property on the person, so every other report splits by variant without extra setup, and a test can be judged on retention and revenue months later rather than on the click. That matters more under a thesis than under a tactical programme, because a belief about your customer should predict what they do next quarter, not what they clicked on Tuesday. Related: calling a winner.

Verdict

A programme of tactics produces real wins for about a year and then plateaus, because the early gains were repairs rather than discoveries and nothing built a model that could generate the next idea.

Write a falsifiable claim about why your customers behave as they do, run three or four tests against it, then judge the claim rather than the variants. Decide in advance what would make you abandon it. The losses are the point: eliminating a wrong belief redirects everything downstream of it, which no green button ever does.

Continue Reading

ab testing strategyexperimentation programmeconversion optimization strategytest hypothesiscro strategy
KISSmetrics

Build your business intelligence layer for free.

KISSmetrics records the variant as a property on the person, so every other report splits by it and a belief about your customers can be checked against retention and revenue rather than the click.