A testing programme without a thesis produces wins that do not compound. Each test answers whether one variant beat another and teaches you nothing transferable, so after forty tests you have forty facts and no model of your customer. The programmes that compound are testing a belief, not a button.
This is about the difference between a backlog of ideas and a thesis, why the first plateaus at around a year, and how to run the second without turning it into a research project.
The newsletter
Join our KISS newsletter
One short read a week on what actually moves revenue, in a free email. Read by 10,000+ operators and founders.
No spam. Unsubscribe in one click.
I.Why a programme of wins adds up to nothing
Each isolated test is locally valid and produces no knowledge that survives it.
A.Forty facts and no model
The standard programme: a backlog of ideas, ranked by some effort-and-impact score, run in order. Green button beat blue. Shorter headline beat longer. Three fields beat five.
Every result is real and none of them transfers. Nothing about the button tells you what to do about the pricing page, because you never learned why it won, only that it did.
After a year you have a list of settled arguments and no better idea what your customer is thinking than when you started.
B.The plateau, which is the diagnostic
The pattern is consistent. A new programme produces good results for two or three quarters, then flattens, and the team concludes the site is optimised.
It is not. What has happened is that the obvious tactics are exhausted. Those first wins came from fixing things that were plainly wrong, which any competent person would have spotted, and testing merely confirmed it. Once the plainly wrong is gone, generating the next idea requires a model of why people behave as they do, and a tactical programme never built one.
What the plateau looks like
Metrics viewII.What a thesis is, and what it changes
A falsifiable claim about why your customers behave the way they do, specific enough to be wrong.
A.Three examples, and one non-example
“Our buyers stall because they cannot estimate what they will pay.” Falsifiable, specific, and it generates a dozen tests: a calculator, typical-usage examples, a simplified unit, a starting price on the enterprise tier.
“People who arrive from comparison searches need proof we are different, not proof we are good.” Generates a different dozen, all about the arrival experience for that segment.
“The blocker is a security review nobody tells us about.” Generates tests about surfacing compliance information earlier.
The non-example: “we want to increase trial signups by 20%”. That is a target. It generates nothing, because it contains no claim about why people currently do not sign up.
B.Where a thesis comes from
Not from a workshop. From behaviour you can observe plus conversations you have had.
The behavioural half is where the sequence lives: which people hesitated, what they looked at before they stopped, what the ones who converted did differently. A drop-off rate cannot produce a thesis because it has no shape. A path can.
The conversational half is unavoidable. Behaviour tells you where people hesitate and never tells you what they were thinking, so somewhere in the formation of any thesis there are five conversations with people who did the thing.
The same programme, run both ways
Comparison view| Tactical | Thesis-led | |
|---|---|---|
| Unit of work | One test | A few tests against one belief |
| A loss means | Try the next idea | Evidence against the belief |
| After 12 months | A list of settled arguments | A model that generates tests |
| Failure mode | Plateau | Attachment to a wrong thesis |
| Who can run it | Anyone with a backlog | Someone who talks to customers |
III.Running it, including the tests you should lose
A few tests per thesis, then judge the thesis. And expect to be wrong about a third of the time, or you are not testing beliefs.
A.Test the belief, not the variant
Take the thesis and design three or four tests that would each be expected to win if the thesis holds. Run them. Then make a decision about the thesis, not about the variants.
If three of four moved the metric, the belief is probably right and there is a whole direction to exploit. If none did, the belief is wrong and the useful output is eliminating it, which is worth more than a green button because it redirects everything that comes after.
B.Why losses only pay under a thesis
In a tactical programme a losing test is a dead end: the idea did not work, cross it off, take the next one. Nothing was learned that helps with the next idea.
Under a thesis, the same loss is evidence. Four tests all aimed at making the price estimable, all flat, tells you the stall is not about estimation, and that eliminates a whole region of the backlog.
This is also the honest defence of testing at a company where most tests lose, which is most companies. A programme where everything wins is testing things that were already known.
C.The two ways this goes wrong
Attachment. A team that has built a thesis will defend it, reinterpreting flat results as poor execution. Write down in advance what result would make you abandon it, and hold to that.
Underpowering. Testing a belief across several tests multiplies your sample size requirement, and running four underpowered tests to judge one thesis is worse than running one properly. See sample size and multiple comparisons.
Kissmetrics records the variant as a property on the person, so every other report splits by variant without extra setup, and a test can be judged on retention and revenue months later rather than on the click. That matters more under a thesis than under a tactical programme, because a belief about your customer should predict what they do next quarter, not what they clicked on Tuesday. Related: calling a winner.
Verdict
A programme of tactics produces real wins for about a year and then plateaus, because the early gains were repairs rather than discoveries and nothing built a model that could generate the next idea.
Write a falsifiable claim about why your customers behave as they do, run three or four tests against it, then judge the claim rather than the variants. Decide in advance what would make you abandon it. The losses are the point: eliminating a wrong belief redirects everything downstream of it, which no green button ever does.
Continue Reading
How to Find the Right Ideas for A/B Testing (Stop Guessing)
The biggest waste in A/B testing is testing the wrong things. This guide shows you how to use analytics, user feedback, and competitive analysis to identify tests that move the needle.
Read articleA/B Test Sample Size: How to Calculate It and Why Most Teams Get It Wrong
Most A/B tests are called too early with too little data. This guide shows you how to calculate the right sample size before you start and why peeking at results destroys statistical validity.
Read articleIntroduction to A/B Testing: How to Run Experiments That Actually Work
A/B testing is the most reliable way to improve conversion rates. But most tests fail because of poor methodology, not poor ideas. This guide shows you how to run tests that produce trustworthy results.
Read article