Statistics · 27 April 2026
The novelty effect: why your winning test faded after four weeks
A variant that wins in week one and evaporates by week six probably didn't stop working. It never worked — it was just new.
It's one of the most common post-mortems in CRO: a test wins convincingly, the change ships, and a couple of months later someone notices the metric has drifted back to exactly where it started. Cue the awkward meeting. Did the market change? Did another release cancel it out? Was the test wrong?
Often, the answer is simpler: your returning visitors noticed something new, interacted with it because it was new, and then stopped. The uplift was real, briefly — it just wasn't an uplift in value. It was an uplift in novelty.
Who novelty affects, and who it can't
The tell is in the segments. New visitors have never seen your old design; to them, nothing is novel. If a variant genuinely communicates or converts better, it should work on people with no memory of the control.
Returning visitors, meanwhile, notice change itself. A moved button gets clicked because it moved. A redesigned banner gets attention because it's different from yesterday. That attention decays on a timescale of days to weeks as the change becomes the new normal.
So the diagnostic is straightforward: split your result by new vs returning visitors. A variant that wins +18% with returning visitors and +1% with new ones isn't an 18% winner. It's a 1% winner wearing a costume — and on sites with loyal, frequent audiences (SaaS dashboards, student accommodation portals mid-booking-season, anything with an account), the costume can be most of the number.
The mirror image: primacy effects
Novelty has an evil twin. Sometimes a genuinely better variant loses early because regular users have muscle memory for the old layout and stumble on the new one. Their friction is temporary — it fades as they relearn — but if your test only runs a fortnight, you'll record a loss for a change that would have won at steady state. Big navigation and layout changes are especially prone to this.
Both effects point the same direction: early data over-represents reaction to change, and under-represents steady-state behaviour, which is the thing you actually care about.
Practical defences
Run tests in full weeks, and long enough to see decay. Two weeks minimum; for audiences with heavy return rates, four. Then look at the trend of the uplift over time, not just the total. A real winner's effect is roughly stable; a novelty winner's effect visibly shrinks week on week.
Always read new visitors separately. Not as post-hoc fishing — pre-register it. "Primary metric overall; novelty check on the new-visitor segment" should be in the test plan before launch.
Be most suspicious of your flashiest tests. Animation, colour, movement, anything that shouts "I'm different!" is precisely the category where novelty inflates results. A copy change on a form field rarely triggers novelty; a redesigned hero often does.
Re-measure after rollout. For big wins, hold back a small long-running control (a 95/5 holdout) for a month or two. It's cheap insurance, and it either confirms the win or catches the fade before the annual report does.
None of this makes early wins worthless. It makes them provisional — which is what every test result is anyway. The teams that get burned are the ones who confuse "significant in week one" with "true forever".