The result you get in week one isn't the result

By Mateo Benitez

Two well-documented psychological effects distort early test data in opposite directions, and most teams only know about one of them.

Two contrasting beams of light crossing a minimal interior, representing familiarity and novelty pulling early test results in opposite directions

Two ways a test can mislead you

A genuinely better redesign can lose an early test to a worse, familiar one. A genuinely worse redesign can win an early test against a better, familiar one. Both are real, well-documented findings, and most teams that ship based on week-one data have only ever heard of one of them.

The familiarity effect

People rate things they've simply seen before more favorably than things they haven't, independent of actual quality, a finding replicated enough times across enough contexts that it has a name: the mere-exposure effect. An interface someone has used for two years carries an advantage no A/B test accounts for by default. Test a genuinely clearer version against it in week one, and some fraction of the users choosing the old one aren't choosing it because it's better. They're choosing it because it's the one they already know how to use without thinking, and that comfort reads, in a survey, indistinguishable from preference.

The novelty effect

The opposite bias shows up just as often and gets talked about far less. A new pattern, badge, or interaction draws attention purely because it's new, and attention alone can lift engagement numbers for a while, independent of whether the change actually serves anyone. A redesign that spikes in week one isn't necessarily good. It might just be interesting, in the specific way anything unfamiliar is interesting for exactly as long as it stays unfamiliar.

Why this matters for the timeline

Run a test for a week and there's no way to know which effect, if either, is doing the work. A win might be real improvement, or it might be curiosity that decays to baseline by week four. A loss might be real regression, or it might be resistance that fades once the new pattern becomes the familiar one. The only way to tell the difference is time, which is exactly the variable most launch timelines are least willing to give a test.

The result in week one isn't wrong, exactly. It's just measuring familiarity and novelty at least as much as it's measuring quality, and neither of those wears off on a schedule that fits inside a sprint.

Follow me to keep in touch

Where I share my creative journey, design experiments, and industry thoughts.

Create a free website with Framer, the website builder loved by startups, designers and agencies.