Home · Method
Measuring a store change without fooling yourself
Retail data is generous. Give it any hypothesis and enough slices, and it will hand you a confirming number — some store, some week, some category where the new thing "clearly worked". This is not because stores are mysterious. It is because store data is noisy, seasonal and abundant, which is the exact recipe for self-deception. The measurement craft is mostly a list of ways people fool themselves, plus the cheap habit that defeats each one.
Seasonality eats naive comparisons
The most common false positive in retail is "after versus before". Install in March, compare April to February, declare victory — having actually measured spring. Same-period-last-year helps but drifts with the calendar (Easter moves, weather doesn't repeat, last year had a stockout). The defence isn't cleverer arithmetic; it's concurrent controls. Matched stores measured over the same calendar absorb the season, the weather, the macro mood — everything except your change. Before/after without controls is a press release, not a measurement.
The novelty bump flatters everything
New fixtures get attention from shoppers and — underrated — from staff, who straighten, restock and evangelise the new thing while it's shiny. Both effects fade. A change that looks brilliant in week two and ordinary in week six didn't "stop working"; it never worked, beyond being new. Run past the bump, and weight the late weeks over the early ones when you judge. If the effect needs to be permanent to pay back, only the boring part of the curve counts.
Window-shopping for results
Decide the measurement window when you design the test, because every retail time series contains a flattering interval if you go looking afterwards. The same applies to metrics (basket, mix, traffic, category units — pick the primary up front) and to store subsets ("it worked in the urban formats" is a new hypothesis to test, not a salvage of the old one). None of this requires statistics beyond honesty; the pre-commitment does the heavy lifting.
Spillovers and sabotage, the quiet confounds
Two field effects that desk plans miss. Spillover: if test and control stores share a catchment, your change can siphon shoppers from the control — inflating the gap from both sides. Keep matched pairs geographically apart. And contamination of the human kind: district managers love sharing "what's working", which means your control stores may quietly adopt half the treatment by week five. Tell operations which stores must stay frozen, in writing, and check.
Small effects on big bases are still money
The last trap is dismissing a result for being undramatic. Store interventions rarely deliver double-digit lifts; the realistic prize is low single digits on a metric with a very large base. A small, durable mix shift, multiplied across a chain and a year, often pays for the entire programme — but only a measurement designed as above can distinguish that small real effect from small real noise. That, in the end, is what the discipline buys: the ability to take modest numbers seriously.