Home · Method

How to run an in-store pilot that proves something

Most retail pilots are not experiments. They are previews — a soft launch in a friendly store, watched optimistically, reported selectively. The store got the new thing, sales did whatever sales do, and a deck was written. Nobody can say what would have happened without the change, because nobody arranged to find out. Then the rollout happens anyway, since the pilot's real function was ceremonial.

An actual experiment is a different animal, and the difference is mostly decided before day one. Here is the skeleton we keep coming back to.

Shrink the question until it fits in one sentence

"Test the new store concept" is a renovation, not an experiment — twenty variables changed at once, attribution impossible by design. A pilot earns its name when one deliberate change stands alone: this fixture, this layout move, this screen content, this staffing pattern. If stakeholders want to test five ideas, that is five pilots or one honest admission that you're doing a redesign and will never know which part worked. Both are legitimate. The blend is not.

Controls are the experiment

The test stores tell you what happened; the control stores tell you what it means. Without matched stores left alone, every result has a dozen alternative explanations — weather, a competitor's closure, a TikTok trend, payday. Match controls on what drives their numbers: format, trading volume, catchment type, promo exposure. And resist the volunteer problem: stores that ask to host pilots are systematically better-run than the chain average, which inflates every result they touch. Choose on similarity; persuade afterwards.

Where the change itself can't be randomised across stores cheaply — heavy fixtures, construction — consider flipping the design: pick the windows randomly instead, alternating periods with and without the treatment in the same store. Slower, but it rescues tests that per-store assignment would price out of existence.

Define winning before you see the score

The most common corruption in retail testing is quiet and post-hoc: the pilot was about basket size, but basket size was flat, so the recap leads with the uplift found in one category, in one daypart, in week three. Choose the primary metric before launch, write it where everyone can see it, and pre-commit to the decision rule — roll, kill, or extend at threshold X. Secondary findings are allowed to be interesting. They are not allowed to be promoted to victory conditions retroactively.

Run long enough to bore everyone

Two weeks is a novelty bump with a start and end date. Staff behave differently around anything new; regulars notice the change; the effect breathes. Most store-level interventions need six to eight weeks before the signal settles — longer if the category is slow-cycle. The practical rule: when stakeholders stop asking about the pilot, it is finally producing trustworthy data.

Report the boring version

The honest pilot report is short: question, design, primary metric, result against the pre-agreed threshold, decision. Confidence intervals beat adjectives. "Inconclusive — extend four weeks" is a respectable outcome, and so is a clean kill: a pilot that stops a bad rollout has paid for the whole testing programme. The deck with seventeen upside slides and a hidden methodology page is how chains end up wallpapered with fixtures nobody can explain.

Screens are the easiest place to practise all of this — variants cost nothing and assignment is software. Start there: the menu board test design →