Statistics · 16 March 2026

The Texas sharpshooter problem, or: why your test 'won with iOS users in Leeds'

Slice any flat test into enough segments and one of them will always be a winner. Post-hoc segmentation is where inconclusive tests go to be laundered into success stories.

A scatter of grey dots with a target drawn around one random cluster of blue dots

The Texas sharpshooter fires his rifle at the side of a barn, walks over, and paints a target around the tightest cluster of bullet holes. Bullseye.

CRO has its own version, and it happens in the meeting after a test comes back flat. Nobody wants to write "inconclusive" after three weeks of work, so someone opens the segment report. Desktop? Flat. Mobile? Flat. Mobile Safari? Down a bit. Returning visitors on Android who arrived from paid social? +34%, significant! The deck gets a slide: "variant wins with our high-intent mobile audience."

The target has been painted around the bullet holes.

Why this is guaranteed to happen

Run a test at 95% confidence and each segment you examine carries roughly a 5% chance of a false positive on its own. Check twenty segments — device × browser × new/returning × channel gets you there fast — and the chance that at least one shows a spurious "significant" result is over 60%. On a genuinely null test. Every time.

And it's worse than the raw multiple-comparisons maths, because segments are small. Small samples produce wild swings, so the most extreme numbers in any segment report will always come from the thinnest slices — which are precisely the ones people find most exciting. "+34% with iOS users in Leeds" is exciting because only 212 people were in that bucket.

The test for whether a segment finding is real

One question: was this segment named before the test launched?

If yes — "we hypothesise this checkout change helps mobile users most, because the current flow requires horizontal scrolling on small screens" — then the segment read is legitimate. It was a prediction, and the data confirmed or refuted it. Ideally it was also powered: you checked the segment would have enough traffic to say anything.

If no, the finding is a hypothesis, not a result. That's the honest reframe, and it's genuinely useful: "this flat test generated the hypothesis that mobile users respond differently — let's design a test to check that." The sharpshooter's sin isn't noticing the cluster. It's claiming he aimed at it.

The corporate incentives that keep the sharpshooter employed

Post-hoc segment fishing persists because everyone in the room is rewarded for finding it. The agency gets a win for the monthly report. The stakeholder who championed the test saves face. The analyst gets to present something other than a shrug. The only party unrepresented in the meeting is the truth — and the future team who'll roll the change out to that segment, watch nothing happen, and quietly lose faith in testing.

This is a place where process beats virtue. Two rules do most of the work:

  1. Pre-register the primary metric, the success threshold, and any segments of interest before launch. Written down, dated, in the test doc. Anything not on the list is exploratory by definition.
  2. Label exploratory findings as hypotheses in every report, in those words. "Hypothesis generated: mobile-first users may respond better — candidate for a follow-up test" is a sentence stakeholders accept surprisingly readily, once it's the house style.

Flat tests are frustrating. But a flat test honestly reported keeps its full value: the change didn't matter, stop investing in it. A flat test laundered through segment fishing costs you twice — once for the rollout that won't work, and again for the real question you never got round to asking.