Skip to content

Blog

A/B testing your screenshots without fooling yourself

Both stores will happily show you a winner that is noise. What a screenshot test can and cannot tell you, how long to actually run it, and the four ways these tests go wrong.

Two smoked glass panels on either side of a chrome balance beam

Your opinion about your own screenshots is the least reliable input available. You have seen the app ten thousand times, you know what every screen does, and you cannot un-know it. Testing is the only way out of that.

But a test you read wrong is worse than no test, because it converts a guess into a confident guess. Here is how these actually behave.

What you have to work with

Apple — Product Page Optimization. Up to three treatments against your current page, split from your existing App Store traffic, reported in App Store Connect with an indication of confidence.

Google — Store Listing Experiments. Same idea in Play Console, with more variants available and the ability to test the icon and feature graphic too.

Both are free, both use real traffic on your real listing, and both are dramatically better than asking your team which version they prefer.

The number that decides everything

Not conversion rate — traffic volume. A screenshot test needs enough visitors to distinguish a real difference from random variation, and the smaller the true effect, the more you need.

The uncomfortable arithmetic: detecting a small improvement takes far more traffic than detecting a large one. If your listing gets a few hundred impressions a week, you cannot detect a 3% lift in any reasonable timeframe — and 3% lifts are the common case. You can still detect a large one, which is an argument for testing bold changes rather than cautious ones.

The four ways these go wrong

1. Peeking

You check the results daily, see your variant ahead on day three, and ship it.

This is the most common error and it is a real statistical problem, not fussiness. Early in a test the numbers swing wildly; if you look repeatedly and stop when you like what you see, you will "find" winners that are noise. Both stores report confidence — wait for it, and decide the stopping rule before you start.

2. Changing more than one thing

A variant with a new first screenshot, new captions and a new palette that wins tells you something worked. It does not tell you what, so you cannot do it again. One change per test, always.

3. Testing panel six

Almost nobody swipes that far. A change there affects a small fraction of visitors, so the effect on overall conversion is tiny and needs enormous traffic to detect. Test the first screenshot. Then the second. Then caption angle. Then order.

4. Seasonality and novelty

A test running across a holiday, a launch, a press mention or a paid campaign is measuring that event as much as your screenshots. Run tests over whole weeks, and be suspicious of anything that overlaps a spike in traffic from a single source.

Two smoked glass panels balanced on either side of a chrome beam
The test is only as good as the discipline around it: one variable, a stopping rule set in advance, and enough traffic to see the effect you are looking for.

What to test, in order

  1. The first screenshot. Largest audience, largest possible effect. Test a genuinely different approach, not a tweak.
  2. The second screenshot. Second-largest audience, same logic.
  3. Caption angle across the set — outcome-led against feature-led is a real, testable difference.
  4. Panel order. Cheap to produce, occasionally surprising.
  5. The icon. Huge effect when it moves, but it also affects recognition for existing users, so treat it carefully.

On the numbers you see quoted

You will find case studies claiming very large lifts from screenshot changes. Some are real. Most come with no traffic figures, no test duration, no confidence interval, and no mention of what else changed that month.

We are not going to quote any of them at you, including the flattering ones, because a number without its methodology is decoration. The useful claim is much more boring: first-impression assets carry the most leverage on a product page, and testing them against real traffic beats arguing about them. That is well supported. The specific percentage you will get is not knowable in advance, and anyone telling you otherwise is selling something.

Why this stalls in practice

Not the statistics. The production.

Each variant needs a full screenshot set, at every size. If producing a set is a multi-day design job, a three-variant test is a month of work nobody schedules, and the test never runs. That is the actual reason most teams have never tested their listing.

The fix is making the second version cheap: fork the set, change the panel under test, re-render that panel rather than the whole thing. Variants stay visually consistent with the original because they came from it — which also keeps the test clean, since you are measuring the change rather than an incidental redesign.

Practical detail on both stores' tools in product page optimization, and on what to change first in ASO screenshots.

ASO

The App Store screenshot optimization playbook

The whole practice in one place — what each panel is for, how to write captions people actually read, what to test and in what order, and the compliance lines you cannot cross. The hub for everything else we have written on this.

Read
ASO

Are your App Store screenshots costing you installs?

A diagnostic you can run in ten minutes, on your own listing, without any tools. Six specific failures that leak installs from traffic you already earned — and which ones are worth fixing first.

Read
ASO

Custom Product Pages: matching the listing to the ad that earned the tap

A CPP is a second product page with its own URL, and the reason to build one is message match. When it is worth the maintenance, how many you actually need, and the mistake that makes them worthless.

Read