Do apps with 4.8-star ratings rank higher? We analysed 1,000 App Store searches

Apple names ratings as a ranking input, so the folklore follows: get to 4.8 and you climb. We took 1,000 result sets, 18,796 individual search results, and looked. The star value explains almost nothing. The number of ratings explains six times more.

Why the question is worth asking properly

Apple states that search results are based on “text relevance (matches for your app’s title, subtitle, keywords, and primary category), as well as user behavior (downloads, ratings and reviews, and more)”. That single sentence is the origin of an enormous amount of advice about chasing a higher star average, usually with 4.8 named as the threshold that matters.

But “ratings and reviews” is ambiguous in exactly the place that matters. It could mean the star value, or the number of people who left one, or both. Those are different things and they imply completely different work: raising an average is a product problem, and raising a count is a distribution problem.

Method

We collect daily search results for the terms our customers track. For this study we took every distinct term and country the collector saw in the seven days to 12 September 2026, each counted once at its most recent observation, and kept result sets where at least eight of the top 20 apps had both a star value and a rating count on file. That gives 1,000 result sets and 18,796 individual search results.

Taking each term once matters. The same search on consecutive days returns almost the same apps in almost the same order, so counting it twice would double the sample without adding any information and would make every figure look more certain than it is.

The raw picture: almost flat

There is a gradient, and it is tiny. The best-rated band is not the best-placed band: 4.5-4.7 edges it. Below 4.5 the difference is about one position out of twenty, and the under-3.0 group places better than the 3.0-3.9 group, which is the kind of inversion you see when an effect is mostly noise.

A negative number means the higher value sits nearer position one. Both do, and the count is roughly six times the stronger signal. Neither is large: even the count explains only a few per cent of where an app sits, which is a useful reminder that relevance, not popularity, does most of the work in a search result.

The confound, and what happens when you remove it

Stars and rating counts travel together, so a raw comparison cannot separate them. In our sample the 4.5-4.7 band has a median of 13,507 ratings and the 3.0-3.9 band has 72. Comparing those two groups is mostly comparing big apps with small ones.

So we split the same placements into bands of similar rating count and looked at the star effect inside each band, where the confound is largely held still.

Read down a column and the effect of scale is obvious: the same star band moves from about position 12 to about position 9 as the rating count grows. Read across a row and the star effect is small, inconsistent, and in the smallest band runs backwards — apps rated 4.8-5.0 with under 200 ratings sat lowest of any group in the table.

That inversion is not mysterious. A 4.9 average from 30 reviews usually belongs to a very small app, and being very small is what shows up in the position. It is a clean demonstration of why the raw gradient should not be read as a star effect.

A cleaner test: compare inside one result set

Comparing a finance term against a puzzle-game term compares two different populations. Comparing the top three against positions 8-20 of the same search does not. We ran that paired test on the 990 result sets with enough scored apps at both ends.

The apps at the top of a result set are rated 0.023 stars higher than the apps at the bottom of it. They have roughly nine times as many ratings. The higher-rated group held the top three in 52.6% of sets and the bottom in 22.8%, so the direction is real — it is the size that is negligible.

What the literature says

This is not a new question academically, and the published work lands in a similar place. Sällberg, Wang and Numminen, writing in the Journal of Marketing Analytics in 2022, tracked 341 apps daily for about two years from their App Store launch and separated rating information from review text. For productivity apps they report that the average rating score had no significant direct effect on downloads, while the volume of ratings did. For gaming apps the volume effect was negative, which is itself a warning against treating any of these as simple levers.

Work on visual assets points the same way about the limits of single-variable thinking: Wang and Li, in Electronic Commerce Research in 2016, extracted colour, complexity and symmetry features from Android icons and found measurable relationships with download behaviour. The broader field of app store mining opened with Harman, Jia and Zhang at MSR 2012, which established that store metadata can be analysed at scale at all.

What this means for your listing

Limits of this study

One week, iOS only, and the 3,564 terms our collector follows, which lean toward the categories our customers work in rather than a random sample of the App Store. Rating count is a proxy for install base and a noisy one, since apps prompt at very different rates. Position is observational: we cannot randomise an app’s star average, so nothing here is a causal estimate. And a correlation this small would be swamped by any category effect we have not modelled.

What survives those limits is the comparison between two signals measured the same way on the same data. Whatever bias sits in our sample applies to both, and the count still outperforms the star value by roughly six to one.

Blog · Free ASO tools · Products · AsoTheory