App Store Ratings by Country: 12 Apps, 5 Storefronts
Back to Blog

App Store Ratings by Country: 12 Apps, 5 Storefronts

We checked 12 apps across five App Store countries. Compare real rating scores and counts, download the 60-row dataset, and build a useful local baseline.

September 10, 20267 min

App Store ratings depend on the storefront you check. On September 10, 2026, we looked up the same 12 apps in the United States, United Kingdom, Germany, Poland and Japan. All 60 app-country pairs returned a matching app in five batched requests. Every app had a different rating across those five storefronts; the largest gap was Calm, at 4.339 in Japan and 4.788 in Poland.

A comparison of observed App Store ratings across five storefronts for twelve selected apps.

This is a small, deliberately chosen comparison that you can inspect and reproduce. It answers a practical question: what goes wrong when a competitor report uses a US rating to describe an app in another country? Download the 60-row public dataset to check each observation.

What we measured

We selected familiar apps across education, music, productivity, fitness, wellbeing, finance and design: Duolingo, Spotify, Notion, Strava, Calm, Headspace, Revolut, Wise, Canva, Todoist, Forest and AllTrails. We chose the names before looking at the country differences. This is a purposive sample, so its percentages and averages cannot estimate the wider App Store.

We verified each numeric app ID against its official Apple product page, then sent five public iTunes Lookup requests, each containing all 12 IDs. We fixed the entity to software and the response language to English, changing only the country. Requests ran from 04:25:19 to 04:25:35 UTC on September 10, 2026, with at least 3.2 seconds between requests. No AppDrift customer data was used.

Apple's archived API documentation identifies the country parameter as the storefront selector and lists an approximate limit of 20 calls per minute.[1]

We used numeric IDs to avoid substituting a similarly named app. Apple's lookup documentation describes ID-based retrieval as a way to reduce false positives.[2]

The fields here are averageUserRating and userRatingCount. We round scores to three decimals for comparison; these are API values, not the one-decimal labels in the App Store interface. Counts are ratings, not written reviews. The snapshot contains neither downloads nor revenue. Every app returned the same version number across its five storefronts, although that does not make its users, feature exposure or rating history identical.

The lowest and highest rating for each app

The counts in this table belong to the storefronts with the lowest and highest scores. They are not necessarily the smallest and largest rating counts. GB means the United Kingdom; DE Germany; PL Poland; JP Japan. Country links open the corresponding Apple product page, whose live figures may have changed since collection.

Public iTunes Lookup observations, September 10, 2026. Spread = highest minus lowest unrounded score.
AppLowest score · country · countHighest score · country · countSpread, stars
Duolingo4.543 · JP
750,827 ratings
4.734 · PL
295,437 ratings
0.192
Spotify4.581 · JP
5,038,483 ratings
4.805 · PL
1,455,172 ratings
0.224
Notion4.646 · JP
50,050 ratings
4.777 · US
90,041 ratings
0.130
Strava4.528 · DE
27,193 ratings
4.811 · US
372,239 ratings
0.283
Calm4.339 · JP
9,913 ratings
4.788 · PL
5,578 ratings
0.449
Headspace4.753 · JP
3,447 ratings
4.854 · PL
7,584 ratings
0.101
Revolut4.532 · JP
6,446 ratings
4.877 · PL
207,408 ratings
0.345
Wise4.642 · US
109,784 ratings
4.791 · GB
168,254 ratings
0.149
Canva4.722 · JP
551,737 ratings
4.902 · PL
87,430 ratings
0.180
Todoist4.612 · JP
21,767 ratings
4.839 · PL
5,917 ratings
0.227
Forest4.698 · JP
10,621 ratings
4.853 · PL
7,527 ratings
0.155
AllTrails4.637 · DE
8,003 ratings
4.889 · US
1,038,380 ratings
0.252
Lowest and highest API rating for each of 12 apps across five storefronts; Calm has the widest spread at 0.449 stars
Each line connects one app's lowest and highest observed storefront rating. These are country differences in one snapshot, not changes over time.

Observation 1: a country omission can change the comparison

Six of the 12 apps had a spread above 0.2 stars. The median spread was 0.208. Calm had the widest difference, 0.449, followed by Revolut at 0.345 and Strava at 0.283. Headspace had the narrowest, 0.101. These are descriptions of the selected apps, not thresholds for deciding whether a rating difference matters.

Consider a finance-app comparison. Revolut's observed score exceeded Wise's in the US: 4.722 versus 4.642. In Japan the order reversed: Wise scored 4.677 and Revolut 4.532. Copying the US numbers into a Japanese competitor brief would give the wrong order for this metric. It would still tell us nothing about which app is better for a particular customer.

Apple explicitly says an app's summary rating is specific to each territory. The country belongs beside the score wherever you store or report it.[3]

Observation 2: the rating count changes much more than the score

AllTrails returned 1,038,380 ratings in the US and 759 in Japan: a 1,368-fold difference. Its scores were 4.889 and 4.740 respectively. Headspace returned 973,855 US ratings and 3,447 Japanese ratings, a 283-fold difference, while the two scores were 4.842 and 4.753.

Across all five storefronts, the largest rating count was more than ten times the smallest for 11 of the 12 apps. Forest was the exception, with a 6.54-fold difference. Those ratios describe how many ratings the endpoint returned. They do not measure the ratio of downloads, customers or market opportunity.

AllTrails rating counts across five storefronts, from 759 in Japan to 1,038,380 in the US, on a logarithmic scale
AllTrails rating counts by country, on a logarithmic scale. Count extremes can differ from score extremes.

A report that keeps the star score but drops the count loses useful context. Write “4.740 from 759 ratings, Japan, checked September 10” so someone reviewing the finding can see what was actually measured. Keep that baseline when investigating subsequent changes through store monitoring.

Observation 3: the US is not a global baseline

The US had the highest observed score for three apps: Notion, Strava and AllTrails. Poland had the highest for eight; the UK had the highest for Wise. That does not establish a national tendency to give higher scores. We did not sample users or control for the app experience, acquisition mix or time spent in each market.

The useful observation is narrower: none of these apps can be fully described by copying one country's score into every country row. Store the numeric app ID and country together. Keep country consistent when moving from ratings to keyword tracking, and treat a ranking result as a separate measurement.

Keep a date beside each observation too. Our study of day-to-day App Store ranking movement shows why a position captured last month should not be presented as a current reading.

A comparison workflow you can use today

  1. Choose the decision and storefront first. For a German launch, compare the apps people would encounter in Germany. Use an explicit country field rather than an unlabeled default.
  2. Match the app and platform. Save the numeric Apple ID, bundle ID, country and software entity. A similarly named app or a different platform listing can produce a plausible but irrelevant number.
  3. Capture the score, count, version and timestamp together. Preserve the raw response. If a request fails, record the failure; if it returns no matching result, record that separately. Neither means zero ratings.
  4. Read local written feedback before proposing a fix. Look for specific, recurring issues and check whether they still affect the current version. Our review-response guide covers turning that reading into a useful response.
  5. Measure the proposed change separately. If local feedback exposes unclear copy, review the localization checklist before using metadata translation. Compare later outcomes within the same country and record concurrent releases. This snapshot cannot predict an uplift.

What this snapshot cannot explain

We did not collect rating histories, written review text, reset dates or acquisition cohorts. Apple allows a rating reset with a new release; it affects all countries simultaneously for the selected platform and leaves written reviews intact. We cannot identify whether a sampled app used that option.[4]

The API may cache results, so retrieval time is not the creation time of the underlying ratings. These observations were sequential, not simultaneous. A returned app confirms a matching public record in that lookup; it does not test installation or product availability for every person. Repeating this method later would produce a second snapshot, not an explanation of why the numbers changed.

Frequently asked questions

Are App Store ratings different by country?

Yes. Apple states that summary ratings are territory-specific. Each of the 12 apps in this September 10, 2026 snapshot returned different scores across the five storefronts checked.

Does userRatingCount mean written reviews?

No. It is the rating count returned by the lookup endpoint. A star rating and a written review are different records; this dataset does not count review text.

Can rating counts estimate downloads by country?

Not from this dataset. We did not measure downloads or the share of users who rate an app. A larger count cannot be converted into a download estimate here.

Does a lower country rating prove a localization problem?

No. A country difference is a reason to investigate local feedback, product experience and release history. This observational sample cannot attribute the difference to translation or any other cause.

How can I reproduce the comparison?

Use the app IDs and request URLs in the downloadable CSV. Keep the software entity and response language fixed, specify each country, and save fresh timestamps and responses. Respect Apple's published API request limit; expect live values to change.

Sources

  1. Apple: public search API parameters and request guidance
  2. Apple: ID-based lookup examples
  3. Apple: ratings, reviews and responses
  4. Apple: reset an app overview rating

App Localization

Translate your app to 60+ languages with AI

  • Cultural adaptation, not just word-for-word translation
  • AI-powered — native-quality results in minutes
  • Localized screenshots for every market
Try AppDrift FreeFree to start · No credit card

Get ASO tips that actually work

Free ASO checklist + weekly insights from 10,000+ developers shipping in 60+ languages. No spam, unsubscribe anytime.