How Do I Track Which Aso Changes Actually Work?
App Store Optimisation is one of those disciplines where it is very easy to feel like you are doing the work without actually knowing if it is working. You change your subtitle, tweak a screenshot, swap out a keyword, and then watch your download numbers bounce around for a few weeks. Sometimes they go up. Sometimes they go down. Rarely do you know which change caused what. That gap between action and understanding is where most ASO efforts quietly fall apart.
We see this clearly on the gifting and wishlist platform we worked on, where the client was initially reluctant to invest in ASO at all. Their logic was that the product had strong referral mechanics built in, so word-of-mouth would carry the growth. Our view was different. Even if referrals bring someone to the App Store listing, the icon, copy, category, and screenshots still need to do real work. People need to understand the product from the first screen, and if they download it and then realise it is not for them, you have gained a download and lost a user. ASO is about ensuring the right people arrive already pre-sold.
What to leave alone during a test
Whilst any single test is running, leave everything else unchanged. This sounds obvious, but the temptation to keep improving is strong, especially when a team is actively engaged with the product. Resist it. A change made mid-test contaminates the result and means you have to start the clock again.
Setting a Baseline Before You Change Anything
A baseline is a four-to-six week window of data collected before you make any changes, giving you a stable reference point for everything that follows. Without it, you have no way to tell whether a movement in downloads after a change is meaningful or just the normal variance your product experiences week to week.
To set a useful baseline, you want to capture at least the following metrics across your chosen window: impressions, conversion rate from product page view to install, downloads, and where possible, retention at day one, day three, and day seven. Together these give you a picture of how well the listing is attracting and converting users, and what the quality of those users looks like.
You also want to note any external factors active during your baseline period. Was there a paid campaign running? Did the product receive press coverage? Was there a seasonal event that typically drives or suppresses downloads in your category? These factors do not invalidate the baseline, but they are context you will need when interpreting results later.
Run your baseline for at least four weeks to account for natural weekly variance. A two-week baseline will often look clean but will not capture the rhythm of your actual audience behaviour across different days and periods.
Running One Change at a Time
The single most common mistake in ASO testing is making multiple changes at once. It feels efficient. You have a list of things you want to improve, the product is live, and waiting four weeks between each change sounds painfully slow. But running multiple changes simultaneously means that when downloads move, you have no way to know which change caused it.
The discipline of one change at a time is the whole basis of being able to learn. Every change you make is a question. If you ask three questions at the same time and get one answer, you still do not know which question it was answering.
This is particularly true for the indexed fields. Because a title or subtitle change resets the stabilisation window, any concurrent keyword change makes it impossible to separate the effects. You end up with a four-week wait and no clear signal at the end of it.
Prioritising your test queue
The practical solution is to build a prioritised list of the changes you want to test, ordered by expected impact, and work through it sequentially. This converts the frustration of slow testing into something more useful: a structured roadmap with a clear hypothesis and expected outcome for each step. Teams that do this tend to learn faster over six months than teams that make ad hoc changes every few weeks, even though the ad hoc teams feel like they are moving faster.
How Long to Run a Test Before Reading the Results
The timing question is one of the most common sources of bad ASO decisions. Read results too early and you are reacting to noise. Read them too late and you are sitting on a change that is either helping or hurting you for weeks longer than necessary.
For changes to indexed fields, title, subtitle, keyword field, the minimum useful window is four weeks, based on the stabilisation period for keyword ranking. For non-indexed changes like screenshots, three weeks is generally enough to see a meaningful conversion signal, assuming your download volume is sufficient to generate reliable data.
Volume matters here more than teams typically appreciate. If your app receives forty downloads a week, a three-week test window gives you a hundred and twenty data points. Statistical patterns from a sample that small are fragile. A product with higher weekly volume will produce clearer signals faster. A lower-volume product may need longer windows, or may need to accept wider confidence intervals on its conclusions.
The practical answer is to set your window before the test starts, stick to it, and resist the urge to check results daily. Daily checking creates the illusion of movement where there is mostly fluctuation. Weekly reviews against your baseline give a cleaner read.
Distinguishing a Real Signal From a Coincidence
A movement in your download numbers after an ASO change is not automatically evidence that the change worked. Several things can produce download spikes that have nothing to do with what you changed: a competitor's app going offline or receiving negative press, an algorithm shift that temporarily boosts visibility in your category, or organic social posts about your product gathering unexpected traction.
The question to ask is whether the movement is consistent or isolated. A real signal from an ASO change tends to produce a sustained shift rather than a spike. Conversion rates improve and hold. A coincidental spike tends to revert over the following two weeks as the external factor passes.
Multiple variables also create compounding false positive risk. GrowthBook notes that with five unrelated metrics tracked in a single experiment, the probability of at least one false positive reaches 41%. This is worth holding in mind if you are watching many metrics simultaneously and tempted to call success on whichever one moved.
The most useful check is to look for corroborating signals. If conversion rate, download volume, and day-seven retention all improve together after a change, you have a much stronger case than if only one of them moved. A single metric moving in isolation warrants scepticism before action.
When Word-of-Mouth Obscures Your ASO Results
Products with strong social mechanics, referral loops, or community-driven growth face a specific ASO tracking challenge. When a significant share of your downloads comes from referrals, your store metrics can move independently of anything you did in the listing. A wave of organic sharing can inflate download numbers for two or three weeks in a way that looks exactly like a successful ASO test.
This was precisely the tension on the gifting and wishlist platform we worked on. The product was social and group-oriented by design, which made referral a plausible growth channel. But referral-driven downloads are not a substitute for ASO clarity. The people arriving via a friend's recommendation still land on the listing. If the icon and copy do not match their expectations of what the product does, you still see churn from users who download and then disengage.
The way to separate referral noise from ASO signal is to track where your downloads are sourced. Apple's App Store Connect and Google Play Console both provide channel attribution data. If a spike in downloads during an ASO test is almost entirely attributable to App Store browse or search, that is a clean signal. If it is coming predominantly from direct or referral sources, the test result is contaminated and the window needs extending.
Segment your download data by source before reading any ASO result. An unattributed download number is far less useful than one split between search, browse, and referral origins.
Category Selection as a Testable Variable
Category is one of the most underused levers in ASO, partly because it feels like a fixed decision made at launch and partly because the consequences of a change are harder to measure than a subtitle swap. In practice, category selection is a strategic variable worth revisiting as a product finds its audience, and it can be tested with the same discipline as any other change.
Our approach when a product could plausibly sit in more than one category is a two-stage process. Launch in the less competitive category first. Lower competition means higher rankings are achievable from a standing start, which generates more organic impressions and downloads. More downloads build usage and awareness. Once the product has traction, consider moving into the more competitive primary category that may be a better long-term fit, but is considerably harder to rank in without an established track record.
The reasoning is straightforward: a product buried at position 80 in a crowded category gets almost no organic visibility. The same product at position 12 in a smaller category gets meaningfully more, and that visibility compounds over time through downloads, ratings, and reviews. Entering the harder category later, with momentum behind you, is a better position than entering it cold.
Testing a category move
To test a category change, treat it like any other ASO variable. Log the change with a clear hypothesis, set a minimum four-week window, and track impressions and download volume as your primary metrics. A successful move to a less competitive category should show an improvement in search ranking positions and a corresponding lift in organic impressions.
Knowing When a Result Is Good Enough to Act On
Perfect certainty is not available in ASO testing, and waiting for it means either never acting or acting so slowly that the product has moved on before you conclude anything. The practical question is not whether you can be certain, but whether the evidence is strong enough to justify the next step.
A result worth acting on shows consistency across at least two primary metrics, sustains over the full length of your test window, and holds when you segment by source to remove referral noise. If those three conditions are met, the result is good enough to carry forward. Log it in your change log with the outcome, and move to the next item in your test queue.
A result that shows partial improvement, where one metric moves and others do not, warrants a different response. Run a second test that isolates the variable further before committing to the change as a permanent update. Partial signals are valuable as direction, but not as decisions.
The bigger risk is acting on too little. A one-week spike in downloads that returns to baseline in week two is noise. Treating it as confirmation that a change worked, then making the next change on top of it, compounds the problem and makes the testing record meaningless over time.
Building a Testing Cadence Your Team Will Actually Follow
A testing system that requires unusual discipline to maintain will not survive contact with a real product team. Launches slip, priorities shift, and if ASO testing depends on someone remembering to log a change or check a result, it will fall apart during the first busy month. The system needs to be simple enough that following it is the path of least resistance.
The elements of a cadence that tends to hold up in practice are these.
- A shared change log updated every time a store listing edit goes live, not retrospectively.
- A fixed weekly review time, even if it is fifteen minutes, where one person checks the current test against its baseline and records what they see.
- A decision rule agreed in advance: at the end of the test window, the result is logged, the change is kept or reverted based on the pre-defined criteria, and the next item in the queue is started.
- A review of the test queue every four to six weeks to reprioritise based on what the product has learned.
The cadence works because it removes the decision of whether to test. The queue exists, the window is set, and the review happens on a fixed schedule. What changes is the content of the tests, not the structure around them.
This also connects to a broader principle we hold about product iteration generally. Launching and iterating based on real user feedback is more effective than trying to perfect a product before release. The same logic applies inside ASO. A working testing system that produces imperfect data is worth considerably more than a theoretically ideal system that never gets used because it demands too much effort to maintain.
Conclusion
ASO tracking is not complicated in theory. Make one change, wait for the data to stabilise, compare against your baseline, record what you found, and move to the next test. The difficulty is in doing that consistently over months, in the face of noisy data, product pressures, and the constant temptation to make one more quick change before the last one has settled.
The tools are secondary to the habit. A spreadsheet used consistently outperforms sophisticated analytics that nobody checks. A change log maintained from day one of a product's store life gives you something genuinely useful: a record of what you tried, what you learned, and what you now know about your audience that you did not know before.
The gifting and wishlist platform taught us that even products with strong organic referral mechanics still need their listings to do real work. Category selection, copy clarity, and screenshot quality all affect whether the right users arrive and stay. Getting those elements right is a testing problem, and testing problems need systems.
If your team is making ASO changes without a reliable way to read the results, the changes are decoration. The system is what turns them into learning.
Let's talk about your ASO testing strategyFrequently Asked Questions
A baseline gives you a stable reference point so you can tell whether any movement in downloads after a change is meaningful or just normal weekly variance. Without four to six weeks of data collected before you start testing, you have nothing reliable to compare your results against. Metrics to capture include impressions, conversion rate, downloads, and early retention figures.
You should run your baseline for at least four weeks, as a shorter window will not capture the natural rhythm of your audience behaviour across different days and periods. A two-week baseline can look deceptively clean but may miss patterns that only emerge over a longer stretch. Six weeks is preferable if your category experiences noticeable seasonal variation.
Making multiple changes at once means that when downloads shift, you have no way of knowing which change caused the movement. Each change is essentially a question, and if you ask several questions simultaneously and get a single result, you still cannot identify which question was answered. Testing one change at a time is the only way to draw clear, actionable conclusions.
Any change made while a test is already running contaminates the results, making it impossible to attribute any movement to a single variable. When this happens, you effectively have to start the clock again and wait for a fresh testing window. The temptation to keep improving mid-test is understandable, but it will cost you reliable data.
Yes, because even users who arrive via referrals still land on your App Store listing and need to quickly understand what the product is. If the icon, copy, screenshots, or category fail to pre-sell the right audience, you risk gaining downloads from people who then churn immediately. A download from someone who is not a good fit costs you more than no download at all.
The core metrics to follow are impressions, conversion rate from product page view to install, and total downloads. Where possible, you should also track retention at day one, day three, and day seven, as this tells you something about the quality of the users your listing is attracting. Together these metrics show both how well your listing converts and whether it is bringing in users who actually stick around.
Yes, you should note anything that might be influencing your numbers during the baseline period, such as a paid campaign, press coverage, or a seasonal event in your category. These factors do not make the baseline invalid, but they provide important context when you come to interpret your test results later. Ignoring them can lead you to misread what is actually driving any changes you observe.
Changes to indexed fields like the title or subtitle reset the stabilisation window that the app store uses to process and reflect those updates. If you make a concurrent keyword change at the same time, you cannot separate the effects of each individual change. You end up waiting the full testing period and still have no clear signal to act on at the end of it.