---
title: How can you test and optimise AI driven personalisation features?
description: Learn how to properly test AI driven personalisation, from behavioural feedback loops to the signals that matter beyond conversion rates.
image: https://weareaffective.com/hubfs/learning-centre-images/how-can-you-test-and-optimise-ai-driven-personalisation-features.webp
---

[Skip to content](https://weareaffective.com/learning-centre/how-can-you-test-and-optimise-ai-driven-personalisation-features#main-content)

[![we\_are\_affective\_logo\_200](https://weareaffective.com/hs-fs/hubfs/we_are_affective_logo_200.png?width=175&height=48&name=we_are_affective_logo_200.png "we_are_affective_logo_200")](https://weareaffective.com)

- [Home](https://weareaffective.com)
- About Us 
  
    - [Our Story](https://weareaffective.com/about)
    - [How We Work](https://weareaffective.com/how-we-work)
- Our Services 
  
    - [App Planning & Strategy](https://weareaffective.com/app-planning-strategy)
    - [App Design](https://weareaffective.com/app-design-agency)
    - [App UX Design](https://weareaffective.com/app-ux-design)
    - [App UI Design](https://weareaffective.com/app-ui-design)
    - [App Technical Architecture](https://weareaffective.com/app-architecture)
    - [Existing App Audits](https://weareaffective.com/app-audit)
- [Case Studies](https://weareaffective.com/case-studies)
- [Pricing](https://weareaffective.com/pricing)
- [Learning Centre](https://weareaffective.com/learning-centre)

- [Get Started](https://weareaffective.com/get-started)

Expert Guide Series

# How can you test and optimise AI driven personalisation features?

 Table of Contents

Personalisation that works feels invisible. The product seems to understand you without announcing that it does, and you move through it faster, stay longer, and come back more readily. Personalisation that fails feels intrusive or arbitrary, and users either ignore it or leave. The difference between those two outcomes is rarely in the algorithm. It is in how the personalisation gets tested, what signals teams choose to measure, and whether the feedback loops are built to catch emotional misfires, not just conversion dips.

> The difference between personalisation that works and personalisation that fails is rarely in the algorithm itself.

The challenge is that AI-driven personalisation does not behave like a static feature. It adapts, which means the thing you tested last month is not quite the thing running today. That makes standard testing approaches inadequate on their own, and it means teams can be sitting on data that looks clean while users are quietly disengaging in ways that no dashboard is tracking.

We work at the intersection of [behavioural psychology and product design](https://weareaffective.com/app-user-research), which means we see testing failures that are interpretive. The data was there, but the team was asking the wrong questions of it, or measuring the wrong moments, or optimising for a signal that was only loosely connected to what users actually needed. This article sets out how to do it properly.

## What AI-driven personalisation actually needs to be tested

A common assumption is that personalisation testing means checking whether the model's recommendations are accurate. That is one layer. But the more consequential question is whether the experience of being personalised to feels right to the user, and that question lives several layers above the model itself.

What actually needs to be tested falls into three broad areas. First, the quality of the input signals: is the product reading the right behavioural data to inform personalisation decisions, and is it reading it at the right moments? Second, the presentation layer: how and when does personalised content surface, in what language, and with what surrounding context? Third, the emotional response: does the personalisation make users feel understood, or does it make them feel watched?

Each of these areas requires a different testing approach. Model accuracy can be assessed quantitatively. Presentation layer decisions respond well to controlled experiments. Emotional response needs [qualitative methods alongside the numbers](https://weareaffective.com/learning-centre/5-user-testing-methods-that-will-save-your-app-from-failure). Teams that only run quantitative tests are flying with one instrument.

Dwell time on certain screens, speed of movement through a flow, return visit patterns, and task completion rates all carry information about how users are responding emotionally, not just functionally. Those behavioural signals are the raw material for understanding whether your personalisation is landing or not, and they need to be treated as such from the start of any testing programme.

## Why A/B tests alone are insufficient for personalisation

A/B testing assumes that a variant either works or it does not, and that the population tested against is reasonably stable. Neither assumption holds well for adaptive personalisation. The variant changes based on user behaviour, so the thing you are testing is a moving target. And the population is segmented by the algorithm itself, which means your control group and your test group may differ in ways that confound the results.

The deeper problem is what A/B tests optimise for. They measure a defined outcome, usually a click, a conversion, or a session length figure. These are downstream signals. By the time a drop in conversion shows up, the emotional damage is already done. Users who felt a product was reading them strangely have already formed a view of it, and that view does not reset when you roll back the variant.

There is also a gap between what self-reported data captures and what behaviour reveals. NPS scores and satisfaction surveys collect feedback after the fact, when users are in a calm, analytical state rather than the emotionally heightened moment of actually using the product. Behavioural data captures the in-context reality. The two need to sit alongside each other, because each tells you something the other cannot.

According to [McKinsey](https://www.mckinsey.com/business-functions/marketing-and-sales/our-insights/the-value-of-getting-personalization-right-or-wrong-is-multiplying), 71% of customers expect personalised communications, which means the bar for what counts as acceptable personalisation is already set by the wider market. Meeting that bar requires knowing whether users converted and whether the experience felt right to them in the moment.

## UX/UI design built around *real* psychology

We design app interfaces around how people actually think and behave. User research, psychology-driven UX/UI design and technical specs delivered as one complete package.

[See how we work](https://weareaffective.com/how-we-work) [Get started](https://weareaffective.com/get-started)

No commitment

## Building a behavioural feedback loop

A behavioural feedback loop connects what users do inside the product to what the personalisation engine does next, and it also connects that usage data back to the team so that human interpretation can sit alongside algorithmic decision-making. Without the second part, the loop is closed in the wrong direction, optimising the model without improving the experience.

The inputs worth feeding into a loop for personalisation go beyond clicks. Dwell time on specific screens tells you where attention is held and where it is lost. Speed of movement through a product gives a rough signal of confidence or uncertainty. Return visit frequency and patterns show whether the product has earned a place in the user's routine. Task completion rates, and particularly whether users struggle with the same steps repeatedly or move fluidly across different tasks, indicate whether the product is reducing or adding to cognitive load.

> Behavioural signals are the raw material for understanding whether personalisation is landing or misfiring.

These are indicators of emotional state, and they allow the team to adapt personalisation strategy, including the [visibility of features, the terminology used](https://weareaffective.com/learning-centre/progressive-disclosure-isnt-just-about-information-its-about-building-confidence), and the tone of communications, in response to how users are actually feeling rather than how the product assumes they feel.

The loop only works if someone is [reading the signals and asking what they mean](https://weareaffective.com/learning-centre/how-to-read-a-user-session-recording-for-emotional-signal-rather-than-task-compl). Analytics tell you what happened. The interpretation of why, and what to do about it, still requires human judgement sitting on top of the data.

Set up a weekly review of behavioural signals alongside your standard analytics. Look for screens where dwell time has changed, flows where drop-off has shifted, and tasks users are repeating. These patterns carry more information about personalisation quality than conversion figures alone.

## Signals worth measuring beyond conversion rates

Conversion is the most common thing teams measure because it is the easiest to define. But for personalisation specifically, it is a late signal. A user who converts despite a poor personalised experience is still a user who has formed a negative impression, and that impression affects retention even if it never shows up in the conversion data.

The signals that give earlier and more accurate information about personalisation quality include the following.

- Return visit rate within the first 72 hours of a new feature being surfaced to a user
- Time to second engagement with a personalised recommendation
- Rate at which users dismiss or hide personalised content
- Notification open rate broken down by content type and timing, not just overall
- Feature adoption depth, meaning how far into a personalised flow users go before stopping
- Qualitative feedback referencing the product feeling relevant or irrelevant

Whether users are sharing the product with others is a particularly revealing signal. Recommendation behaviour reflects genuine enthusiasm, and it is one of the harder signals to fake through dark patterns. A product that users are recommending to friends is doing something right emotionally, not just functionally.

The combination of hard behavioural data and softer sentiment signals creates a more complete picture than either alone. Teams that treat sentiment as a soft and therefore optional metric are typically the ones surprised when strong conversion data is followed by poor retention. The sentiment was telling them something the conversion numbers were not.

## How expectation-setting before a flow changes what your data tells you

One of the most reliable ways to misread drop-off data is to ignore what users knew before they entered a flow. If a user starts a multi-step personalisation questionnaire without any sense of how long it will take, drop-off at step three tells you something different than if they were told upfront what to expect and still left at step three. The data point is the same. The meaning is opposite.

On a health and wellbeing product we worked on, we tested two versions of a multi-step flow. One primed users upfront with a clear indication of how long the process would take. The other simply showed a progress bar without any time expectation. Without priming, drop-off rates sat at around 80 to 85%, almost always within the first three or four questions. After introducing upfront priming, completion rates rose to approximately 95%. The same flow, the same questions, a completely different outcome driven by what users knew before they started.

This matters for personalisation testing because onboarding flows and preference-setting screens are where personalisation gets its foundational data. If users are abandoning those flows, the model is working with incomplete information, and every downstream personalisation decision is built on a thin foundation.

Before testing drop-off within a personalisation flow, check what expectation-setting happens before it starts. Poor completion rates are often a priming problem rather than a flow problem, and fixing the wrong thing wastes the whole test cycle.

Expectation-setting also changes what your completion data tells you about user intent. A user who completes a flow they knew would take five minutes has demonstrated a different level of engagement than one who stumbled through it without context. Both show as completions. Only one represents genuine investment in the personalised experience.

## Using notification timing as a testable personalisation variable

Notifications are one of the most directly testable parts of a personalisation system, and they are also one of the most commonly mishandled. The default approach is to treat notification preferences as a binary setting: users either receive them or they do not. That binary produces almost no useful data about what would actually serve users well.

A more useful framing is to ask, for each type of notification, whether the user has implicitly or explicitly requested that information through their behaviour in the product. Notifications aligned with what users are already doing feel like helpful prompts. Notifications sent independently of context feel like interruptions, and users learn quickly to ignore or disable them.

On a concierge app for a premium property development, we pushed for drip-fed, context-sensitive notifications rather than the default of delivering all information upfront. The client accepted the recommendation, and subsequent A/B testing and research with real users confirmed it worked well. Users who received the delayed, contextual notifications reported measurably lower stress levels than those who did not. The notifications came to feel like timely assistance rather than demands for attention.

Timing is a variable that almost no team tests systematically. According to [Leanplum](https://izooto.com/blog/best-time-of-day-to-send-push-notifications), personalised push notifications have four times the open rate of generic ones, and timing is a core part of what makes a notification feel personalised rather than broadcast.

Run a granular notification preference audit before optimising timing. Ask whether users can reduce notification frequency without losing genuine value from the product. If they can, the current cadence is higher than user benefit requires, which means timing tests will be optimising the wrong baseline.

## Mapping the emotional arc to interpret behavioural data correctly

Behavioural data describes what users did. It does not, on its own, explain why. The same drop-off rate at the same screen can mean the task felt too hard, the content felt irrelevant, the user was interrupted, or the personalisation surfaced something they found off-putting. Reading that as a single problem leads to the wrong fix.

One of the artefacts we produce from research debriefs is a document mapping the emotional arc of users through a feature or product experience, capturing [how users feel at each stage](https://weareaffective.com/learning-centre/how-to-tell-the-difference-between-a-user-problem-and-a-user-preference), where their confidence is high, where anxiety creeps in, and where the experience either matches or violates their expectations. These emotional arc maps sit alongside brand personality documents to ensure that implementation decisions are emotionally coherent, not just functionally justified.

An emotional arc map changes how teams interpret behavioural data because it gives each data point a context. A high dwell time on a screen that the emotional arc identifies as a moment of decision is a different signal than high dwell time on a screen that should feel smooth and effortless. Both are delays. One is productive, one is a friction point. Without the emotional context, teams typically treat all dwell time as positive engagement, which leads to misguided conclusions.

The practical step is to build emotional arc mapping into your research process before you run personalisation tests, so that when data comes back, you have a framework for interpreting it that goes beyond surface behaviour. This is how you avoid optimising for the wrong moment.

## Common mistakes teams make when optimising personalisation

The most frequent mistake is treating personalisation as a feature to be shipped rather than a system to be calibrated over time. A team builds a recommendation engine, runs an A/B test, sees a lift in the primary metric, and considers the work done. The system then drifts as user behaviour changes, model outputs shift, and the emotional quality of the experience degrades quietly without triggering any of the metrics being watched.

#### Optimising for platform metrics over user benefit

A useful audit question here is what we think of as the removal test: what would happen to engagement if a specific personalisation feature were turned off for a segment of users? If the honest answer is that the team would never do that, that reaction is itself informative. It signals the feature is retained to serve platform numbers rather than users, and that the team already knows it. Features that genuinely serve users can survive the question.

#### Conflating early adoption with sustained value

A second common mistake is conflating novelty with genuine value. New personalisation features often produce an engagement spike on launch because users are curious. That spike is not evidence that the feature is working as intended. The signal worth watching is engagement at 30 and 60 days, when novelty has worn off and the feature is competing purely on relevance. Teams that optimise on week-one data are building on a reading that is not representative of ongoing user experience.

A third mistake is building personalisation tests without first knowing what the emotional baseline is. If you do not know how users feel before a personalisation feature is introduced, you cannot reliably attribute changes in sentiment to the feature. Research that establishes the [emotional starting point is what makes the test results interpretable](https://weareaffective.com/learning-centre/how-to-run-a-concept-test-that-doesnt-just-confirm-what-the-team-already-believe) on a personalisation project.

## Turning test findings into durable personalisation improvements

Test findings become durable improvements only when the team treats them as updates to a model of [user behaviour rather than as isolated results](https://weareaffective.com/learning-centre/what-a-development-team-actually-needs-to-know-about-the-user-before-sprint-one) to act on. A single test that shows a lift in completion rates tells you something about one flow at one moment. A pattern of tests interpreted together tells you something about how your users think, what creates friction for them, and what creates confidence. The second kind of knowledge compound. The first kind does not.

The process for turning findings into durable improvements has a specific shape.

1. Capture the behavioural signal alongside the emotional context from qualitative research
2. Identify whether the finding reflects a user expectation mismatch, a friction problem, or an emotional misfire
3. Make the change at the right layer, whether that is model inputs, presentation, timing, or language
4. Retest with the emotional arc as the frame, not just the conversion metric
5. Document the reasoning, not just the result, so the next team member can build on it rather than repeat it

On the fitness social network where we redesigned the location-sharing flow, introducing approximate proximity rather than precise location disclosure and sequencing conversation before location sharing, conversion from sign-up to successfully meeting another user rose from around 20% to approximately 60 to 70%. That was a single behavioural flow change, but it was grounded in an understanding of the emotional state users were in at that moment in the product. The model was secondary. The emotional insight drove the decision.

Improvements built on that kind of understanding hold because they are solving the right problem. Improvements built purely on optimising a metric tend to erode as the context around them shifts, because they were never grounded in why the metric was moving in the first place.

## Conclusion

Testing and optimising AI-driven personalisation is less about the sophistication of the testing apparatus and more about the quality of the questions being asked of the data. Teams that ask only whether conversions went up are answering a narrow question that misses most of what determines whether personalisation builds long-term engagement or quietly erodes it.

The work is in building feedback loops that capture emotional signals alongside behavioural ones, setting expectations before flows rather than only tracking what happens inside them, testing notification timing as a genuine variable rather than a binary, and reading behavioural data through an emotional arc that gives each data point its proper meaning.

The 21-point gap between how retailers perceive their personalisation and how consumers actually experience it, identified by [Contentful](https://www.contentful.com/blog/ecommerce-personalization-statistics/), is a measurement and interpretation gap. Teams believe their personalisation is working because their metrics say it is. Users experience something different. Closing that gap requires measuring the right things, in the right moments, with the right framework for interpretation.

The practical starting point is to review what your current testing programme actually measures, and to ask honestly whether any of it captures how users feel during the personalised experience, not just what they do at the end of it. If the answer is no, that is where the work starts.

[Let's talk about your personalisation testing approach](https://weareaffective.com/get-started)

## Frequently Asked Questions

Why is AI-driven personalisation harder to test than standard product features?

Unlike static features, AI-driven personalisation adapts continuously based on user behaviour, which means what you tested last month may not reflect what is running today. This makes standard testing approaches inadequate on their own, and means disengagement can go undetected even when your data appears clean.

What are the three main areas that need to be tested in any personalisation system?

The three areas are the quality of the input signals being used to inform personalisation decisions, the presentation layer covering how and when personalised content surfaces, and the emotional response of users to being personalised to. Each area requires a different testing approach, ranging from quantitative assessment to qualitative research methods.

Why are A/B tests not enough when testing personalisation features?

A/B tests assume a stable population and a fixed variant, but adaptive personalisation breaks both of those assumptions. The variant shifts with user behaviour, and the algorithm itself segments the audience in ways that can confound your results.

How can you tell whether personalisation is having a negative emotional impact on users?

Behavioural signals such as dwell time, speed of movement through a flow, return visit patterns, and task completion rates all carry information about how users are responding emotionally. Teams should treat these signals as indicators of emotional response, not just functional performance, from the very start of a testing programme.

What is the difference between personalisation that feels helpful and personalisation that feels intrusive?

Personalisation that works feels invisible, helping users move through a product faster and return more readily without drawing attention to itself. When it fails, it feels either intrusive or arbitrary, and users tend to disengage or leave rather than flag the problem directly.

Do you need qualitative research methods when testing personalisation, or is quantitative data sufficient?

Quantitative data alone is not sufficient because it cannot reliably capture emotional misfires or tell you whether users feel understood versus watched. Qualitative methods are essential alongside the numbers, particularly for assessing the emotional and psychological dimensions of the personalised experience.

What kinds of input signals should a personalisation system be drawing on?

A personalisation system should be reading the right behavioural data at the right moments, rather than simply collecting as much data as possible. Testing should include a check on whether those input signals are genuinely informative, as poor signal quality upstream will undermine personalisation quality regardless of how good the model is.

What is the most common reason personalisation testing fails in practice?

The most common failure is not a technical one but an interpretive one. Teams often have the data they need but are asking the wrong questions of it, measuring the wrong moments, or optimising for a signal that is only loosely connected to what users actually need.

## Related Articles

[![We Are Affective](https://weareaffective.com/hubfs/we_are_affective_logo_mark.svg)](https://weareaffective.com)

20-22 Wenlock Road  
London, N1 7GU  
United Kingdom

+44 20 4572 8062  
[hello@weareaffective.com](mailto:hello@weareaffective.com)

<https://linkedin.com/company/weareaffective> <https://instagram.com/weareaffective> <https://facebook.com/weareaffective>

Services

[App planning & strategy](https://weareaffective.com/app-planning-strategy) [App design](https://weareaffective.com/app-design-agency) [App UX design](https://weareaffective.com/app-ux-design) [App UI design](https://weareaffective.com/app-ui-design) [App technical architecture](https://weareaffective.com/app-architecture) [Existing app audits](https://weareaffective.com/app-audit)

Legal

[Privacy policy](https://app.termly.io/policy-viewer/policy.html?policyUUID=b8fa9921-7518-4fb5-8ddd-9dc7f5977ed2) [Terms](https://app.termly.io/policy-viewer/policy.html?policyUUID=8b6a6ad5-91bd-4176-a5f7-6d36b0398f70)

Case studies

[TravAI](https://weareaffective.com/case-studies/travai) [Meditech](https://weareaffective.com/case-studies/harley) [WorkingWeight](https://weareaffective.com/case-studies/workingweight) [SkinSync](https://weareaffective.com/case-studies/skinsync) [Three Lochs](https://weareaffective.com/case-studies/three-lochs) [Drift](https://weareaffective.com/case-studies/drift)

About us

[Our Story](https://weareaffective.com/about) [How We Work](https://weareaffective.com/how-we-work)

Guides

[Creating an app](https://weareaffective.com/how-to-create-an-app) [Building an MVP](https://weareaffective.com/building-an-mvp) [Cost and budgeting](https://weareaffective.com/app-development-cost) [App technology](https://weareaffective.com/app-development) [Planning and strategy](https://weareaffective.com/app-planning-strategy) [User research](https://weareaffective.com/app-user-research) [Onboarding design](https://weareaffective.com/app-onboarding-design) [User psychology](https://weareaffective.com/user-psychology-app-design) [Launch and growth](https://weareaffective.com/app-launch-growth)

 Copyright © 2026, weareaffective.com. All rights reserved.