Why Do Users Say One Thing but Do Something Completely Different?
Ask users whether they trust a product, and most will say yes. Ask whether they found the experience easy, and they will rate it highly. Ask whether they would recommend it, and they will nod. Then watch what they actually do: they hesitate at the checkout, they abandon mid-flow, they never return after the first session. This gap between what people say and what they do is one of the oldest problems in product design, and it is still causing teams to build the wrong things with a great deal of confidence.
People are feeling okay with stuff until they're not, the emotional state changes, and the trust has to be even higher there.
The problem runs deeper than dishonesty. Users are not lying. They genuinely believe what they report. The issue is that self-reported satisfaction is collected in the wrong conditions, at the wrong moment, from a version of the user who is calm, detached, and not actually doing anything at stake. That person has a different psychology to the one who is about to share their location, enter their card details, or commit to a subscription. "People are feeling okay with stuff until they're not, " as Simon Lee at We Are Affective puts it. The transition from observer to participant changes everything.
What follows is an exploration of why this gap exists, what sustains it, and how to design research and products that account for the user who actually shows up rather than the idealised one who exists in a survey response.
The Gap Is Real: What the Data on Self-Reported Satisfaction Actually Shows
Self-reported satisfaction scores are not useless, but they carry a structural flaw. NPS, CSAT, and similar measures are collected after the fact, when users have stepped back from the product and are reflecting on it in a calm, analytical state. That is a fundamentally different mental mode to the one they were in while actually using it.
The correlation between self-reported satisfaction scores and actual behaviours like retention and conversion is consistently modest, meaning survey results alone are a poor predictor of what customers will actually do. That is a weak to moderate relationship at best. It means a product can score well on user satisfaction and still haemorrhage users at key points. McKnight's research on online trust found something similar: stated trust scores diverge substantially from actual willingness to transact or share data, particularly when friction or perceived risk increases at the moment of commitment.
Part of the reason these scores persist in product teams is organisational, not methodological. Positive scores travel well up through a company. They reflect well in board reports. A number that looks good creates pressure to keep using the measure that produced it, regardless of whether it is telling the full story. The gap between reported satisfaction and real behaviour does not appear in the slide deck, so it does not get addressed.
To genuinely understand how a product is perceived, both consciously and subconsciously, teams need real-time, in-context behavioural data alongside the scores. The scores alone are a partial picture of a partial moment.
Why Testing Conditions Produce Cleaner Answers Than Real Life
Usability testing has a setting problem. When a participant sits down in a session to review a checkout flow or a data-sharing consent screen, they are reviewing a simulation: no money changes hands, no location is revealed to a stranger. They are evaluating the design from a position of safety, which produces a fundamentally cleaner reading than the real world gives you.
In a test environment, users assess flows functionally and rationally. They notice whether a button label is clear, whether the layout makes sense, whether the copy reads naturally. They are not flooded with the anxiety that arrives when a payment is genuinely pending, or the wariness that accompanies a request for personal data from a service they have known for forty seconds.
According to AppTweak's ASO Blog, when users are asked whether they would use an app, between 60 and 80 percent typically respond positively, but actual usage is often only 10 to 20 percent. That gap is partly a function of testing conditions producing cleaner, more optimistic answers than real-world conditions sustain.
The fix is to be precise about what usability testing can tell you, and where its reach ends. Functional clarity, information architecture, and layout decisions are well-suited to a controlled setting. Trust, commitment behaviour, and the micro-moments where users decide to drop off require live data from the actual product, in the actual moment, with actual stakes.
UX/UI design built around real psychology
We design app interfaces around how people actually think and behave. User research, psychology-driven UX/UI design and technical specs delivered as one complete package.
The Role of Emotional State at the Moment of Decision
Emotional state at the moment of use is not background noise. It shapes the decision in front of the user more directly than the design of the interface. A user who arrives at a product already anxious processes information differently, tolerates uncertainty less, and reaches drop-off thresholds faster than a calm one. The design might be identical. The outcome is not.
We worked on a teardown of a meditation app that illustrated this precisely. The app promised calm in under a minute, then immediately asked new users to define their goals. The assumption baked into that design was that someone opening a meditation app for the first time was arriving calm, curious, and ready to engage rationally with onboarding questions. The reality is that most first-time users open a meditation app because they are anxious and need help. They are not in a state to consider their long-term goals. They need immediate relief, not a goal-setting exercise.
Nobody opens a meditation app for the first time because they're already calm and curious, they open it because they really need help.
This is a pattern we see reflected across many products. The design is built around the ideal user arriving in an ideal state, and the actual user, arriving with actual emotions, is served something that does not fit. The gap between stated satisfaction in research and real-world behaviour often traces back to exactly this mismatch: the design was tested on a calm participant, but it is used by someone in a very different place.
Before designing an onboarding flow, map the emotional state your user is most likely in when they first arrive. Build from that state, not from a hypothetical calm one.
When Trust Collapses at the Point of Commitment
Trust in a digital product is not a fixed property. It shifts at specific moments, and those moments are almost always the ones where something real is being asked. A user can move through a product feeling broadly positive, rating it well if asked, and then reach the point where they are asked to share their precise location, enter payment details, or grant access to personal data, and the whole emotional picture changes.
We worked on a map-based fitness social network designed to connect people for runs and cycle rides. Users were dropping off at a specific point in the flow: the moment they were asked to share their precise location with a potential match. The drop-off was not because the location-sharing mechanism was confusing or the design was unclear. It was because the request came too early. Users had not yet had any opportunity to converse with or learn about the other person. They were being asked for a high-trust action before trust had been given time to form.
Moving the location-sharing prompt to after an initial exchange changed the context of the request entirely. The action was the same. The moment was different. That difference was the thing users could not have articulated in a survey, because in a survey they were not actually being asked to share their location with a stranger. The emotional stakes were absent, so the problem was invisible.
This is the pattern: a product looks fine in testing because the participant is not genuinely at stake, and then live analytics reveal drop-off at exactly the points where the emotional cost of commitment is highest.
Why Guidance Can Feel Like Friction to the Wrong User
The assumption that guidance reduces anxiety and abandonment is reasonable for most contexts. But it fails for specific user types, and building without accounting for those users produces something that works against them.
We worked on a product in the financial trading space. The product had technically strong interaction design, with contextual guidance, layered information, and clear pathways through the flow. A more established competitor in the same space had a denser, more bare interface with fewer affordances. The expectation would be that the guided product felt easier. For experienced commodity traders, it felt worse.
Experienced traders need to reach time-sensitive trades quickly. Every additional step, every contextual tooltip, every progressive disclosure mechanism is a delay they did not ask for. The bare competitor interface felt more intuitive to them because it mapped directly to how they already thought. Stripping away guidance was not creating confusion for this audience. Adding it was creating friction.
Test your product with the users who will actually use it at the level of expertise they will actually bring. Novice-facing design assumptions embedded in expert-facing products create friction that does not show up until live use.
In most products, confusion leads to anxiety and anxiety leads to abandonment. That relationship holds until you are designing for an expert audience with a specific task and a time constraint, at which point the inverse is true. Understanding the use case and the actual user, rather than a generalised one, is what separates a design that works in testing from one that works in the world.
Context Cannot Be Separated From Behaviour
The same design decision can produce completely different behaviour depending on the context in which a user encounters it. This is not a minor variable. Context, including the user's situation, their stress level, their familiarity with the task, and the stakes they perceive, shapes behaviour more reliably than the interface does in isolation.
We worked on a concierge product for people moving into properties with furniture deliveries. Understanding the emotional state of the user in that moment did not require any intrusive monitoring or sentiment analysis. The context itself told us what we needed to know: someone managing a moving day with deliveries arriving is inherently under pressure. That stress is a given, not a hypothesis. Designing for that emotional state meant the product could respond to it without needing to ask.
Context as Signal
Context works as a signal in both directions. A user browsing a fitness app on a Sunday afternoon is in a different state to the same user opening it before a Monday morning session they are not motivated to do. The product, the user, and the interface are identical. The context is not, and it changes what the design needs to do.
Behavioural patterns within a product give strong indicators of emotional state without requiring users to report anything. Dwell time, speed of movement through a product, return visit patterns, and whether users struggle with the same thing repeatedly or move fluidly across different tasks all provide real-time signals that self-reported data cannot. The user who lingers without progressing is telling you something their survey response would not.
Situation Over Self-Report
When teams design for a decontextualised average user rather than the actual person in a specific situation, they build something that works in no context particularly well. Designing for the situation means starting with what the user is carrying before they even open the product.
What Observed Behaviour Reveals That Users Cannot Tell You
Users cannot accurately report their own behaviour for a simple reason: they are not in the moment when they are reporting. Memory flattens emotional texture. Retrospective accounts of a checkout experience, a sign-up flow, or a confusing navigation path are reconstructions, and they leave out the hesitation, the micro-moment of distrust, and the friction that nearly caused abandonment.
Session length is a good example of what observed data can and cannot tell you on its own. A long session looks like a positive signal. It could indicate that the user is genuinely finding value and resonating with the product. It could equally mean the product is confusing and they cannot find what they came for. Or it could mean the gamification mechanisms are doing their job of keeping the user present regardless of whether they are getting anything from it. The number is the same across all three explanations.
What separates these interpretations is the pattern of behaviour within the session. A user who completes different tasks across a long session is likely exploring with intent. A user who attempts the same action repeatedly and fails, or who drifts across the product without completing anything, is telling a different story. Neither of those users could give you the real account in a post-session survey, because the frustration and the confusion were in the moment, not in the memory.
Pair session length data with task completion patterns and failure loops. A long session with repeated failures at the same point is a problem wearing the mask of engagement.
How to Design Research That Closes the Gap
Closing the gap between stated and actual behaviour requires using research methods that match the conditions being studied. Self-reported satisfaction belongs in the toolkit, but it needs live behavioural data alongside it, collected from users who are genuinely at stake rather than evaluating from a safe distance.
Match the Method to the Question
The research method needs to match what is being investigated. For functional clarity and layout decisions, controlled testing is well-suited. For understanding trust, commitment behaviour, and emotional responses at high-stakes moments, live product analytics from real users in real situations give you information that no test session can replicate.
A useful starting structure for closing the gap looks like this:
- Map the moments in the product where real commitment is required: payment, data sharing, location access, subscription confirmation.
- Instrument those specific moments with behavioural tracking: hesitation patterns, back-navigation, and time-on-step alongside completion rates.
- Collect self-reported feedback separately, and compare it against the behavioural data at those same moments.
- Where the two diverge, treat the behavioural data as the more reliable signal.
- Iterate on the moments of divergence first, because that is where the gap between perceived and actual experience is largest.
Design for the Arriving User
Research needs to include the actual user arriving in their actual emotional state, not a participant who has been briefed, seated, and asked to evaluate. Diary studies, experience sampling during real use, and analytics from live products all capture what a lab session cannot: the person who opened the app on a bad day, who hesitated at the payment screen because something in the copy felt off, and who left without being able to say exactly why.
Conclusion
The gap between what users say and what they do is a structural feature of how emotional state interacts with decision-making, and it means the conditions under which you collect data matter as much as the data itself.
A user who rates their experience positively in a post-session questionnaire is telling you about a calm, retrospective version of themselves. The user who hesitated at the checkout, who dropped off when asked for their location before building any trust, who found guidance patronising because they were an expert with a time-sensitive task, that user is visible only in the behavioural data produced during actual use.
The meditation app assumed calm, curious users and built an onboarding flow for them. The fitness network asked for location before trust had formed. The trading product added guidance that expert users experienced as obstruction. In each case, the design was built for an idealised user rather than the actual one arriving at the product in a specific emotional state, with specific needs and a specific level of patience.
Getting this right means pairing self-reported measures with real-time behavioural data, treating context as a core design input rather than background noise, and researching the moments of highest commitment with the seriousness they deserve. If that sounds like a shift from how your product team currently works, let's talk about your research approach.
Frequently Asked Questions
Self-reported satisfaction is collected after the fact, when users are calm and detached from the experience. This reflective state is psychologically different from the one users are in when they are actually making decisions under pressure, such as entering card details or sharing personal data.
No, users genuinely believe what they report at the time of answering. The problem is that the person completing a survey is not the same psychological version of themselves as the one navigating a high-stakes moment in a product.
These measures are collected after users have stepped away from the product, meaning they reflect a calm, analytical mindset rather than the emotional state present during actual use. Research consistently shows only a modest correlation between satisfaction scores and real behaviours like retention and conversion.
Positive scores travel well within organisations and reflect favourably in board reports, which creates internal pressure to keep using the measures that produced them. Because the gap between reported satisfaction and real behaviour rarely appears in a slide deck, it tends not to get addressed.
In a test environment, no real money changes hands and no actual personal data is shared, so participants evaluate designs from a position of safety. This produces more positive and more considered responses than users would give when genuinely at risk of losing something or exposing sensitive information.
The gap is most pronounced at moments of commitment, such as sharing a location, entering payment details, or signing up for a subscription. Users may feel comfortable with a product right up until that point, at which stage their emotional state and level of required trust shift considerably.
Teams need real-time, in-context behavioural data collected alongside self-reported scores rather than instead of them. This combination gives a more complete picture by capturing how users actually respond in the moment, not just how they recall or rationalise the experience afterwards.
Research on online trust shows that stated trust scores can diverge substantially from actual willingness to transact or share data, particularly when friction or perceived risk increases. Users may report trusting a product in the abstract but hesitate or drop off entirely when a genuine commitment is required.