How to Run a Concept Test That Doesn't Just Confirm What the Team Already Believes
Teams commission concept tests because they want to know if an idea will work. That is the honest version. The less honest version is that they commission concept tests because they want evidence it will work, and they build the test to get it. The stimuli show the concept at its best. The questions are framed to invite approval. The participants are recruited from a pool that already leans towards the category. And when the results come back positive, everyone feels good and nothing changes.
This is not deliberate deception. Most teams genuinely believe they are being rigorous. The problem sits in the small decisions made before a single participant walks through the door — the prototype fidelity, the question wording, the screener criteria. Each one introduces a small tilt. Together, they produce a test that is structurally incapable of surfacing a serious problem.
A concept test built to confirm is worse than no concept test at all. It consumes budget, creates false confidence, and gives the team permission to stop asking hard questions. The product launches into the market carrying the same flaws the research was supposed to catch.
Running a concept test that actually challenges the idea takes a different kind of deliberateness. It means designing for disconfirmation, not for validation. This piece walks through how to do that.
Why Most Concept Tests Are Built to Succeed
The pressure to validate an idea starts long before the research brief is written. By the time a concept reaches testing, a team has spent weeks or months on it. There is a business case, a roadmap slot, a stakeholder who championed it in a board meeting. The emotional investment is real, and it shapes every decision that follows.
So the prototype gets polished. The team shows the concept in its ideal form, with carefully chosen photography, clean copy, and a flow that makes the value proposition obvious. Real products do not look like this at launch, but the test version does. Participants respond to what they see, not to the rougher reality they will eventually encounter.
The questions follow the same pattern. Teams naturally frame questions around what they want to explore. "What do you like about this?" produces very different data from "What would stop you buying this?" Both are legitimate questions. Most concept tests lean heavily on the first.
There is also a documented reluctance to look closely at what the research actually says. In a State of User Research report published by Dovetail in 2022 and 2023, approximately 60% of researchers said their work was sometimes or rarely acted upon, and around 40% of those cited lack of stakeholder engagement as the primary reason. The test produces findings. The findings sit in a slide deck. The product ships more or less unchanged.
The Three Structural Biases Undermining Your Results
Most concept test problems are not about attitude. Teams do not consciously decide to ignore bad news. The bias is structural — it gets baked into how the test is designed, and it produces distorted results regardless of how sincerely everyone wants to learn.
Confirmation Through Stimulus Quality
A high-fidelity prototype tells participants how much effort the team has put in. Without meaning to, it signals that the concept is already nearly ready. This raises the social cost of negative feedback and suppresses the kind of honest reaction that lower-fidelity stimuli tend to produce. When everything looks polished, participants search for things to endorse rather than problems to name.
Question Order as a Nudge
Starting with "what do you like?" primes participants before they encounter harder questions. Their earlier positive answers become an anchor. When a question later asks about concerns, participants soften their responses to stay consistent with what they said before. The sequence itself has introduced a bias, and the data carries it all the way through to the debrief.
The third bias is participant selection. Teams often recruit people who already engage with the category, partly because they are easier to find and partly because they seem like natural prospects. But heavy category users bring sophisticated priors. They already know what to expect, they evaluate the concept on informed criteria, and they are more likely to give the team the kind of feedback that sounds useful but reflects a minority view.
Design that understands your users
We build app experiences around real user behaviour, not assumptions. Research, psychology-driven design and technical specs that turn users into loyal advocates.
Designing Stimuli That Don't Whisper the Answer
The stimuli presented in a concept test communicate more than the concept itself. They communicate how much the team believes in it, how far along it is, and — subtly — what kind of response the team is hoping for. Getting stimuli right means stripping out those signals before a single participant sees the work.
One principle worth holding to is deliberate roughness. When stimuli look unfinished, participants feel comfortable saying so. They treat the concept as genuinely in development, which is the frame you want, because it gives them permission to be direct. A sketch-level wireframe or a rough storyboard tends to produce more honest reactions than a pixel-perfect prototype, and the data is correspondingly more useful.
Stimuli should also avoid loading the concept with its own best arguments. Landing pages and polished decks often front-load benefits and social proof. In a concept test, that kind of pre-persuasion contaminates the results. Present the core idea as neutrally as possible, describe what the product does, and let the participant form their own view of whether that matters to them.
A practical approach is to test multiple variants, including at least one that omits features the team considers central. If enthusiasm drops sharply when a particular element is removed, that is informative. If enthusiasm stays broadly the same, that is also informative, and the team has learned something about what is actually driving appeal.
Use rough wireframes or paper sketches rather than polished prototypes to reduce the social pressure on participants to give approving feedback.
Writing Questions That Welcome Bad News
Good concept test questions assume the participant has reservations and create space for those reservations to surface. Bad questions assume the participant is interested and invite them to explain why. The wording of a question determines the kind of data it can return, and most teams write questions that cannot receive the answers they most need to hear.
Start by replacing approval questions with usage questions. "What do you think of this?" invites a verdict. "Walk me through what you would actually do with this" produces behaviour. People can express polite enthusiasm for a concept while simultaneously having no intention of using it, and a usage-framed question makes that contradiction visible in a way that an opinion question never will.
Probing Inaction Directly
Ask explicitly about barriers. Questions like "what would make you hesitate before signing up?" or "what would you want to know before spending money on this?" invite the participant to voice genuine concerns rather than suppress them. The key is framing these as expected and normal — participants hold back if they sense their doubts are unwelcome.
Avoid questions that produce consensus. "How likely are you to recommend this to a friend?" is a well-worn measure, but in a concept test context it tends to produce inflated numbers because recommending is a lower-stakes commitment than actually buying. If you use likelihood-to-purchase questions, ask the participant to explain their score immediately, and listen for the hesitation under a positive number.
Ordering matters as much as wording. Open questions about concerns should come before questions about benefits. Once a participant has committed to articulating what they liked, they have already shaped the frame for everything that follows.
Ask "what would stop you using this?" before you ask "what appeals to you about this?" The sequence changes what participants feel comfortable saying.
Recruiting Participants Who Might Actually Say No
A concept test is only as honest as the people in it. Recruit participants who already lean towards buying the category and you will get data from people predisposed to find reasons to approve. That is a sample that makes concepts look better than they are.
The useful participants are the ones sitting on the edge of the target demographic — people for whom the category is relevant but not habitual, people who have looked at similar products and walked away, and people whose relationship with the problem the concept solves is ambivalent rather than active. These participants bring real friction. They have actual reasons for not engaging, and a concept test is the right place to hear those reasons.
Screener criteria tend to filter these participants out, often without meaning to. A screener that selects for regular category users, active interest in the problem, or prior purchase behaviour will reliably exclude anyone who bounced before buying. Including a segment of bouncers or considerers alongside active users produces a much more informative spread of responses.
- Include people who considered similar products and chose not to buy
- Recruit across a range of engagement levels, not just heavy users
- Avoid over-screening for prior brand familiarity
- Check that your screener does not inadvertently select for enthusiasm about the category
Weighting feedback by participant profile at the analysis stage is also worth doing deliberately. Not every participant's view carries equal relevance to every feature, and reviewing who said what — their background, their relationship with the problem, their prior behaviour — changes how the findings should be read and reported.
Review screener criteria against your target demographic before fieldwork begins. If every criterion selects for enthusiasm, the test cannot surface genuine scepticism.
A Worked Example: Testing a Wellbeing Subscription Concept
Consider a team developing a subscription product in the personal wellbeing space. The concept centres on a personalised programme that tracks habits, delivers guided content, and generates weekly progress data for the user. The team has strong internal belief in the data layer — the science behind the tracking is solid and the stakeholders who know it best want it front and centre in the experience.
A standard concept test, run in the usual way, would likely produce encouraging results. Participants in the target demographic tend to respond positively to the idea of tracking and progress. The prototype, shown at a polished fidelity, looks trustworthy. Questions about interest and likelihood to subscribe return reasonable scores. The team feels validated.
Where the Standard Test Fails
The problem surfaces when the stimuli change. When the same concept is presented with lower-fidelity screens, without the polished visual framing, and questions are reordered to ask about hesitations before benefits, a different picture emerges. Participants find the data presentation early in the journey overwhelming. For people who are not already familiar with the science behind the tracking, seeing results without context raises anxiety rather than reassurance. The emotional experience of encountering the concept shifts from encouraging to stressful.
This is the kind of finding that a confirming test cannot return. The concept is not wrong, but the sequencing is, and the team's own expertise in the science makes that very hard to see from the inside. Testing with participants who have little prior knowledge of the category, presenting stimuli at a rougher fidelity, and asking about emotional responses before asking about intent produces the data the team actually needs to make good decisions.
Conclusion
Running a concept test that challenges the idea rather than confirming it requires decisions made before the research starts. Stimulus fidelity, question order, participant screeners — each one either opens the door to honest feedback or closes it. Most concept tests close it, and the product goes to market carrying the same unexamined assumptions the research was supposed to surface.
The goal is not to find reasons to kill a concept. A well-designed concept test serves the team's actual interests far better than a flattering one does, because it produces findings the team can act on. A positive result from a rigorous test carries real weight. A positive result from a test that could only ever return positive results carries none.
Research that surfaces problems early is cheaper than the same problems discovered at launch. It is also more useful to the people building the product, because it leaves room to respond. The team that builds a concept test designed to welcome bad news is the one that ends up with a better product, not because they were more cautious, but because they were more honest with themselves about what they did not yet know.
If you are building a concept test and want a second opinion on how it is structured, let's talk about your research design.
Frequently Asked Questions
A concept test is a piece of user research designed to evaluate whether a product idea will work before it is fully built or launched. Teams use them to gather evidence about whether an idea resonates with their target audience, though the article notes that in practice many teams unconsciously use them to confirm decisions they have already made.
The bias is largely structural rather than intentional — it gets built into small decisions made before research begins, such as using polished prototypes, positively framed questions, and participant screeners that favour enthusiastic respondents. By the time a concept reaches testing, teams have often invested significant time and emotional energy into it, which subtly shapes every design choice that follows.
A concept test designed to confirm is argued to be worse than running no test at all, because it consumes budget and creates false confidence whilst giving the team permission to stop asking difficult questions. The product then launches carrying the same flaws the research was supposed to identify.
A highly polished prototype signals to participants that the concept is nearly finished, which raises the social cost of giving negative feedback and suppresses honest reactions. Lower-fidelity stimuli tend to produce more candid responses because participants feel less pressure to be encouraging about something that appears unfinished.
Designing for disconfirmation means deliberately structuring a concept test to surface problems, weaknesses, and reasons the idea might fail, rather than seeking approval. It matters because a test that is only capable of producing positive findings cannot be trusted, regardless of what results it returns.
The way questions are framed significantly shapes the data collected — for example, asking 'What do you like about this?' will produce very different responses from 'What would stop you buying this?' Most concept tests lean heavily towards positively framed questions, which naturally skews results in favour of the concept being tested.
According to a State of User Research report published by Dovetail in 2022 and 2023, approximately 60% of researchers said their work was sometimes or rarely acted upon. Around 40% of those cited a lack of stakeholder engagement as the primary reason findings were not used.
The article identifies confirmation through stimulus quality, biased question framing, and flawed participant recruitment as the three key structural biases. Each introduces a small tilt towards positive results, and together they can produce a test that is incapable of surfacing a serious problem with the concept.
Related Articles
How does anchoring shape the way people perceive price?
The number 99 appears everywhere. Shops price items at £9.99 rather than £10. Software...
How Much Does It Cost to Build a Mobile App?
One of the most asked questions we get at We Are Affective is "How much it will cost to develop my...