How to Test a Product Concept With Behavioural Realism Rather Than Interview Enthusiasm
People say kind things in research sessions. They lean forward, they nod, they tell you the concept is exactly what they've been looking for. Then the product launches and nobody buys it. This gap between what people say and what people actually do is one of the oldest problems in product development, and it keeps catching teams out because the methods used to test concepts are often designed, whether intentionally or not, to attract agreement rather than reveal truth.
Concept testing, at its best, gives teams a genuine read on whether a product idea maps onto a real problem that real people experience. At its worst, it becomes a performance: participants play the role of enthusiastic early adopter, researchers play the role of neutral observer, and everyone leaves feeling good about something that has not actually been tested at all. The distinction between those two outcomes comes down almost entirely to method.
According to the Product Development and Management Association, organisations with the strongest testing programmes have a 24% product failure rate, compared to 46% for those without. That gap is large enough to matter commercially, and it points directly at the quality of testing rather than the quality of ideas. The question is not whether to test, but how to design tests that produce behavioural signal rather than polite enthusiasm.
Concept Collisions
Another artefact-based method worth using is concept collision, where you place your concept alongside competing solutions the participant already uses and ask them to explain how they would decide between them. This is not a question about your concept in isolation. It is a question about real-world choice under realistic conditions, where your product would need to displace or complement something that already exists in the person's life. The friction in that conversation is exactly the friction your product will face at the point of adoption.
What makes artefact-driven sessions harder to run than standard interviews is that the conversation is less predictable. Participants take you places you did not plan to go. That is the value of the method, because those unplanned places are often where the genuinely useful information sits.
Ask participants to bring three things they currently use to solve the problem your concept addresses. The combination of those three things tells you more about the real problem than any question you could ask directly.
Introducing Real Stakes Into Post-Session Decisions
One of the clearest signals available in concept testing is what people choose when something is actually at stake. Stated preference and revealed preference diverge sharply as soon as real consequences are introduced, and the gap between them is often larger than teams expect.
There are several practical ways to introduce real stakes without creating ethical problems. The most straightforward is a willingness-to-pay exercise that asks participants to make a real financial commitment, even if only a small one. Asking someone to put £2 on a product concept they say they love produces a very different response from asking them to rate how much they would pay on a scale of one to ten. The moment money is involved, even a trivial amount, the rational mind engages and hypothetical enthusiasm gives way to genuine preference.
Trade-Off Tasks
Trade-off tasks work on a similar principle. Present participants with a scenario where they must give something up in order to get the product, whether that is time, money, data, or an existing habit, and observe what they are and are not willing to exchange. If your concept requires people to share their location data continuously, the question is not whether they say they are comfortable with that. The question is whether they are willing to turn it on when you put the screen in front of them and ask them to do it.
This matters because app user research involving real decisions produces qualitatively different data from testing involving imagined decisions. When users are genuinely about to commit to something, even elements of the experience that looked fine in a standard session can surface as friction points significant enough to cause abandonment. The emotional state in a real decision is simply not the same as the emotional state in a hypothetical one, and the gap between those states is where many product assumptions quietly break down.
An Energy Monitoring Product in Practice
Consider how these methods work together in the context of an energy monitoring product, a tool designed to help households reduce their consumption by showing real-time usage data and offering personalised recommendations. This is a product category where stated interest is reliably high and actual adoption is reliably low, which makes it a useful illustration of the gap these methods are designed to close.
A standard concept test might present participants with a description of the product, walk them through a prototype, and ask whether they would download and use it. Most people would say yes. The concept is easy to support in the abstract: saving money and reducing environmental impact are things people want to be seen as caring about. The research session is a social context, and in social contexts, people present the version of themselves they are proud of.
What Behavioural Methods Surface
A diary study run in the weeks before concept testing, asking participants to log their energy-related decisions and behaviours, would surface something different. It would show how rarely people check their current usage data, how little most people understand their bills, and how energy sits at the very bottom of their daily attention unless something goes wrong. That context reframes the product question entirely: the challenge is not getting people to want better energy data, it is reaching them at a moment when energy is actually in their attention.
Artefact-driven sessions, where participants bring in their most recent bills and talk through what they noticed and what they ignored, would add another layer. Trade-off tasks asking people to choose between the product and keeping their current supplier, or asking them to share their smart meter data with a third party as part of signing up, would introduce real stakes and reveal where the actual friction sits. The combination produces a much more honest picture of the adoption challenge than enthusiasm in a room ever could.
Conclusion
The methods in this article are not complicated. None of them require specialist technology or long timelines. What they do require is a willingness to design research around what people do rather than what they say, and to treat the gap between those two things as the most useful information available.
Behavioural realism in concept testing is a discipline of subtraction as much as addition. It means removing the hypothetical framing that invites performance, removing the abstract questions that reward articulate optimists, and replacing them with tasks, artefacts, stakes, and observation. The result is research that feels less tidy but produces far more honest signal.
Teams that build this discipline into their testing process before a product launches are working with a genuine read of the problem. According to the Product Development and Management Association, the difference between strong and weak testing programmes corresponds to nearly halving your failure rate. That is not a marginal improvement. It is the difference between a product that finds its people and one that launches into silence.
The goal is always to understand what people will actually do, in the actual conditions of their actual lives, before the product exists. Every method in this piece is in service of that goal. If your current concept testing process produces a lot of enthusiasm and not much friction, that is almost certainly a signal about the method rather than the idea.
Let's talk about building behavioural realism into your concept testing.
Frequently Asked Questions
People naturally tend to be polite and agreeable in research settings, often playing the role of an enthusiastic early adopter without meaning to mislead anyone. The problem is that saying you like something costs nothing, whereas actually buying it involves real commitment, effort, and trade-offs that interview situations never replicate.
Concept collision involves placing your product idea alongside competing solutions the participant already uses in their daily life, then asking them to explain how they would choose between them. This mirrors the real conditions your product would face at the point of adoption, and the friction that emerges in that conversation is genuinely useful signal.
One practical approach is a willingness-to-pay exercise where participants make a small but real financial commitment, such as pledging a couple of pounds, to a concept they claim to support. Even a modest financial stake causes stated preference and revealed preference to diverge, giving researchers a much more honest read on genuine interest.
Polite enthusiasm is what participants express when they nod, agree, and say a concept is exactly what they have been looking for, without any real consequence attached to that response. Behavioural signal emerges when participants are asked to make decisions, trade-offs, or commitments that reflect how they would actually behave in the real world.
The combination of things a person already uses to address a problem reveals far more about the true nature of that problem than any direct question could. It shows researchers what the real competitive landscape looks like from the user's perspective, including workarounds and habits that might never come up in a standard interview.
Research from the Product Development and Management Association found that organisations with strong testing programmes experience a 24% product failure rate, compared to 46% for those without such programmes. That difference is large enough to have a meaningful impact on commercial outcomes, pointing clearly to the importance of testing quality rather than simply testing volume.
Artefact-driven sessions are less predictable because participants follow their own associations and take the conversation into territory the researcher did not plan for. This unpredictability is actually the method's strength, as those unplanned directions are often where the most genuinely useful information is found.
Concept testing becomes misleading when it is designed, even unintentionally, to attract agreement rather than reveal truth. When both the researcher and the participant are effectively performing their roles rather than genuinely probing the concept, the session produces confidence rather than insight, which can lead teams to invest in ideas that have not actually been tested.