Skip to content
Expert Guide Series

How to Run a Behavioural Kill-Criteria Workshop Before a Build Kicks Off

Most builds fail before a single line of code is written. The decisions that doom a product tend to happen in the weeks before development begins, when teams are still excited, still aligned, and still certain the thing they are about to build is the right thing to build. The trouble is that certainty and clarity are very different things, and confusing the two is how teams end up six months in, halfway through a build they should have stopped three months ago.

A behavioural kill-criteria workshop changes that. It forces a team to agree, while everyone is still rational and unattached, on exactly what would need to be true for them to stop. Not what would need to be wrong, but what specific, observable behaviours from real users would tell them the direction is not working. Done well, it is one of the most protective things a team can do before committing serious time and money to a build.

The process draws on behavioural psychology rather than gut feeling. Human beings are poor at calling things off once they have started. We experience the sunk cost effect, we rationalise failure, and we convince ourselves that the next sprint will fix what the last five did not. Kill criteria exist to protect teams from their own optimism. They are agreed in advance, written in observable terms, and held to even when the pressure to carry on is high.

Teams that agree on observable exit conditions before a build starts protect themselves from their own sunk cost thinking.

This article walks through how to run that workshop: who to invite, how to frame the problem, how to write criteria that actually hold under pressure, and how to defend them when development momentum makes stopping feel impossible.

Why Builds Fail Without Exit Conditions

Teams rarely enter a build expecting it to fail. They enter it expecting to learn as they go, to iterate based on feedback, and to course-correct if things go wrong. The problem is that without pre-agreed exit conditions, "going wrong" is almost impossible to define in the moment. Every struggling metric has an explanation. Every piece of negative user behaviour has a rationalisation. Teams default to carrying on because carrying on always feels more defensible than stopping.

This is not a character flaw. It is a predictable feature of how human decision-making works under pressure and attachment. Once a team has invested time, built relationships, and committed publicly to a direction, the psychological cost of reversing that commitment rises sharply. Research into loss aversion tells us that the pain of losing what we have built feels roughly twice as heavy as the pleasure of gaining something equivalent. Teams do not just weigh the evidence, they weigh the evidence through the lens of what stopping would cost them personally.

Exit conditions written in advance sit outside that distortion. They are agreed when no one has anything at stake, when the build has not started and the ego has not yet attached itself to a particular outcome. A criterion that says "if fewer than 30% of users complete the onboarding flow within their first session after three weeks of live testing, we stop and reassess" is not a judgment in the moment. It is an earlier, clearer version of the team speaking to a later, more compromised version of itself.

Without that mechanism, teams tend to keep building long past the point where the evidence justified stopping. The costs of that are not just financial. They include the trust of users who encounter something that was not ready, and the morale of a team that senses the problem but cannot name it.

Who Belongs in the Room

The composition of a kill-criteria workshop matters more than almost any other variable. Get the wrong people in the room and the criteria produced will either be too safe to be useful or too politically shaped to be honest. The goal is to get the right people in the room at the right moment in the process.

The right people share a few characteristics. They are genuinely open to the research and the evidence. They care about the product improving rather than about being right. They have enough context to understand what the product is trying to do but have not yet formed such a fixed view of how it should work that evidence to the contrary will not reach them. People who have already decided what they want to build will not engage honestly with the question of what would make them stop. Including them too early, before there is enough evidence to work with, means their assumptions will fill the space where criteria should be.

In practical terms, a well-balanced workshop room for this kind of session tends to include a product owner or lead, a UX or behavioural designer, someone with access to analytics and data, and at least one person who has sat with real users recently, whether through research sessions, support tickets, or direct contact. If a specialist partner like WAA is running the workshop, they bring the framework and the facilitation. The client team brings the context and the decision-making authority.

Keep the group to six people or fewer. More than that and competing opinions start to dilute the criteria into vague, unchallenging language that no one will ever actually use to stop anything.

The one role that is sometimes missing but should always be present is someone who can say no. Kill criteria without executive backing are decorative. Someone in the room needs the authority to act on them.

UX/UI design built around real psychology

We design app interfaces around how people actually think and behave. User research, psychology-driven UX/UI design and technical specs delivered as one complete package.

See how we work Get started

No commitment

Framing the Right Problem Before Writing Any Criteria

The single most common mistake in a kill-criteria session is jumping straight to writing criteria before the team has properly defined what the product is supposed to do for the person using it. Teams tend to describe their product in terms of its features and functionality. That framing produces criteria based on feature adoption, which is the wrong unit of measurement. Features are means. The user's goal is the end. Kill criteria should be written around the end.

Before any criteria are written, the workshop needs to spend time answering a small set of behavioural questions. What is the user trying to accomplish? What does success look, feel, and behave like for that person? What would change in their daily life or routine if this product worked exactly as intended? And, critically, what would their behaviour tell you if it was not working, without them having to say so?

That last question is the one that matters most, and it is the hardest to answer well. Users rarely tell you directly that a product has failed them. They exit quietly. They abandon flows at specific points. They return to screens they already visited because they did not understand what they were being asked. They take longer than they should on decisions that should feel straightforward. These are the behavioural signals that belong in kill criteria, and they can only be identified if the team has first agreed on what "working" actually looks like in behavioural terms.

Users rarely say a product has failed them. They exit quietly and leave traces in the data behind them.

A useful exercise here is to ask the team to describe what a user would be doing differently in their life three months after using the product successfully. Anchoring to that outcome, rather than to the product's features, produces far sharper criteria.

Ask the team to describe three specific user behaviours they would expect to see if the product were working well. Then ask what the opposite of each of those looks like in the data. Those opposites become the foundation of your kill criteria.

Writing Observable Kill Criteria

A kill criterion only works if it is observable. That means it must be expressed in terms of something the team can actually measure, without interpretation or debate about what the number means. "Users are not engaging" is not a kill criterion. "Fewer than 25% of users return to the product within seven days of their first session, measured over a four-week period" is a kill criterion. The first invites argument. The second does not.

Observable criteria tend to cluster around a few categories of behavioural signal. Completion rates through key flows tell you whether users are reaching the outcomes the product was designed to help them reach. Return behaviour tells you whether the product is delivering enough value to warrant a second visit. Time-on-screen at specific decision points tells you whether users are confused, hesitant, or anxious about what they are being asked to do. Error rates and repeated actions, such as going back and re-entering information, tell you about comprehension failures or trust gaps.

Setting Thresholds That Hold

The number attached to a criterion matters as much as what it measures. Thresholds set too loosely will never be triggered even by a clearly failing product. Thresholds set too tightly will trigger on normal early-stage variance and create unnecessary panic. The best approach is to ground thresholds in one of three places: comparable benchmarks from the sector the product operates in, the team's own data from earlier iterations or related products, or a reasoned first-principles estimate of what the product would need to achieve to be commercially viable.

  • Write each criterion as a complete sentence that names the metric, the threshold, and the measurement window.
  • Agree in advance on the data source that will be used to measure each criterion, so there is no room for dispute about which number applies.
  • Include both leading indicators, which signal early trouble, and lagging indicators, which confirm it.
  • Limit the total number of kill criteria to between five and eight. More than that and the team will lose track of what they are actually watching.

Avoiding Vanity Measurements

Standard engagement figures such as session length and monthly active users tell you very little about whether a product is genuinely serving people. A user who spends a long time in a flow is not necessarily engaged. They may be confused. Writing kill criteria around metrics that can flatter a failing product builds in a bias toward carrying on. The criteria that matter are the ones tied to the user's actual goal, not to the product's activity figures.

Running the Two-Hour Workshop

Two hours is enough time to run this workshop well, provided the session is structured tightly and the facilitator holds the group to the task. The temptation in these sessions is to drift into broader product strategy conversations, which are worth having but belong in a different meeting. The facilitator's job is to keep pulling the group back to the specific question of what observable evidence would cause them to stop.

A workable structure for the two hours looks like this. The first 20 minutes are spent on context-setting: a brief description of what the product is trying to do, for whom, and what the team currently believes success looks like. The next 30 minutes are spent on the behavioural framing exercise described in the previous section, working through what the user is trying to accomplish and what their behaviour would look like if the product were and were not working. The middle 40 minutes are the core writing exercise, where the group works through each behavioural signal category and drafts specific, observable criteria with thresholds and measurement windows. The final 30 minutes are for pressure-testing: the facilitator challenges each criterion in turn, asking whether it is genuinely observable, whether the threshold is realistic, and whether the group would actually act on it if it were triggered.

Run the pressure-test round with a simple question for each criterion: "If this threshold were hit at week six of the build, would we actually stop?" If anyone hesitates, the criterion needs to be sharpened or the threshold needs to be renegotiated before anyone leaves the room.

At the end of the session, the criteria are documented in a single reference sheet that every person in the room signs off on. That document does not live in a shared folder that no one looks at again. It is reviewed at each sprint review and referred to explicitly at any point where the team is discussing whether to continue a direction.

Holding the Line When Development Pressure Mounts

The real test of kill criteria is not whether they get written. It is whether they get applied when the pressure to ignore them is highest. Development builds momentum. Teams invest in code, in design, in stakeholder relationships, and in the narrative they have been telling about the product. When a kill criterion is triggered mid-build, all of that investment becomes a reason to argue for continuing rather than a reason to reflect carefully.

The most common challenge to kill criteria under pressure is not direct rejection. It is reframing. "The metric is low because we have not marketed it yet." "The drop-off is expected at this stage." "The sample size is too small to be conclusive." Each of these statements may occasionally be true. But they are also exactly the kind of rationalisations that kill criteria were designed to protect against. If the group accepts them without scrutiny, the criteria cease to function.

Protecting against this requires two things. First, the criteria need to have been written with enough specificity that reframing them is genuinely difficult. A criterion that names the data source, the measurement window, and the exact threshold leaves very little room for moving the goalposts without everyone in the room noticing. Second, the group needs at least one person whose role in each review is to represent the criteria rather than the build. That person asks the same question each time: does the data meet the threshold we agreed, and if so, what do we do now?

Stopping a build is not failure. A team that stops based on clear evidence has done something genuinely disciplined, and that discipline is more valuable in the long run than the feature they did not ship.

Conclusion

Kill criteria are one of the more practical things a team can bring to a build, and one of the least commonly used. Most teams rely on intuition, on internal pressure, and on the hope that problems will resolve themselves if they keep going. Those are not unreasonable instincts, but they are not a system. A behavioural kill-criteria workshop, rooted in solid app planning and strategy, gives teams a system.

The process does not require a large investment of time. A well-run two-hour session, with the right people in the room and a clear facilitation structure, produces something genuinely useful: a small set of specific, observable criteria that the team has already agreed to honour. That agreement, made before anyone has anything at stake, is what makes the criteria work. It removes the negotiation from the moment when emotions and sunk costs are highest.

The behavioural framing is what separates good kill criteria from checkbox criteria. When teams anchor their exit conditions to real user behaviour rather than to feature adoption figures or engagement counts, the criteria become much harder to argue away. A user who abandons a flow, returns to a screen repeatedly, or never comes back after their first session is telling you something concrete. The criteria just give you permission to listen.

Teams that build this habit early, before the first sprint begins, tend to make better decisions throughout the build. They are clearer about what they are watching for, more honest with themselves when the signals are not encouraging, and faster to act when the evidence points in a direction they did not expect. If you want to bring this kind of thinking to your next build, let's talk about your pre-build process.

Frequently Asked Questions

What is a behavioural kill-criteria workshop?

A behavioural kill-criteria workshop is a structured session held before a build begins, in which a team agrees on the specific, observable user behaviours that would indicate a project is not working and should be stopped. Rather than relying on gut feeling or vague concerns, it produces written exit conditions based on real user behaviour. The goal is to make the decision to stop rational and pre-agreed rather than emotional and reactive.

Why should the workshop happen before the build starts?

Holding the workshop before development begins means participants are not yet emotionally or professionally attached to a particular outcome, making it far easier to agree on honest exit conditions. Once a build is underway, psychological biases such as sunk cost thinking make it increasingly difficult to stop, even when the evidence suggests the team should. Pre-agreed criteria act as a message from a clearer-headed team to a future, more compromised version of itself.

What are kill criteria, and how are they different from general project risks?

Kill criteria are specific, observable, pre-agreed conditions that would trigger a team to stop or reassess a build, written in measurable terms rather than subjective ones. Unlike general project risks, which are often broad and open to interpretation, kill criteria must describe actual user behaviours that can be tracked and verified. For example, a criterion might state that if fewer than 30% of users complete onboarding within their first session after three weeks of testing, the team stops and reassesses.

What is sunk cost thinking, and why is it a problem for product teams?

Sunk cost thinking is the tendency to continue investing in something simply because of the time, money, or effort already spent on it, rather than because the evidence supports carrying on. For product teams, this means continuing to build long after the signals suggest they should stop, because stopping feels like an admission of failure. Kill criteria help protect teams from this bias by establishing exit conditions before any investment has been made.

Who should be invited to a behavioural kill-criteria workshop?

The article indicates that the workshop covers who to invite as part of its practical guidance, suggesting that the right participants are those who have decision-making authority over the build and a stake in its direction. Including people from across disciplines, product, design, engineering, and business leadership, helps ensure the criteria are realistic, measurable, and genuinely binding. Keeping the group focused and avoiding too many voices helps the session remain productive.

How do you write kill criteria that will actually hold under pressure?

Effective kill criteria must be written in observable, measurable terms that refer to specific user behaviours rather than internal metrics or team sentiment. They should be unambiguous enough that no one can reasonably argue the criterion has not been met when the data says otherwise. The article notes that the workshop specifically addresses how to write criteria robust enough to withstand the pressure that builds when development momentum makes stopping feel costly.

What happens if a team ignores its kill criteria once a build is under way?

Ignoring pre-agreed kill criteria exposes a team to the full force of the biases the workshop was designed to prevent, including loss aversion, rationalisation, and sunk cost thinking. The financial costs of continuing a failing build are significant, but the article also highlights broader consequences including damaged user trust and reduced team morale. The value of kill criteria depends entirely on the team's willingness to honour them even when carrying on feels more comfortable.

Is this approach based on established psychological principles?

Yes, the kill-criteria process draws explicitly on behavioural psychology, particularly research into loss aversion, sunk cost effects, and how human decision-making becomes distorted under pressure and emotional attachment. The article references findings suggesting that the pain of losing something already built feels roughly twice as heavy as the pleasure of an equivalent gain, which explains why teams default to continuing rather than stopping. Structuring exit conditions around observable behaviour rather than subjective judgement is designed to counteract these well-documented cognitive tendencies.