Skip to content
Expert Guide Series

What a Behavioural Success Criterion Looks Like and Why Acceptance Criteria Aren't Enough

Most product teams can tell you whether a feature was built correctly. Fewer can tell you whether it worked on the person using it. Acceptance criteria answer the first question well. They confirm that a button is in the right place, that a form submits without errors, and that the right data appears on screen. What they cannot tell you is whether a user felt confident pressing that button, understood why the form was asking what it was asking, or left the screen feeling informed rather than overwhelmed.

This gap matters more than most teams realise. A product can pass every acceptance criterion on the board and still produce anxious, confused, or disengaged users. The build is clean but the behaviour is broken. And because the team has no language for measuring that gap, they often do not notice until something more dramatic happens: a drop in retention, a spike in support contacts, or a quiet drift away from the product that shows up in the numbers weeks later.

Behavioural success criteria give teams a way to describe what good looks like at a human level. They sit alongside acceptance criteria rather than replacing them, and they force a different question onto the table. What emotional and cognitive state should the user be in when this interaction is done? Writing that down changes how teams design, test, and evaluate the work.

The Limits of Acceptance Criteria

Acceptance criteria were designed to define when a piece of work is complete. In that role they are useful and often precise. They describe conditions that can be tested, checked, and closed off. A screen loads within two seconds. A user receives a confirmation email. A filter returns the correct results. These are all verifiable and, on their own terms, meaningful.

The problem is that they measure the product, not the person. They describe what the system does, and they are silent on what the user experiences while it does it. A checkout flow can satisfy every acceptance criterion and still generate enough anxiety at the payment step that a meaningful number of users abandon it. A data-entry screen can pass QA completely and still produce an error rate that signals users are confused and overwhelmed. The criteria were met. The interaction failed.

This is not a flaw in acceptance criteria as a tool. It is a limitation in scope. They were never designed to carry emotional or cognitive intent. But when teams treat them as the only measure of success, that limitation becomes a blind spot. Work that is technically complete gets shipped into a real human context that the criteria never accounted for.

Engagement metrics do not close this gap either. Session length, for instance, tells you how long someone stayed. It does not tell you whether they stayed because the product was working for them or because they were lost inside it. A number that looks healthy in a review meeting can be obscuring something quite different in practice.

Defining Behavioural Success Criteria

A behavioural success criterion describes the emotional or cognitive state a user should be in at a defined point in an interaction. Where an acceptance criterion asks whether the system performed correctly, a behavioural success criterion asks whether the user responded in the way the design intended. Both matter. Only one of them gets written down with any regularity.

Writing a behavioural success criterion starts with identifying the emotional problem the interaction is trying to solve, not just the functional one. A password reset flow is functionally about recovering account access. Behaviourally, it is about reducing the frustration and embarrassment that often accompany the moment of being locked out. If the criterion only describes the functional outcome, the team has no shared definition of what a good emotional outcome looks like.

Before writing a behavioural success criterion, name the emotional state the user is likely arriving with. A user landing on a cancellation page arrives differently to a user landing on a new feature announcement. The starting point shapes everything that follows.

A good behavioural success criterion is specific, observable, and tied to a testable moment. It names the emotional or cognitive state the design is aiming for, describes what that state looks like in practice, and connects it to a point in the user journey where it can be assessed. Vague intent, such as "users should feel comfortable", becomes a criterion when it is grounded. Comfortable doing what, at which step, and what would discomfort look like if it appeared?

This level of specificity is what makes behavioural criteria useful in testing. Teams can design research questions around them, observe for them directly, and build a shared understanding of whether the interaction worked at a human level, not just a technical one.

UX/UI design built around real psychology

We design app interfaces around how people actually think and behave. User research, psychology-driven UX/UI design and technical specs delivered as one complete package.

See how we work Get started

No commitment

The Emotional and Cognitive States That Matter

Not all emotional states deserve equal attention across a product. The states that matter most are the ones attached to moments where the product is asking something of the user. Browsing a content feed is low-stakes. Granting access to a contact list, entering payment details, or completing a medical history form is a different situation entirely. These are the moments where emotional and cognitive load spikes, and where behavioural success criteria do the most work.

Emotional states worth naming in criteria include confidence, anxiety, trust, frustration, confusion, and relief. Cognitive states include comprehension, recall load, and decision fatigue. Each of these produces observable signals in research and in live product data. A user who is anxious at a high-stakes request point hesitates, re-reads content, or abandons the flow. A user who is experiencing high cognitive load makes more errors, takes longer than expected on simple tasks, and often cannot describe what the screen was asking them to do.

Simon Lee describes a consistent pattern in high-stress environments where the issue is lower comprehension rather than an inability to find things on screen. Users operating under emotional pressure lose their grasp of the overall process they are going through. They abandon logical thinking and respond from a more reactive, emotional place. When comprehension of tasks drops significantly, and not just the practical execution but the fundamental understanding of what is being asked, that is a signal worth designing around explicitly.

Designing for emotional and cognitive states means knowing which states the interaction produces, not just what it delivers.

A product that accounts for this designs differently at those moments. It reduces the amount of information on screen, sequences steps carefully, and builds confidence before asking for anything sensitive. Writing that intent into a criterion makes it a design requirement rather than an aspiration.

A Practical Template for Writing Behavioural Success Criteria

A behavioural success criterion follows a consistent structure. It names the user, the moment, the target emotional or cognitive state, and the observable signal that confirms it. That four-part structure turns an intention into something a team can design toward and a researcher can test against.

The four-part structure

The structure runs as follows. First, name who the user is and what context they are arriving from. Second, define the moment in the journey the criterion applies to. Third, state the emotional or cognitive state the interaction is designed to produce. Fourth, describe what that state looks like when it is present, and what the absence of it looks like in testing or in data.

  • Who is the user and what is their emotional starting point at this moment?
  • What specific step or screen does this criterion apply to?
  • What emotional or cognitive state should the user be in when this step is complete?
  • What observable signals confirm that state is present or absent?

Keeping criteria testable

A criterion that cannot be tested is closer to a value statement than a success condition. "The user should feel confident" becomes testable when it specifies what confidence looks like in this interaction. Does the user proceed without re-reading the content? Do they complete the step within a certain time range? Do they rate their clarity highly in a post-task question? The signal does not have to be a number, but it does have to be something a person can observe and report on.

Write the "absence" signal as clearly as the "presence" signal. Knowing what failure looks like behaviourally makes it far easier to spot in testing before it reaches a live audience.

Teams that are new to this often find it easier to write the criterion after they have written the acceptance criteria for the same feature. The functional definition is already clear. The behavioural question then becomes: given that this is what the system does, what should the user feel and understand when they encounter it?

Worked Example: A Veterinary Practice Management Tool

A veterinary practice management tool handles clinical records, appointment scheduling, and billing across a busy practice. The users are vets and nurses who switch between tasks rapidly during the working day. The product has a feature that surfaces treatment reminders for animals with chronic conditions, prompting a staff member to contact the owner before an appointment window closes.

An acceptance criterion for this feature might read: the system displays a treatment reminder notification when an animal's scheduled review date falls within the next five days. That is clear, testable, and complete as a technical condition.

A behavioural success criterion for the same feature addresses what happens to the user when that notification appears. In a busy clinical environment, any additional alert competes with a high cognitive load. The user arriving at this notification is likely mid-task, focused on something else, and short on time. The behavioural criterion needs to account for that.

A well-formed criterion here describes a user who, on seeing the reminder, immediately understands which animal it refers to, what action is required, and how time-sensitive it is, all without needing to navigate away from their current task. The observable signal of success is that they take the correct action without opening the full record or asking a colleague for context. The signal of failure is hesitation, a mis-timed action, or the notification being dismissed and not returned to.

In high-frequency professional tools, the goal is cognitive speed with accuracy. Behavioural criteria for professional users often focus on comprehension and decision confidence rather than emotional comfort.

Worked Example: A School Parent-Communication App

A school parent-communication app lets teachers send messages, share updates, and flag concerns directly to parents. A new feature allows teachers to mark a message as requiring a read receipt, so they can confirm a parent has seen time-sensitive information, such as a trip permission reminder or a welfare check note.

The acceptance criterion confirms that a read receipt is triggered when the parent opens the message, and that the teacher sees the confirmation in their sent items. Functionally, this is complete.

The behavioural picture is more layered. For teachers, the moment of sending a read-receipt message carries a degree of professional weight. They want to feel confident that the request is appropriate in tone and that it will not feel punitive or surveillance-like to the parent. For parents, receiving a message flagged for a read receipt lands in a different emotional place to a standard update. Some parents will feel mild anxiety, particularly if they associate formal acknowledgement requests with bad news.

Behavioural success criteria here operate on both sides of the interaction. For the teacher, success means they send the message without second-guessing the framing, which suggests the feature's language and positioning did its job. For the parent, success means they open the message without apprehension escalating before they have read the content, and that the read-receipt mechanism feels like a straightforward administrative step rather than a test of their responsiveness.

Observable failure signals include teachers editing or deleting read-receipt messages before sending, and parents opening then immediately closing a message before re-opening it, a pattern that suggests initial alarm followed by a decision to re-engage.

Conclusion

Acceptance criteria describe what a product does. Behavioural success criteria describe what that product does to a person. Both belong in the same conversation, and the absence of the second kind leaves a team with only half the picture.

Getting this right is a discipline rather than a one-time exercise. It asks teams to name emotional and cognitive intent at the same point in the process where they name functional requirements, and to carry that intent through design, testing, and evaluation. It changes how research questions are written, how testing sessions are structured, and how shipped work is assessed against its original purpose.

The two worked examples above cover very different contexts, a time-pressured clinical environment and an emotionally loaded communication tool, and they need different kinds of behavioural criteria because the users and their emotional starting points differ. That context-dependence is the whole point. There is no generic emotional success condition that transfers across products. The criteria have to be written for the specific person, the specific moment, and the specific stakes involved.

Teams that build this habit find that it surfaces design problems earlier, makes research more productive, and gives everyone involved in a product a shared language for something that previously went unspoken. The functional work and the human work stop being separate conversations and start informing each other from the beginning.

If your team is working through what behavioural success criteria look like in your specific product context, let's talk about your product and where to start.

Frequently Asked Questions

What is a behavioural success criterion?

A behavioural success criterion describes the emotional or cognitive state a user should be in at a specific point during an interaction with a product. Unlike acceptance criteria, which measure whether a system performed correctly, behavioural success criteria measure whether the user responded in the way the design intended.

How do behavioural success criteria differ from acceptance criteria?

Acceptance criteria measure the product itself, confirming that technical conditions have been met, such as a form submitting without errors or a screen loading within a set time. Behavioural success criteria measure the human experience of that product, capturing whether users felt confident, informed, or at ease during the interaction.

Do behavioural success criteria replace acceptance criteria?

No, behavioural success criteria are designed to sit alongside acceptance criteria rather than replace them. Both types serve distinct purposes, and a well-rounded definition of success requires both the technical and the human dimension to be captured.

Why aren't acceptance criteria enough on their own?

A product can satisfy every acceptance criterion and still produce anxious, confused, or disengaged users, because acceptance criteria were never designed to carry emotional or cognitive intent. When teams treat them as the sole measure of success, they create a blind spot that can lead to poor user experiences going undetected until they surface in retention or support data.

Can engagement metrics fill the gap left by acceptance criteria?

Not reliably, as engagement metrics such as session length do not reveal the quality of a user's experience. A user may spend a long time on a screen because they are lost or confused rather than because the product is working well for them, making such metrics potentially misleading when reviewed in isolation.

What kinds of problems can arise when behavioural success is ignored?

Teams may ship technically complete work that generates user anxiety, confusion, or disengagement, which often goes unnoticed until more visible problems emerge. These can include drops in retention, spikes in support contacts, or a gradual drift away from the product that only becomes apparent in the data weeks later.

How does writing behavioural success criteria change how teams work?

Articulating the emotional and cognitive state a user should be in when an interaction is complete forces teams to think differently about design, testing, and evaluation. It introduces a human-centred question into the process that would otherwise go unasked, shaping decisions from the outset rather than being considered only after the fact.

At what stage of a project should behavioural success criteria be introduced?

Behavioural success criteria are most effective when written alongside acceptance criteria during the definition stage of a piece of work, before design and development begin. Establishing them early ensures that the human intent of an interaction informs how the work is shaped, tested, and ultimately judged as complete.