Skip to content
Expert Guide Series

How to Tell if Your Product Spec Is a Decision or a Wish List

Most product specs look like decisions on paper. They have headings, feature descriptions, acceptance criteria, maybe even a priority column. But if you read them carefully, you start to notice something: a lot of them are really just lists of things people hope will be true. They describe features rather than outcomes. They say what the product will do without explaining why that thing matters or how you would know if it worked.

A wish list and a spec can look identical at the top of a document. The difference lives in what each one demands of the person writing it. A real decision requires you to name the behaviour you want to change, state the assumption underneath it, and commit to what would prove you wrong. A wish list asks you to do none of those things. It just asks you to keep adding rows.

This matters because the gap between those two things is where products go wrong. Teams build features that made sense in a meeting but solve no real problem for a real person. They ship things confidently, then find that users respond in ways nobody predicted. The spec looked complete. The thinking behind it was not. At We Are Affective, we see this pattern across product work of all kinds, and we have a simple three-part diagnostic we run to close the gap. This article walks through how it works.

Why Specs Drift Into Wish Lists

The drift happens gradually, and it usually starts from a reasonable place. Someone has an idea for a feature. They write it down. Someone else adds a related idea. A stakeholder review adds three more. Before long, the document reflects a collection of opinions rather than a set of grounded decisions, and nobody has stopped to ask whether the thinking underneath each item is actually sound.

Part of what drives this is that writing down a feature feels productive. It looks like progress. A spec with twenty items in it feels more thorough than one with five, even if the twenty-item version has never been tested against a real user need. The document grows because growth feels like momentum.

A 2022 survey by UserZoom and Ipsos found that approximately 72% of product decisions were made without any user research informing them at all. That figure is worth sitting with. Not research that was conducted and ignored, but decisions made with no user data present in the room. In that context, it is easy to see how specs become wish lists. When no external check exists, the only thing shaping the document is what the people in the room believe, and belief without evidence tends to accumulate rather than sharpen.

The fix is a set of questions you run against every item before it earns its place on the list.

The Three-Part Diagnostic

The diagnostic we use at WAA has three parts, and each one does a specific job. The first asks you to name the behaviour change the feature is designed to produce. The second asks you to surface the assumption underneath that expectation. The third asks you to define what evidence would prove that assumption wrong.

These three questions work together because each one depends on the previous. You cannot sensibly identify an assumption if you have not yet named a behaviour. You cannot define a falsifying condition if you do not know what you are assuming. The sequence matters, so the diagnostic runs in order every time.

What makes it useful is also what makes it uncomfortable. Most features, when you run them through these questions, cannot answer all three cleanly. They can describe what the feature does, but they struggle to name a specific behaviour change. They can state an assumption, but they cannot say what would prove it false. That failure is diagnostic in itself. An item that cannot pass all three parts is a wish, not a decision, and it does not belong in a committed spec until it can.

Before adding any item to a spec, write one sentence describing the specific user behaviour this feature is designed to change. If you cannot write that sentence, the feature is not ready to be specced.

Running this diagnostic across a spec rarely reveals that every item is wrong. More often, it reveals that some items are well-grounded, some are close but need more work, and a few are carrying assumptions nobody has ever tested. That sorting process alone is worth the exercise.

Design that understands your users

We build app experiences around real user behaviour, not assumptions. Research, psychology-driven design and technical specs that turn users into loyal advocates.

See how we work Get started

No commitment

Naming the Behaviour Change

Every feature in a product spec exists, at some level, because someone believes it will change how users behave. They will complete more bookings, read more articles, submit fewer support requests, return more often. But most specs do not state this explicitly. They describe the feature itself and leave the behaviour change implied.

Leaving it implied is a problem. When a team builds something without agreeing on the specific behaviour it is meant to produce, they have no shared reference point for evaluating whether it worked. After launch, one person points to session length as evidence of success. Another points to support volume. A third looks at return visits. All three are measuring different things, because nobody agreed upfront on what the feature was actually supposed to do.

Naming the behaviour change forces that agreement before the build. It asks you to state, precisely, what a user will do differently because this feature exists. Not "users will find it easier to manage their account" but "users who currently abandon the account setup flow at step three will complete it." The specificity is the point. Vague outcomes cannot be measured, and things that cannot be measured cannot tell you whether a decision was right.

This step also surfaces features that have no clear behavioural purpose. If a team cannot agree on what behaviour a feature is meant to change, the honest question is why the feature is in the spec at all. Often the answer is "because a competitor has it" or "because it seemed like a good idea in the meeting." Those are not good enough reasons to build something.

A feature without a named behaviour change is an opinion dressed up as a product decision.

Once the behaviour is named, the rest of the diagnostic has something concrete to work with.

Surfacing the Underlying Assumption

Every product decision rests on at least one assumption. The assumption might be about what users currently find difficult, about how they make choices, about what would motivate them to act differently, or about how they feel in a particular moment. These assumptions are always present. The only question is whether they have been made visible.

When assumptions stay hidden, teams treat them as facts. A feature gets built on the belief that users will notice a new prompt and respond to it, without anyone checking whether users in that context are actually paying attention to that part of the screen. A new flow gets designed on the belief that simplifying a process will increase completion rates, without anyone asking whether complexity was actually the reason people were dropping off.

For each feature in your spec, write down the single most important thing that has to be true about your users for this feature to work. That is your core assumption. Write it in plain language, not as a hypothesis.

Surfacing the assumption does not mean the assumption is wrong. It means the team knows what they are betting on. A lot of product work involves making reasonable bets based on incomplete information, and that is fine. What is not fine is making those bets without knowing you are making them, because then nothing prompts you to check whether you were right.

The most common assumptions we find hiding in specs are about emotional state. A feature might assume that users feel anxious at a particular point in a journey, or confident, or motivated. Those emotional states are real and important, but they are rarely written down as assumptions. They sit underneath the spec, invisible and unexamined, shaping every decision above them.

Defining What Would Prove It Wrong

This is the part most teams skip. Naming a behaviour and surfacing an assumption takes effort, but it still feels relatively safe. Defining what would prove the assumption wrong is harder, because it means committing to a falsifying condition before you know whether you will hit it.

The question is simple to ask and genuinely difficult to answer well. If your feature assumes that users find the current account setup flow too long, what data would tell you that length is not actually the problem? If your feature assumes that adding social proof to a booking page will increase completions, what result would tell you the assumption was wrong? A drop in completions is obvious. But what about flat results? What about completions that rise but support requests that rise too?

Defining a falsifying condition requires the team to think through what they would actually do if the feature did not produce the expected behaviour change. That is a useful discipline because it reveals whether the team genuinely believes the assumption is testable, or whether they have framed it in a way that makes it impossible to challenge.

  • A good falsifying condition is specific and observable, not subjective.
  • It names a threshold, not just a direction. "Completion rates fall below 60%" is a condition. "Users don't like it" is not.
  • It is agreed before the feature ships, not invented after to explain a result.
  • It does not require a perfect controlled experiment, just a plausible real-world check.

When a team cannot define what would prove an assumption wrong, that is often a sign the assumption has been framed too broadly, or that the feature is carrying more certainty than the evidence supports.

Write your falsifying condition as a simple "if this happens, we were wrong" statement and agree it with the whole team before the feature goes into development. Revisit it at your next review after launch.

Running the Diagnostic Across Your Spec

Once you understand how the three parts work individually, running the diagnostic across a whole spec is a matter of applying the same questions to every item in order. The process surfaces a natural sorting of features into three groups: those that are genuinely ready to build, those that need more work before they can be committed to, and those that are essentially guesses.

How to run it in practice

Take each item in the spec and write three short statements alongside it. The first names the behaviour change. The second names the core assumption. The third describes what would prove it wrong. If any of the three cannot be written in plain, specific language, the item goes into a holding category until the gap is filled.

This does not need to be a long process. For a spec with fifteen to twenty items, a working session of ninety minutes with the right people in the room is usually enough to sort the list. The session tends to surface healthy disagreement, because different team members often have different implicit assumptions about why the same feature is in the spec.

What to do with the results

The features that pass all three parts cleanly are your committed spec. The features that have a named behaviour but an untested assumption move to a research or validation phase before any build commitment is made. The features that cannot answer the first question clearly get removed from the spec entirely or returned to an ideas list.

The result is a shorter document than you started with, but a more honest one. Every item in it represents a real decision, with a named outcome and a known assumption, and the team knows in advance what would tell them they were wrong.

Conclusion

A product spec that reads like a wish list is not a planning problem. It is a thinking problem. It reflects what happens when features are added faster than the reasoning behind them is examined, and when the people writing the document have not been asked to commit to anything beyond the description of the feature itself.

The three-part diagnostic asks for three things that should be easy to provide if a decision is real. A named behaviour change. A visible assumption. A condition that would prove the assumption false. When all three are present, the feature belongs in a spec. When they are not, it belongs in a conversation first.

Most teams find that running this diagnostic does not slow them down. It saves time, because it stops build work beginning on features that would eventually have been questioned anyway. The reckoning that happens after launch, when a feature fails to produce the expected result and nobody can agree on why, takes far longer than the twenty minutes it would have taken to run the diagnostic before a single line of code was written.

The goal is a spec where every item is something the team genuinely decided, with clear reasoning and a known way to check whether they were right. That kind of document is worth building from. If you want to think through how this applies to your own product work, you can find us at weareaffective.com.

Frequently Asked Questions

What is the difference between a product spec and a wish list?

A genuine product spec requires you to name the behaviour you want to change, state the assumption underneath it, and commit to what would prove you wrong. A wish list simply describes features without explaining why they matter or how you would measure success.

Why do product specs so often drift into wish lists?

The drift typically happens because writing features down feels productive and gives the impression of progress, encouraging teams to keep adding items. When no user research is present to act as an external check, the document ends up reflecting accumulated opinions rather than grounded decisions.

How common is it for product decisions to be made without user research?

A 2022 survey by UserZoom and Ipsos found that approximately 72% of product decisions were made with no user research informing them at all. This absence of external data is a key reason why specs so frequently become wish lists.

What is the three-part diagnostic described in the article?

The diagnostic involves three questions run against every item in a spec: what behaviour change is the feature designed to produce, what assumption underlies that expectation, and what evidence would prove that assumption wrong. The questions must be answered in sequence, as each one depends on the previous.

Why does the order of the three diagnostic questions matter?

Each question builds directly on the one before it, so the sequence cannot be skipped or rearranged. You cannot meaningfully identify an assumption without first naming a target behaviour, and you cannot define a falsifying condition without knowing what you are actually assuming.

What happens when teams build features based on wish lists rather than real decisions?

Teams often end up shipping features that made sense in a meeting but solve no genuine problem for a real user. They may release a product with confidence, only to find that users respond in ways nobody anticipated because the underlying thinking was never properly tested.

Is writing a longer or more detailed spec the solution to this problem?

No — the article is clear that a longer spec or a more elaborate template will not fix the issue on its own. The solution is a set of rigorous questions applied to every item before it earns its place in the document.

How can you tell whether a feature in your spec is a real decision or just an addition to a wish list?

Run the feature through the three-part diagnostic: if it cannot cleanly answer what behaviour change it produces, what assumption underpins it, and what would prove that assumption wrong, it is likely a wish list item rather than a grounded decision. Most features, the article notes, struggle to pass all three questions.