There is a specific kind of failure that looks like success. You run eight customer interviews, fill a document with quotes, present a deck titled "What We Heard", and the roadmap does not change by a single item. The research happened. The decision did not.
This is almost never caused by talking to the wrong people. It is caused by asking questions whose answers cannot change a decision.
The rule that fixes most interviews
Ask about past behaviour, not future intent. People are unreliable narrators of what they will do and reasonably good reporters of what they did.
Compare two questions about the same feature:
- "Would you use a bulk export?" — invites politeness. Almost everyone says yes.
- "Walk me through the last time you needed data out of this system." — produces a story with dates, workarounds, and the actual cost.
The second question cannot be answered with a courtesy. It either produces a specific incident or it reveals that the problem does not happen often enough to remember — which is itself the answer you needed.
If nobody can recall a recent instance of the problem, you are not looking at an underserved need. You are looking at a hypothetical.
The question set that does the work
Five prompts will carry most discovery conversations:
- "Tell me about the last time you [did the job]." Anchors to a real event.
- "What did you do just before that? And after?" Reveals the workflow either side, which is usually where the real friction lives.
- "What made you deal with it that day rather than the week before?" Surfaces the trigger — the single most useful thing for positioning.
- "What did you try first?" Names your actual competition, which is frequently a spreadsheet.
- "What happened as a result?" Quantifies the pain, or exposes that there wasn't much.
Notice none of them mention your product. The moment you describe what you are building, the interview stops being research and becomes a pitch with a politeness filter attached.
After someone finishes an answer, wait. Three seconds feels unbearable and produces the most useful sentence in most interviews — the qualification, the exception, the thing they assumed was too obvious to mention.
Five formats, and how each one breaks
Every research format has a predictable failure mode. Choosing well is mostly a matter of deciding which failure you can afford today.
| Format | Good for | How it fails |
|---|---|---|
| One-to-one interview | Depth, sensitive topics, understanding the why | One person's tidy version of reality; exceptions get forgotten |
| Usability test | Whether people can operate what you built | Answers "can they" and never "should we" |
| Survey | Sizing something you already understand | You can only learn what you thought to ask; responders differ from non-responders |
| Session replay / analytics | What actually happens, at scale | Shows behaviour, never motive. You will invent the why |
| Sales and support calls | Free, constant, already recorded | Skewed to complainers and to deals in progress |
The strongest combination is the cheapest one: mine existing support tickets and sales calls to form a hypothesis, then run five interviews to understand the motive behind the pattern. That sequence costs a couple of days and beats a month of unfocused conversations.
One articulate, senior, enthusiastic customer will distort a roadmap more than any other single force in product management. They are memorable, they follow up, and they sound like the market. Check every strong signal from one account against behaviour across the base before you build.
How many interviews
Stop when you stop being surprised. In practice that is around five to eight for a reasonably narrow segment, and it climbs sharply if you mix segments — five interviews spread across five different customer types teaches you almost nothing, because you have one data point per group.
Interview one segment at a time. Saturation within a segment is the signal; volume across segments is noise.
Synthesis, which is where most research dies
Raw notes are not findings. The step everyone skips is converting observations into something with a decision attached.
A synthesis that works has three levels:
- Observation — "Four of six exported to CSV and rebuilt the same pivot table each month."
- Interpretation — "The reporting view does not answer the question they are actually asked in their monthly review."
- Decision — "Build the monthly-review view; deprioritise the export improvements we scoped."
If your research output stops at the first level, you have produced a document. If it reaches the third, you have produced a decision. Only the third changes anything.
Making it routine
The single highest-leverage change most teams can make is moving from occasional research projects to a standing weekly slot — one or two conversations, every week, forever. It removes the negotiation about whether to do research, keeps the sample fresh, and means you are never more than seven days from a real customer.
Once you have a steady stream of conversations, the next problem is what to do with them. That is where an opportunity solution tree earns its keep, and where jobs to be done gives you a way to phrase what you found.