User Interviews And Focus Groups: The Scoring Rubric
A behavioral 1-to-5 rubric that scores your user interviews and focus groups on real signal, using The Mom Test, with an annotated example and an importable Lens.

User Interviews And Focus Groups: The Scoring Rubric
Last quarter you judged a user interview by how it felt on the way out: they nodded, they called the idea great, you wrote "strong signal" and moved on. This quarter you score the same conversation on what was actually said, and about half of those "great" sessions turn out to have produced nothing you can build on.
The reason is uncomfortable. A user interview or focus group that everyone enjoyed is usually the one where you pitched and the other person was being polite. Compliments, a quick "yeah, I'd use that," and feature requests all sound like data, and none of them are. The cost of mistaking the first for the second is not small: in CB Insights' March 2026 analysis of 385 failed venture-backed startups, poor product-market fit was the most-cited underlying reason for failure, at 43%. A good research conversation does not sound enthusiastic. It sounds like the participant's real past, a problem they already spent money trying to solve, and a next step that costs them something, and whether you got those is audible in the recording.
This piece scores a research conversation the way The Mom Test grades one. Rob Fitzpatrick's book runs on a single move: get people describing their own recent behavior rather than reacting to your idea, chase concrete specifics rather than promises about the future, and keep your own mouth shut. The rubric below turns that move into six criteria, each pairing what a weak answer sounds like against what a strong one sounds like, so two researchers scoring the same interview land on the same number. It is also how talk2bud reads a session. The app captures the conversation from your Mac's system audio, with no bot joining the call, and reads the transcript against a product-research Lens that returns the six scores and flags every compliment and hypothetical from what was said. No bot in the room is a capture choice, not a way to record in secret: you tell the participant and get their yes before you start, the way user research already asks you to. The rubric on this page and that Lens are the same object.
Everything here was verified in August 2026 against the sources linked inline: The Mom Test, Nielsen Norman Group's guidance on interviews and focus groups, CB Insights' 2026 startup-failure analysis, and the qualitative-research literature on social desirability bias.
The One-Page Play
The method fits on an index card. Open your last research session. Give each of the six criteria a 1 to 5 based only on what was said, and write the quote that earned the score. The two lowest criteria are your next session's homework. Any high average that hides a 1 on criterion 2 or 5 is a session that flattered you rather than taught you.
The signal tally is the number to watch. Count the facts and commitments you pulled, then count the compliments and hypotheticals you wrote down as if they were findings. A session can average 4.0 on rapport and still be worthless if it produced three compliments and zero facts. The full playbook, with anchors for all five score bands and a 30-day rollout, is in the download section below, as a PDF and an importable Lens.
What A Good User Interview Sounds Like
A checklist asks "did you avoid leading questions?" and a researcher who asked one open question ticks the box. A rubric asks what a 2 sounds like against what a 5 sounds like, with the phrasing attached, so the score survives a second reviewer. That property is the only thing that makes a rubric worth running twice.

Here is the full rubric, in a form you can copy into your own notes. Score each row 1 to 5, and quote the line that earned it.
| Criterion | Listen for | A 1-2 sounds like | A 4-5 sounds like |
|---|---|---|---|
| 1. Grounded in their past, not your idea | You open on how they handle this today, in recent specifics, before describing what you are building | "We're building a tool that scores research calls. Would that be useful?" | "Walk me through the last round of interviews you ran. What did you do after the calls?" |
| 2. Compliments deflected to a fact | When they praise the idea, you dig past the praise to a behavior or a number | "That's amazing, thank you!" written down as a yes | "Appreciate that. When did this last cost you enough that you went looking for a fix?" |
| 3. Specifics and numbers, not opinions | You pull concrete facts (last instance, frequency, time or money spent) instead of accepting "usually" | "Do you generally struggle with analysis?" answered "Yeah, sometimes." | "How many interviews did you run last quarter, and how long did the write-up take?" |
| 4. You talked less, and did not pitch | You let silence work, keep your idea out of the question, and they speak more than you | Every pause becomes another feature explanation | "Tell me more about that." Then you stop talking. |
| 5. The problem proven by prior effort | You establish they already spent time, money, or effort trying to solve this | "Would this be useful?" answered "Definitely," and never tested | "What have you already tried to fix this? What did that cost, and why did you stop?" |
| 6. A real commitment, not a warm goodbye | The session ends with a next step that costs them time, an intro, access, or money | "Super helpful, thanks so much!" | "Would you introduce me to the budget owner and run next month's study on the beta?" |
Criteria 2 and 5 are the two that decide whether you learned anything. They are the signal criteria, and they are the reason a research rubric looks different from a generic interview-quality checklist. Most guidance treats a leading question as the only sin. The subtler failure is a compliment you accept as evidence and a problem everyone agrees with that nobody has ever spent an hour trying to solve. Both are audible, and both are things you can score.
How The Compliment Trap Sounds
Fitzpatrick calls compliments the fool's gold of customer conversations: they feel like progress and pay out nothing. The trap is that a compliment and a fact are said in the same warm tone, so the only way to tell them apart is to read the words back.
Below is a short excerpt from a discovery interview. It is an illustrative example written to show the scoring, not a real transcript, and it is dated the way a real review would be: a discovery interview from 12 August 2026. The founder is "You"; the participant is a product manager who runs research. The margin notes are what the rubric writes.
You: So we're building a tool that records your user interviews and scores them for signal quality. Would something like that be useful to you? Participant: Oh, that sounds really useful. I'd definitely use that. You: Amazing, so glad to hear it. We think it could save you a ton of time. Participant: Yeah, for sure. Analysis is such a pain.
Criterion 1, score 1. The interview opened on the idea instead of the participant's last real project, so everything after it is a reaction to a pitch. Criterion 2 is a 1 too: "I'd definitely use that" is a compliment sitting on top of a hypothetical, and it went in the notebook as validation. The line that would have earned a 5 is short: "Thanks, when did analysis last cost you enough that you did something about it?" On criterion 5, another 1, because "analysis is such a pain" is agreement in the abstract and nobody checked what this person had already tried, so there is no evidence the problem is real.
Now the same participant, same forty minutes, run well.
You: Before I show you anything, take me back to the last user study you ran. What happened once the interviews were done? Participant: Last month I ran eight. I spent maybe two full days re-watching recordings and tagging quotes in a spreadsheet. You: Two days on eight. What did you try to make that faster? Participant: I paid for an AI notetaker for a while, but it summarized instead of pulling the exact quotes I needed, so I cancelled it and went back to the spreadsheet. You: If you could get the tagged quotes without the two days, would you put next month's study through it? Participant: If it actually pulls verbatim quotes, yes. Send it and I'll run the December round on it.
Criterion 1, score 5. It opened on a real, recent project. Real numbers put criterion 3 at a 5: eight interviews, two days. Criterion 5, score 5. The prior effort is on the record, a tool they paid for, why it failed, and that they cancelled it. The close earns criterion 6 a 4: a commitment with a cost attached, and no compliment in sight. Same request, same warmth, and the difference between a notebook full of encouragement and a fact, a number, a proven problem, and a next step you can hold someone to.
Calibration: Two Reviewers, One Number
A rubric earns its keep only when two people scoring the same session land within a point of each other. When they do not, the argument is almost always about an adjective and not a fact. Calibration fixes that by forcing the disagreement down to the quote.
Try it. Score criterion 6 on this closing exchange, then read the key.
You: This was so helpful, honestly. I'll pull my notes together and share what we're building once it's further along. Participant: Sounds great, can't wait to see it.
Score it before reading on. The answer is a 1. A reviewer who enjoyed the call wants a 4, because it ended warmly and you promised to follow up. Criterion 6 is not "did you intend to stay in touch." It is "did the participant commit to something that costs them." Nobody agreed to an intro, a trial, or a next date; "can't wait to see it" is enthusiasm, and enthusiasm is free. A score reads the transcript; a memory reads the mood, and only one of them holds up in a week. That is also why a talk2bud Lens and a careful human reviewer can land on the same number: both are scoring the words on the page, not the feeling in the room.
The rule that keeps calibration honest is the quote rule: no score without a line from the session under it. A score with no line under it is a guess wearing a number, and the moment you argue about the actual quote instead of your impression, the disagreement tends to settle itself. It is the same discipline that separates a performance self-assessment built on evidence from one built on adjectives: an artifact you can point to beats a feeling you can only assert.
How The Rubric Shifts For Focus Groups And 1:1s
The six criteria hold across research, but the weights move with the method.
On a one-to-one discovery interview, criteria 2 and 5 carry the most risk, because a single polite participant can fill forty minutes with pleasant hypotheticals and you will leave convinced. On an early customer-discovery call, criterion 1 is nearly the whole score: if you open on your idea, everything you hear is a reaction to a pitch, which is exactly the failure Nielsen Norman Group warns about when it notes that "if the interviewer asks many leading questions, the validity of the data will be compromised" (User Interviews 101, updated July 2026). A usability test, where you watch someone use the thing, flips the weights: the pitch risk drops and criterion 3 rises, because the score lives in what they did, not what they said about it.
A focus group is the method where this rubric bites hardest, and where you should discount the compliments most. Group settings amplify the exact biases criteria 2 and 5 exist to catch. Social desirability bias, saying what you think the room wants to hear, is stronger with an audience, a pattern documented across qualitative research (see "Everything Is Perfect, and We Have No Problems", Qualitative Health Research, 2020). Jakob Nielsen made the structural point in 1997: focus groups "can only assess what customers say they do and not the way customers actually operate the product," and the moderator "must avoid letting one participant's opinions dominate" (Focus Groups). So in a group, criterion 4 becomes the moderator's whole job, and a consensus that one confident voice built is not signal. Score the group's compliments at a discount and lean on criterion 5: what has each person actually spent trying to solve this.
The Four Ways Researchers Score This Wrong
A rubric removes most bias, but four failure modes survive if you let them, and the first one is specific to research.
Confirmation drift. You score the session by how much it validated the idea you walked in with. This is the founder's version of a halo, and it is the most expensive error on the list, because it rewards exactly the sessions that taught you least. Fix it by scoring criterion by criterion, in order, before you form any overall verdict, and by scoring criteria 2 and 5 first, since those are the two that punish validation-seeking.
The likable participant. An articulate, enthusiastic person makes you drift every criterion upward, even when they handed you opinions instead of facts. Charm is not a score. A delightful participant who never named a real past instance is a 2 on evidence wearing a smile.
Recency. You remember the last two minutes and grade the whole session on the goodbye. A warm sign-off inflates a conversation that decided nothing, which is why the closing exchange in the calibration section reads as a 1 once you score the words.
The dominant voice. In a focus group, one confident participant's opinion gets recorded as "what the group thought," even when the quiet four disagreed. That is the conformity Nielsen flagged, and it is why a group transcript needs each line attributed to a speaker before you score anything at all.
Rolling It Out In 30 Days
Week one is a pilot. Score your own last five interviews without changing how you run them, just to get a baseline and find your weakest criterion. In week two, calibrate: have a colleague score two of the same sessions, and anywhere you disagree by more than a point, argue it to the quote until the anchor settles it. By week three, every live session goes through the rubric, and you report the signal tally to yourself each Friday, facts pulled against compliments banked. The final week is for the trend rather than the averages. If criteria 2 or 5 stay low, the fix is not more interviewing; it is rewriting the discussion guide to open on the last time it happened instead of on the idea.
Running The Rubric On The Interview You Just Did
You will not score sessions by hand forever, and the point of the rubric is to run on every conversation without adding an hour of admin to a week that is already full. Here is the path from a live interview to a score.
1. Capture the session without a bot in the room. talk2bud records from your Mac's system audio, so nothing joins the call as a participant and no bot shows up in the attendee list. That is a recording choice, not a secret one. Capturing without a bot is not the same as capturing without consent, so you still tell the participant you are recording and get their yes, which research ethics asked of you already. On a research call the line is short: "I'm recording so I can quote you accurately instead of paraphrasing, is that okay?" How talk2bud captures a call from your Mac's system audio, and why that beats a screen recorder that loses the other side's voice, is worth a read on its own.
2. Pick the research Lens. In Modes, switch to Lenses and choose Client discovery or Interview, or duplicate one and paste in the rubric from this page as its playbook. The Lens is the rubric: the same six criteria, the same quote rule.

3. Read the score, not the recording. When the session ends, talk2bud scores the transcript against the Lens and hands back the six criteria with the quote that earned each, plus every compliment and hypothetical flagged and whether you converted it into a fact. You read a page, not a forty-minute recording. Whatever scored lowest is the next session's discipline.
The question that stops most teams is not which criteria to use; it is who re-listens to a week of interviews on Friday to score them. Nobody does, which is why research quality quietly slides. A human reviewer forgets which "yes" was a compliment by the time the call is over. A Lens reading the transcript does not, and it quotes the line back to you before you build on it.
Download The User-Interview Playbook
The full playbook is free, under a CC BY 4.0 licence, so you can adapt it and re-share it with attribution.
- The User-Interview Playbook (PDF, 10 pages). The argument, all six criteria with anchors for every score band, the annotated excerpts, the calibration exercise, and the 30-day rollout. Download the PDF.
- The importable Lens (Markdown). The same six criteria as a file talk2bud imports as a call Lens. In the app, go to Modes, then Lenses, then New, and import it. Download the Lens file.
There is no email gate. The playbook and the Lens are the same six criteria, so the score you give a session by hand and the score the app gives it come from one rubric.
Frequently Asked Questions
What is the difference between user interviews and focus groups?
A user interview is a one-to-one conversation between a researcher and a single participant, usually 30 to 90 minutes, designed to explore that person's behavior and experience in depth. A focus group brings 5 to 10 participants together with a moderator for a guided group discussion, typically 1 to 2 hours. The trade-off is group dynamics: focus groups gather many views quickly, but they are prone to groupthink and dominant voices, so one-to-one interviews usually produce more honest data about what people actually do.
How many participants should a focus group have?
A focus group should have roughly 5 to 10 participants. Fewer than 5 can feel thin and stall, while more than 10 makes it hard for the moderator to give everyone airtime and tends to leave quieter participants unheard. The moderator's job is to manage the dominant voices and draw out the quiet ones, because a consensus built by one confident participant is not a reliable finding.
What is the Mom Test for user interviews?
The Mom Test is Rob Fitzpatrick's method for talking to customers so that even your mom cannot lie to you about your idea. Its three rules are: talk about their life instead of your idea, ask about specifics in the past instead of opinions about the future, and talk less and listen more. Applied to user interviews, it means treating compliments and hypotheticals as noise and scoring the conversation on concrete past behavior and real commitments.
What is a leading question in a user interview?
A leading question is one that suggests its own answer, such as "Don't you find this frustrating?" or "Wouldn't a tool like this help?" It biases the participant toward agreeing with you and produces data that only confirms what you already believed. Nielsen Norman Group notes that many leading questions compromise the validity of interview data; the fix is a neutral, past-tense question like "Tell me about the last time you did this."
Are focus groups reliable for product research?
Focus groups are useful for discovering what users want to discuss and generating early ideas, but they are unreliable as a sole source of truth about behavior. As Jakob Nielsen observed, a focus group surfaces what people say they do, which often differs from how they actually use a product, and group settings amplify social desirability bias. For decisions about what people actually do, one-to-one user interviews or direct observation give more trustworthy signal, and any focus-group finding should be discounted for the room effect and confirmed against real prior behavior.
Verified August 2026 against The Mom Test by Rob Fitzpatrick, Nielsen Norman Group's "User Interviews 101" (updated July 2026) and "Focus Groups" (Jakob Nielsen, 1997), CB Insights' 2026 startup-failure analysis (385 companies), and Bergen & Labonté's 2020 study on social desirability bias in qualitative research. Figures are quoted from those sources; the interview excerpts are illustrative examples written to show the scoring, not real research sessions.