Stop Guessing What Users Want: User Interviews and Need Validation for Vibe Projects
AI made building cheap and building right expensive. A field-tested user interview playbook: the five-person recruiting rule, Mom Test questioning, 20 ready-to-use questions, a 30-minute interview structure, affinity mapping and a pain-ranking matrix — plus a lightweight 4-interviews-a-month SOP for solo teams.

AI Made "Building It" Cheap. "Building the Right Thing" Got Expensive.
Let me start with a scene you've probably lived through: inspiration strikes on a Friday night, you vibe-code an MVP over the weekend, ship it to your community on Monday asking for feedback, and by Wednesday you realize — nobody uses it. Not because the code is bad or the UI is ugly, but because the "problem" barely exists, or isn't painful enough for anyone to switch tools.
This is the most expensive failure mode for vibe coders. AI solved the cost of "turning an idea into software." It did nothing about "how much the idea itself is worth." It used to take three months to build an app, so you'd think twice before starting. Now you can ship in a weekend, and the thinking-twice step gets skipped — the wrong direction executed brilliantly just wastes more.
User interviews are the cheapest form of need validation. The price of a coffee and a 30-minute call buys you three months of not building the wrong thing. It has the highest ROI of anything in the product process — it just looks the least glamorous, so nobody wants to do it.
Let me clear up one misconception first: an interview is not market research, not a survey, not a focus group. It's talking to real users about things they actually did, digging needs they didn't know they had out of their past behavior. You're not after opinions. You're after evidence.
Recruiting: Who to Find, Where to Find Them, What to Offer
The five-person rule: don't start by trying to interview fifty people
Jakob Nielsen's famous finding from usability research: five users uncover roughly 85% of usability problems. Need-finding interviews work the same way — you're not doing statistically significant sampling, you're looking for patterns. If three out of five people mention the same pain point, that's a signal. That's enough to act on.
The practical rule: after five interviews, if you're hearing nothing new, this round is done — go build. If all five people told you completely new things, you recruited the wrong crowd; swap in a fresh batch of five. Don't interview twenty people from the wrong audience. Interview volume can't rescue bad recruiting.
Where to find people, ranked by conversion rate
Tier one is your existing users. In-app popups, your email list, your paid-user group — these people have already spent time or money on your product, so response rates are highest. Don't write "please fill out my survey." Write "I want to hear about the specific trouble you ran into last time you used XX. 30 minutes, and there's a small gift card afterward." Concrete, sincere, no hype.
Tier two is relevant communities. Twitter, niche Discords, Slack groups, forums in your vertical. Don't post a recruiting thread right away — lurk for a week, watch who's complaining about problems adjacent to your direction, then DM them. Someone mid-complaint is the best interviewee you'll ever get: their pain is fresh and emotional, no extraction required.
Tier three is a cold-start list. People who left one-star reviews for existing solutions on Product Hunt, Reddit, or app stores — DM them directly. Low conversion, but they're gold for validating "what's wrong with the current options."
Paid recruiting platforms are the last resort. Paying for interviews saves hassle, but platform respondents are professional interviewees — their answers skew polished, and the information density is lower than the first two tiers.
Write recruiting criteria as behavior, not labels
This is the line beginners get wrong most often. Bad example: "25-35, big-city white collar, interested in productivity tools." Good example: "someone who manually organized expense reports three or more times in the past month." Behavioral descriptions self-filter — people matching the behavior almost certainly have a real pain; people matching the label might just be "interested."
What to offer: buy their time, not their answers
For a 30-minute interview, a $5-10 gift card or coffee voucher is a respectable range. In B2B, offer a lifetime free tier or early access to new features. Remember the logic: the reward buys their time, not the "right" answers.
One detail: give the reward after the interview, not before. Pay upfront and they'll feel they owe you, and every answer turns polite — you paid for flattery and got nothing useful. A terrible trade.
Designing the Guide: Ask About Past Behavior, Not Future Intentions
Rob Fitzpatrick nailed this in The Mom Test: good questions ask about specific past behavior; bad questions ask about future opinions. The book's title comes from the fact that if you ask your mom about your idea, she'll always say "that's wonderful, sweetie" — friends and polite praise are equally bad data sources.
Compare. Bad question: "If there were an app that auto-organized expenses, would you use it?" — everyone says yes, and that sentence carries zero information. Good question: "When was the last time you did expenses? How long did it take? Walk me through exactly what you did." — they can only describe something that actually happened. It's very hard to fabricate.
One core principle: task observation beats direct questions, and within direct questions, behavior beats opinion. If you can watch them do it, don't just ask them to describe it. If you can only ask, ask what specifically happened last time — never ask them to predict the future.
Twenty questions you can steal outright
These twenty are the ammunition I've reused across interviews for tool-type products, organized in four groups. Note: you don't ask all of them in one session — six to eight deep dives in 30 minutes is plenty. Order doesn't matter either; follow their thread and grab whichever fits.
Group A: background (set context, 5 questions)
- 1. How do you normally handle XX? (Open-ended opener; gets them into storytelling mode.)
- 2. When was the last time you did this? (Locks onto a concrete event; no vague generalities.)
- 3. How long did it take? (Quantify — time is the best pain ruler.)
- 4. Besides you, who else is involved in or affected by this? (Surfaces hidden stakeholders.)
- 5. How important is this in your work/life? Rate it 1 to 10. (Importance calibration.)
Group B: behavior deep-dives (the core ammo, 6 questions)
- 6. Can you walk me through exactly what happened last time, starting from step one? (Timeline reconstruction — the devil is in the details.)
- 7. What tool did you use? Why that one instead of the alternatives? (An honest review of existing solutions.)
- 8. Which step was the most annoying? Annoying exactly how? (Pain localization — ask "exactly how" at least twice.)
- 9. Did it ever go wrong? What happened? (Failure stories have the highest information density.)
- 10. Did you try other approaches? Why did you give up on them? (Tests whether they truly care — people who care tinker.)
- 11. If this step could magically disappear, how much would you pay? (Price probe — the intensity ruler.)
Group C: task observation (highest information density, 4 questions)
- 12. Can you show me right now? (Screen share or in person — watching once beats asking ten times.)
- 13. Why did you click here instead of there? (The real logic of a decision instant.)
- 14. You hesitated there — what were you thinking? (Hesitation marks friction.)
- 15. If you could keep only one feature, which one? (Forces ranking — ranking is prioritization.)
Group D: wrap-up and referrals (don't waste the last 3 minutes, 5 questions)
- 16. Is there anything I didn't ask that you think I should know? (The most frequent source of surprise insights.)
- 17. If you were recommending a solution to a friend with the same problem, how would you describe it? (Their natural language is your landing page copy.)
- 18. Do you know anyone else who struggles with this? Could you introduce me? (Snowball recruiting — the best recruiting channel there is.)
- 19. Can I come back in a month to share our progress? (Builds a long-term relationship — interviewees become your advisory board.)
- 20. If we build a solution, would you be our first tester? (The commitment test — saying yes is easy; watch what they actually do.)
Question 11 (how much would you pay) is the litmus test of the whole interview. If they hem and haw with "well... it depends," the pain's intensity is probably fake. People in real pain blurt out a number — or ask "when can I buy it, I want it now."
Running It: A Standard 30-Minute Structure
Minutes 0-5: warm up, don't dive straight in
Open with a bit of their background: "What do you do? Do you use our product? How did you find us?" The goal is to relax them, switching from "being interrogated" mode to "telling stories" mode. State the ground rules clearly: I'll record this, the recording is for internal notes only, you can stop anytime, and there's a small gift card at the end. People only tell the truth when the rules are clear.
Minutes 5-20: behavior deep-dives + task observation — the core 15 minutes
For these 15 minutes, the other person should be talking 80% of the time. Your entire script is three lines: "And then?" "Can you be specific?" "What were you thinking at that moment?" Resist the urge to give advice. Resist the urge to say "we're actually building exactly that" — the moment you reveal your hand, everything after is them telling you what you want to hear.
If they're willing, have them demo how they do it today. People auto-beautify when describing their own behavior ("I usually categorize things first"), but hands don't lie — you'll watch them dump everything into one folder without categorizing at all. That gap is the opportunity.
Minutes 20-27: pain confirmation — play back what you heard
"So what you're saying is: every month-end you spend two hours reconciling manually, and the worst part is the invoices never match the numbers in the system. Right?" Then shut up and let them correct you. This is the highest-density part of the whole interview — every correction is a precise need statement. Write it down verbatim.
Minutes 27-30: wrap-up — referrals + long-term contact
Ask Group D questions 16-20. Get referrals on the spot — don't say "ask around and let me know later." Get the WeChat/LinkedIn/email right there. Confirm the follow-up permission on the spot too. When you come back a month later with progress, they'll feel respected — that's how your first batch of seed users is born.
How to record: audio + transcribe within 24 hours + verbatim cards
Always record. Phone, Tencent Meeting, Zoom — anything works; informing them and getting consent is the non-negotiable baseline. Don't trust your notes — note-taking distracts you, and worse, notes capture your interpretation, not their words.
Transcribe within 24 hours of the interview. AI transcription is cheap and good now; 30 minutes of audio costs essentially nothing to turn into text. Past 24 hours, your memory starts editing to flatter your preferences — you'll only remember the parts that support your idea. That's confirmation bias, and it means the interview was wasted.
After transcribing, make "verbatim cards": pull out key quotes one by one, one quote per card, tagged with speaker ID and timestamp. For example: "Month-end reconciliation takes two whole evenings; the invoices never match the system — #3, 12:40." Don't just record conclusions — conclusions are your interpretation and they lose information. Verbatim quotes are the asset; they're the raw material for the next step, clustering.
Affinity mapping example: spread verbatim cards from five interviews on the table and group them by "these are really about the same thing" — let the categories grow out of the cards
From Interviews to Needs: Three Steps
Step one: affinity mapping — let the categories grow themselves
Spread all the verbatim cards out — physical sticky notes across a wall, or an online whiteboard like FigJam or Miro. Then group them by "these are really about the same thing." The key rule: no preset categories. Don't draw three boxes labeled "usability / performance / missing features" and stuff cards into them — that's validating your own hypotheses, not clustering.
Five interviews on a solo team typically produce 80-150 cards, clustering into 6-10 themes. Name each theme in the users' own voice — "invoices never match up" rather than "data consistency issues." The former reminds you of a user's face every time you see it; the latter reminds you of a Jira ticket. Which one would you rather work overtime for?
Step two: pain ranking — the frequency × intensity matrix
Clustering tells you what users are saying; ranking tells you what to build first. Draw a 2×2 matrix: the horizontal axis is frequency (how many people mentioned it, how often), the vertical axis is intensity (how emotional they got, what they'd sacrifice — your question-11 answers are the intensity ruler).
high intensity
│
low-freq, │ high-freq,
high-intensity │ high-intensity ← build v1 here
(niche opportunity,│ (core needs,
park in backlog) │
────────────────────┼────────────────────
low-freq, │ high-freq,
low-intensity │ low-intensity
(ignore outright) │ (polish backlog,
│ do when free)
│
low intensity
low freq high freq
The top-right quadrant (high frequency, high intensity) is your v1 feature list, no debate. Don't trash the top-left (low frequency, high intensity) too fast — few people mention it, but each would pay real money. That's often a high-willingness-to-pay niche: park it in the backlog marked "to validate" and revisit when you have more users. Bottom-left gets ignored outright; bottom-right goes into the experience-polish backlog.
Pain ranking matrix example: frequency on the horizontal axis, intensity on the vertical — top-right is v1 must-build, top-left is a niche opportunity awaiting validation
Step three: the needs pool — one table to rule every idea
Clustering and ranking results must be written down, or next month you'll forget why you deprioritized something. A solo team doesn't need a complex PRD — one table is enough:
Needs pool template (one row per need)
─────────────────────────────────────────────
Need: Auto-detect invoices and fill expense forms
Sources: #2, #3, #5 (3 of 5 mentioned it)
Frequency: High (3/5)
Intensity: High (#3: "would pay $8/month", #5: "overtime every month-end over this")
Quadrant: High-freq, high-intensity
Status: To build
Notes: #2 adds: only VAT special invoices, doesn't care about regular ones
─────────────────────────────────────────────
Need: Auto-generate English versions of expense reports
Sources: #4 (1 of 5 mentioned it)
Frequency: Low (1/5)
Intensity: High (#4: "required for foreign-company audits, would pay separately")
Quadrant: Low-freq, high-intensity
Status: To validate
Notes: Possibly a foreign-company-employee niche; re-interview #4 when user base grows
─────────────────────────────────────────────
Only five statuses: to validate / validating / to build / shipped / dropped.
Keep the dropped ones too, with the reason written down — three months from
now you'll forget why you dropped it and excitedly rebuild the same thing.
A Lightweight Cadence for a Team of One: Four Interviews a Month
The biggest enemy of interviewing isn't bad method — it's not sticking with it. A solo team doesn't need a dedicated researcher. It needs a rhythm so light it can't fail:
- Week 1: recruit. Send 10 invites (email/DM/community post), aiming to book 2 people. Use a fixed invite template; personalize one line each time. Never rewrite from scratch.
- Weeks 2-3: two interviews each. Fixed slot — say, Wednesday 8pm. On the calendar, non-negotiable. Don't do "when I have time" — you will never have time.
- Week 4: cluster + update the needs pool. Two hours: cluster the four sessions' verbatim cards, update statuses and quadrants in the pool. This is the step that turns interviews into decisions. Skip it and the interviews were just chatting.
Total investment: about 6-8 hours a month — roughly two fewer TV episodes a week. The payoff: every item on your roadmap can be traced to "#2, #3, and #5 said this verbatim" instead of "I think users would like it."
Interview debt: how long without interviewing becomes dangerous?
My personal red line: six straight weeks without talking to a user is a debt default. The symptom is easy to spot: you start saying "I think users would like" instead of "I heard users say." The moment that sentence appears, book interviews — no matter how full the roadmap is.
Another quantifiable signal: more than five needs sitting in "to validate" status that nobody has discussed yet means you're building behind closed doors. Interview debt works like tech debt — the longer you carry it, the higher the interest. The difference: when tech debt blows up, the system crashes; when interview debt blows up, you spent three months building something nobody wanted.
Common Pitfalls — All Ones I've Stepped In
Pitfall 1: leading questions. "Don't you find manual reconciliation annoying too?" — don't. You've fed them the answer; their nod is just politeness. Instead: "How do you feel about reconciliation?" Hand the judgment back to them.
Pitfall 2: only interviewing people you know. Friends, coworkers, ex-colleagues say nice things out of politeness — that's literally what The Mom Test is named after. Your mom will always think your idea is wonderful. Friendly interviews are fine for practicing the process, never for validating the direction. Validation needs strangers, ideally strangers currently suffering from the problem.
Pitfall 3: mistaking polite praise for validation. "This idea is so cool!" "I'll definitely use it when it launches!" — Fitzpatrick put it bluntly: compliments are the drug; stay away. The only valid validation is real money or real time spent: prepayment, signing up for a trial, referring a friend, agreeing to a follow-up in a month. Words don't count. Actions do.
Pitfall 4: using interviews as a substitute for shipping. This one is the sneakiest. Interviews can only falsify, never prove — they can tell you "this road is wrong" (nobody has this pain), but they can't tell you "this road is right" (people say it hurts, but will they pay?). Interviews are a filter, not a crystal ball. After interviewing, still build the smallest shippable version and look at real data.
Pitfall 5: asking "why" too much. Chain-asking "why" makes people invent reasons on the spot — humans don't actually understand their own motivations that well. Ask "what" and "how" for behavior; go easy on motives. If you really want the why, watch what they do (task observation), don't ask what they think.
Pitfall 6: no debrief after the interview. If you don't transcribe and extract verbatim cards within 24 hours, the interview might as well not have happened. Your brain auto-beautifies memory, keeping only the parts that support your idea. The dumb-but-effective countermeasure: make it a rule — no verbatim cards produced the same day, no next interview scheduled.
A Copy-Paste Interview Guide Template
Here's a complete guide template — copy it into your notes app and it's ready to use. Square brackets are facilitator notes; delete them during the actual interview.
[User Interview Guide Template] 30-minute version
Interviewee ID: ___ Date: ___ Product/direction: ___
[Opening, 0-5 min]
- Intro + purpose ("I want to understand how you normally handle XX")
- Ground rules: recording for internal notes only, can stop anytime,
small gift card afterward
- Warm-up: What do you do? Do you use our product?
[Behavior deep-dives, 5-20 min] (the core! They talk 80% of the time)
- How do you normally handle XX?
- When was the last time? How long did it take?
- Walk me through exactly what happened last time, from step one.
- What tool did you use? Why that one?
- Which step was the most annoying? Annoying exactly how?
[ask "exactly how" at least twice]
- Did it ever go wrong? What happened?
- Tried other approaches? Why give up?
- If this step could magically disappear, how much would you pay?
[If willing: Can you demo it for me right now?]
[On hesitation: You paused there — what were you thinking?]
[Pain confirmation, 20-27 min]
- Play back: "So what you're saying is..., right?"
[shut up, wait for corrections]
- If you could keep only one feature / fix one hassle, which?
[Wrap-up, 27-30 min]
- Anything I didn't ask that I should know?
- Know anyone else struggling with this? Intro please?
[get contact info on the spot]
- Can I come back in a month to share progress?
- Would you be our first tester if we build a solution?
[Within 24 hours after]
□ Transcribe recording □ Extract verbatim cards (one quote + ID + timestamp each)
□ Log into needs pool □ Send thanks + gift card
Closing: Interviews Are a Vibe Coder's Cheapest Moat
AI has driven the cost of "building it" to the floor, which means competition only gets fiercer — and the only axis left to compete on is "building the right thing." The information about what's right isn't in your head. It's in users' mouths — in the late nights they spent doing expenses overtime last week, in the moment they cursed at a pile of mismatched invoices.
Interviews won't hand you the answer key. They'll just help you make fewer mistakes. And for a team of one, avoiding one big mistake is worth more than nailing ten small features. Go book four conversations this month — start with question 18 (the referral). You'll find users far more willing to talk than you expected. After all, people who genuinely listen to their complaints are genuinely rare.
Related articles

You shipped invite links in a weekend. A month later: 200 invites sent, 3 signups back, 2 from your own alt accounts. A referral program isn't 'adding a link' — it's an economic system. From the R < CAC x L unit-economics formula: one-sided vs two-sided rewards, pricing four reward types (cash, discount, credits, feature unlocks), runnable Next.js attribution and anti-fraud code, 5 fraud-detection moves, a three-stage launch cadence, K-factor metrics, and failure-mode postmortems.

A waitlist is not a wishing well — it's a conversion funnel. This guide covers the four waitlist models and a decision table, a one-day Next.js + Supabase implementation (signup API, double opt-in, queue ranking, referral scoring), a week-by-week 30-day operating rhythm, anti-fraud tactics, five launch-day moves, metric thresholds for validation, and how to fix the classic failure of a long queue with nobody showing up.

Low traffic means you need experiment design more, not less. A field guide for vibe builders: the three myths (sample-size illusion, peeking, testing only UI colors), a one-sentence hypothesis template with north-star vs. guardrail metrics, a sample-size lookup table and run-length formula, a 30-line Next.js feature-flag middleware with a three-stage rollout, five classic traps with real crash stories, a PostHog/GrowthBook/DIY cost comparison, and a one-page experiment retro template.