In this article
TL;DR — Validating a startup idea is not scoring it. It means finding evidence that the problem exists for people other than you, and that they have already spent money or hours trying to escape it. Across four fixed WedgeScout research runs collected in July 2026, 67 of 198 evidence units came from sources labelled vendor marketing (34%). A 78/100 hides that ratio. A source trail shows it.
Idea validation is the process of testing whether a problem is real before you build a product for it. Most guides give you a checklist. This one walks four real ideas through four questions: one held up, and three died at different points. Every quoted source below is a fixed report citation, not a live market claim.
The four ideas came from our own research runs. Read the public solo-SaaS report to inspect a longer claim-to-source trail, including evidence that cuts against the product idea.
Idea validation, defined
Idea validation is the work of testing whether a problem is real before you build a solution for it. Two words carry the definition: evidence and already. QuickBooks, Keeper, and r/Bookkeeping appear below because validation without names is just opinion with paragraphs.
Evidence is not encouragement, a survey response, or a paragraph a model generated. Already means somebody described a past behaviour: the spreadsheet they maintain, the contractor they hired, or the tool they pay for and complain about.
That distinction separates validation from adjacent work. Customer discovery is one way to learn from people. TAM, SAM, and SOM sizes a market you have already established exists. Neither converts an unsupported conclusion into evidence.
What four questions does validation have to answer?
Validation has to answer four questions in order, and an idea can die at any one of them:
- Who has this problem besides me? Find people who described it, unprompted, in their own words.
- Have they already spent money or hours escaping it? A workaround points to a cost that already exists.
- Can you find them in one place? If you cannot name where the segment gathers, you cannot interview, reach, or sell to it. That is a finding, not a delay.
- Who is already selling to them, and where does it stop? No competitors is usually bad news. Ask which part of the job current tools already do and which part they leave on the floor.
The rest of this guide applies those questions to four real ideas. The order matters: each failure below exposes what a particular question is actually testing.
What does a validated idea look like? A worked example
A validated idea looks like a chain you can click. One July 2026 run explored explainable expense-categorization exceptions for small accounting firms across r/Bookkeeping, r/quickbooksonline, and r/smallbusiness. It is the only example here that survived all four questions—and it still earned only one more test, not a build decision.
Who else had this problem?
The problem belonged to small-business bookkeepers. One r/Bookkeeping poster described the constraint before anyone asked:
“Say I have a client who runs 200–300 transactions per month. Many of these are gas stations and convenience stores, travel, restaurants, Home Depot, Amazon, etc. I feel like it's unrealistic for him to give me information on every single receipt.”
The number is not market size. It is one firm's caseload. What it carries is a named constraint: the bookkeeper needs information, but the client's time is scarce.
Had they already spent money or hours on it?
Yes. The same person described a workaround and its failure:
“I've tried sending a spreadsheet to my client but it gets ignored because it is too long and he probably thinks that I am dumb if I don't understand that restaurants are meals.”
That sentence contains a workaround, its failure mode, and a social cost. It is stronger than a generic agreement that the task is annoying.
Could we find them in one place?
Yes. r/Bookkeeping, r/quickbooksonline, and r/smallbusiness carried related complaints in different vocabularies. A reachable segment has rooms where practitioners already ask peers for help; this is one of the cheapest questions to answer.
What does QuickBooks already do—and where does it stop?
QuickBooks Online already does much of this job. It has bank rules and recurring-transaction categorisation; Xero and Sage offer equivalents. The question is not whether an incumbent exists, but which part of the workflow still breaks.
One QuickBooks Online user made the distinction explicit:
“I know QBO bank rules can speed up categorization … But my recurring pain isn't the category, it's job/customer assignment.”
That evidence narrowed our own hypothesis. The strongest finding argues against building a new categorisation tool. On Keeper, a purpose-built alternative, a bookkeeper wrote:
“I've heard of Keeper and such but you need to have a client that is willing to keep up with it.”
Keeper is not necessarily the blocker; client compliance may be. Evidence that changes the idea is the point of collecting it.
What the r/Bookkeeping run does not establish
The run does not establish willingness to pay, segment size, or a competitive gap. It accepted 12 sources and rejected 7 from the three communities above. Its strongest counter-signal is the Keeper line: a verdict of worth pursuing means run one more test, not start building.
What do ideas look like when they fail?
Ideas fail at specific questions, not in general. Three runs—r/Laundromats, r/pressurewashing, and r/ProductManagement—failed for three different reasons.
Laundromats: real pain, explicit refusal to pay
The pain was real. A r/Laundromats source said:
“Laundry owners tend to work on extremely tight margins … and balk heavily at having to pay a trip charge.”
This kills the idea at question two: pain exists, but there is a documented refusal to pay for relief. A product here may need to be free at the point of use or bundled into something already bought.
The run also found The Laundry Boss pricing page: “One flat fee per machine. No hidden costs.” That vendor page is evidence that The Laundry Boss sells a product, not evidence that laundromat owners want it.
Pressure washing: the pain was not the pain we assumed
The hypothesis was weather-driven rescheduling and small-crew coordination. The evidence pointed to a harsher issue. A mobile detailer in r/Detailing wrote:
“The week leading up to July 4th, I was picking up steam … and now 8 days of rain in the forecast.”
“I have been in business for 11 years and it's becoming extremely difficult to make a living from mobile detailing.”
An r/Entrepreneur operator added that living on pressure-washing income would be hard without year-round weather. The weather is not merely a scheduling problem software can solve; it is part of the business model. This run's judgment was weak because independent evidence for the original rescheduling hypothesis was missing. The sharpest quote came from an adjacent trade, so it is a real adjacent finding and weak direct evidence.
AI meeting notes: the workaround is good enough
This idea failed at question four. People had workarounds and were not complaining about them. A r/projectmanagement source said:
“Everyone has 48 hours to suggest corrections/additions to it in case anything is missed or if there are misunderstandings.”
A r/ProductManagement source said:
“I keep the email in my inbox till the tasks are complete or I no longer need to reference the info.”
A workaround people are content with is a hard competitor: it costs nothing and does not require adoption. Category saturation may matter too, but the evidence here supports the good-enough workaround first.
How do you do market research on Reddit?
Reddit works for market research because r/Bookkeeping, r/Laundromats, and r/pressurewashing were not written for your roadmap. That distance reduces leading questions; it does not make one post a market.
Find the rooms before the keywords
Start where the buyer already asks peers for help. For the accounting run, that was r/Bookkeeping and r/quickbooksonline—not r/Entrepreneur or r/startups. The specific room's vocabulary is what you will need later for interviews and ads.
Search for the workaround, not the complaint
Search for “spreadsheet,” “manual,” “still doing this by hand,” “I built a script,” and named incumbents. Problem words alone often return beginners asking what the problem is. The accounting quote above surfaced through a workaround-oriented query, not simply “expense categorization.”
Grade specificity, not volume
Ten posts that name a tool, step, cost, and failed workaround tell you more than one hundred vague complaints. Do not compare raw complaint counts between communities of different sizes; a count measures subreddit size as much as pain.
Log counter-signals in the same file
The Keeper line changed the accounting-run verdict. Put such evidence beside the support, not in a discard pile. Use the market-research template to keep the source, quote, context, interpretation, and limit together.
Reddit is not the only room. In the four fixed runs, 107 evidence units came from Reddit, 83 from the open web, and 8 from Hacker News. Reddit is a majority, not a monopoly.
The evidence hierarchy, ranked by whether you were in the room
Evidence is strongest when you were not in the room shaping it. Friends want to support you; surveys reflect your phrasing; interviews reveal sequence and context, but you are still present.
| Evidence source | Were you in the room? | What it proves | Trust |
|---|---|---|---|
| Asking friends and family | Yes — and they love you | Little beyond support | Lowest |
| Surveys | Yes — you wrote the questions | Your phrasing | Low |
| Customer interviews | Yes | What they remember doing | Medium |
| Unprompted public complaints | No | What they care about when nobody is selling | High |
| Money already spent | No | A past willingness to pay | Highest |
There is another category the table cannot safely rank: evidence written by someone selling the solution. In the accounting run, an accepted r/Bookkeeping post read:
“Our accounting firm … is developing AI in data entry and expense categorization … If you want to be an early adopter and get a great price and service—simply contact us.”
That is an advertisement inside a practitioner forum. It was labelled vendor_marketing. Across the four runs, 65 accepted sources included 18 vendor-marketing sources (28%); 67 of 198 evidence units traced back to a vendor source (34%). A tool must show that proportion rather than quietly count it as “the market talking.”
Can you validate an idea with a ChatGPT prompt?
A ChatGPT prompt can give you directions and cannot give you evidence. A model can produce a persona, competitor list, market analysis, and score from priors without retrieving anything a person said. It looks like research because research has a familiar shape.
What prompts are genuinely good for
Prompts can turn a vague idea into testable directions, generate unfamiliar search phrasing, and draft interview questions for checking against The Mom Test. In the accounting run, model-generated workaround phrasing helped find a strong source; the source—not the prompt—earned the claim.
What prompts cannot establish
Prompts cannot establish whether a specific person paid, switched, or complained. The test is simple: can you click the claim and reach a human who wrote it? If not, it is a hypothesis wearing the costume of a finding.
Want the source trail without a week of reading? Scout one direction with WedgeScout →. Each report should link its claims to the public posts that support or limit them.
Why is a score the wrong output?
A score is the wrong output because it is not falsifiable. A 78/100 does not show which evidence produced it, what would change it, or what you should do differently at 78 rather than 71.
It also merges two questions with different answer types. Does the problem exist? is researchable. Can you win it? depends on distribution, timing, and execution. A tool that folds them into one number is guessing about you.
A validation conclusion needs a visible chain from verdict to claim to evidence to original post, plus a flip condition: the finding that would make you stop, narrow the segment, or change the hypothesis.
Where our own method falls short
Our own pipeline consistently fails question four: competitive gaps. Across nine runs, not one established a competitive gap. The failure is mechanical: competitor queries run before practitioner sources reveal competitor names, so named alternatives in the evidence never receive a second, specific research pass.
QuickBooks, Keeper, Uncat, Xero, and Ramp appeared in collected evidence without being queried by name. The required fix is a second competitor-query round seeded from names discovered in the first. It is not built yet.
Two other limits matter. Of 198 evidence units across four runs, 50 could not be re-verified against their source at delivery time. Of 51 rejected sources, 33 were dropped for source-pool capacity rather than irrelevance. Rejection count is therefore not a quality score. These limits do not erase the quotations above; they bound what the runs support.
Validation ends on a condition, not a date
Validating an idea should take days, not months, and end on a condition rather than a calendar date. Stop when new threads and conversations stop adding surprise. For a narrow segment, that often means roughly 15 to 30 interviews per persona; larger research programs may go much further. NSF I-Corps, for example, requires National Teams to complete at least 100 potential-customer interviews during its seven-week program (NSF I-Corps).
Stop earlier when the segment is unreachable or the strongest sources say the problem is tolerable. Continue only while a specific uncertainty could still change the next decision.
What we still don't know
Several questions raised by these runs remain open because they need measurement, not better prose.
Does a documented refusal to pay ever reverse? The laundromat trip-charge quote does not tell us whether the same relief would sell when bundled into an existing machine or software invoice. Only a price test can.
How stable is one report across two runs of the same idea? Live source collection can return different evidence. We have not measured agreement between repeated runs.
What share of “no evidence found” is absence versus search failure? An empty counter-signal query may mean nobody said it, or that the query was poor. We cannot yet separate those explanations.
Does adjacent-industry evidence predict anything? The pressure-washing run's sharpest quote came from mobile detailing. We label and discount it; we have not measured whether such evidence is predictive.
Do ideas that die at question four differ from ideas that die at question one? Four runs are not a sample. Ask again at forty.
What can't validation tell you?
Validation cannot tell you whether you will win. It can establish that a problem appears real and that people spend time or money avoiding it. Pricing, timing, platform changes, and your own ability to execute are outside that evidence.
It also cannot convert a proxy into a fact. A community's membership is not budget. A complaint about an incumbent is not proof that a replacement will be adopted; Keeper illustrates that trap.
FAQ
How do you validate a startup idea?
You validate a startup idea by answering four questions with evidence you did not create: who else has the problem, what they already paid or did to escape it, where they gather, and where existing alternatives stop.
Can you validate a business idea for free?
You can validate a business idea for free. Reddit threads, reviews, and Hacker News comments cost nothing. The costly input is attention: reading the evidence that argues against you.
What is the strongest evidence for a startup idea?
The strongest evidence is a past behaviour that cost somebody something: a tool they pay for and complain about, a manual process they maintain, or a prepayment. A maintained spreadsheet beats a survey response.
How many people do I need to talk to?
You often need roughly 15 to 30 people per persona, but surprise is the stop condition, not the number. Continue until new conversations stop changing your understanding of the workflow.
Should I validate the problem or the solution first?
Validate the problem first. A solution test before you know who has the problem, how they handle it, and what would make them switch mostly measures politeness toward your proposal.
Is no competition a good sign?
No competition is usually a bad sign. It can mean people will not pay to fix the pain. The laundromat case shows real pain paired with an explicit refusal to pay a trip charge.
What should I do after validating an idea?
Pick the smallest test that addresses the largest remaining uncertainty. If you have pain but no spending, test a paid concierge version. If people spend but are satisfied with QuickBooks or Keeper, test a switch trigger before building a replacement.
The short version
Validating a startup idea means finding evidence that exists whether or not you are in the room. Ask four questions in order: who has the problem, what they already did to escape it, where they gather, and where incumbents stop. An idea dies at a specific question; identifying that question is the output a score cannot give you.
Start a WedgeScout research run →