Can AI Replace User Research?

Can AI Replace User Research?
Open the last set of user stories your team wrote and try to trace one back to its source.
If you can name the customer who said it, the ticket it came from, or the session where someone got stuck, that story is evidence.
When the trail ends at a prompt, you have a guess in the format of a finding.
AI got very good at producing the artifacts of research. It writes clean personas, plausible quotes, and journey maps that look like three weeks of team effort.
We use AI across our own process and it does real work in discovery. It cannot replace the research itself, because it has no way to know where your customer got stuck in your portal or why your ops team stopped using the tool you paid for.
The trouble starts when a team uses it to fill in research they never did.
What can AI actually do in user research?
AI does its best work in research after the evidence already exists.
Once you have interview notes, support tickets, call recordings, and survey responses sitting in a folder, AI is genuinely good at making that pile usable:
- Clustering themes across hundreds of tickets
- Summarizing complaints that repeat across departments
- Cleaning up messy session notes into something readable
- Drafting first-pass categories for a team to argue with
- Comparing what sales hears against what support hears
- Turning raw notes into a format a workshop can run on
A review of twelve peer-reviewed papers published this April lands in the same place, concluding that AI participants may be practical for deriving insights from data already collected rather than for generating new research.
That is a narrow job and a genuinely useful one, because synthesis is the part of discovery that used to take three weeks.
What counts as real evidence in product discovery?
Real evidence is anything a person can trace to a source.
A finding should point back to a quote, a ticket number, a recorded session, a metric, or something a person on your team watched happen:
- User interviews and usability sessions
- Support tickets and call logs
- Sales calls and customer emails
- Watching someone do the work in their actual environment
- Analytics on what people do rather than what they say
- The workarounds your own team built to survive the current system
Assumptions are allowed in this business, and every project runs on a few.
They just have to be labeled as assumptions before they shape a build. We believe that business is built on transparency and trust, and that good software is built the same way.
That starts with a plan that says out loud which parts are verified.
Are synthetic users reliable for research?
Synthetic users fail at the exact thing research is for.
Across the twelve papers reviewed, only nine of twenty-three findings were encouraging, and the failures land in places that matter: synthetic responses show less variance than real people, they exaggerate effects, and the qualitative narratives come back shallow and stripped of context.
Among 150 research professionals surveyed this May, 97% use AI somewhere in their workflow for synthesis, transcription, and coding, while only 8% regularly use tools that generate synthetic participants. Exactly one respondent out of 150 called synthetic users valid on their own.
The concern that ranked highest after accuracy was stakeholders over-trusting AI-generated findings, at 80%.
The stakeholder those researchers are worried about is usually the person approving the build.
Why does a bad assumption cost more in custom software?
A bad assumption costs more in custom software because it stops being a slide and becomes structure.
A weak assumption in a marketing campaign produces a message nobody responds to, and you write a new one.
The same assumption inside a custom build gets encoded into:
- Workflow logic and the exceptions it does or does not handle
- Status labels your customers will read
- Permission models and who sees which record
- Reporting definitions that two departments will later disagree about
- Approval paths and notification timing
- The data structure underneath all of it
Unwinding any of those after launch costs a multiple of what checking the assumption would have cost during discovery.
A status label is an afternoon of copywriting before the build and a data migration afterward.
A permission model that assumed one kind of user is a schema change plus a security review once real accounts are attached to it.
We run a free two to four hour Exploratory before writing a proposal so these questions get asked while the answer is still a whiteboard sketch.
Our design process starts the same way, getting deep into your business, audience, and goals before anyone opens a design file.
How should you use AI in the discovery process?
Use AI in the middle of the process instead of at the front of it.
The sequence that works:
- Collect the evidence first. Interviews, tickets, call notes, analytics, and time spent watching people use what they have now.
- Let AI organize it. Cluster the themes, summarize the repeats, compare across user groups.
- Trace every insight to a source. If a finding has no quote, ticket, or observed behavior behind it, mark it as a hypothesis.
- Review with the people closest to the work. Support, ops, and admins will tell you which summary is missing the context that makes it wrong.
- Validate before it becomes a requirement. A follow-up interview or a prototype walkthrough costs a day and saves a sprint.
Step four is where a generated summary meets its first real correction.
Your support team can tell you inside a minute which theme matches what they hear all day.
For RTI we designed the MDI Align mobile app around a complex daily health regimen, mapping the screen architecture, education content, and survey flows. Focus group testing informed iterative refinements throughout the process, and the client reported a 100% success rate in their pilot study.
Those refinements came from watching real people work through a real regimen.
What should AI never generate in user research?
AI should never generate anything meant to be a fact about a real person. Keep it away from:
- User quotes and customer emotions
- Observed behaviors and usage patterns
- Support complaints that nobody filed
- Workflow exceptions and the business rules behind them
- Accessibility needs
- Product requirements and decision criteria
Use AI to write up a finding after a person has gone and gotten it.
What should you ask before an insight becomes a requirement?
An insight is ready to become a requirement when it survives six questions:
- Where did this come from? Name the ticket, the session, the call, or the metric.
- Did a user say it, do it, or show it? What someone did is stronger evidence than what they told you they would do.
- Is this a finding or a hypothesis? Label it. Projects run on both, and the label is what keeps a budget honest.
- Who closest to the workflow has read this? Your support lead will catch what a summary flattened.
- What decision does it affect? Higher stakes should demand stronger evidence.
- What still needs verifying? Thin evidence earns a follow-up conversation before it earns a line in the backlog.
Do you have enough evidence to make this decision?
Stop asking whether AI can produce the research artifacts faster, because it can and the artifacts were never the point.
Ask whether you have enough real evidence to make the decision in front of you.
Take the requirements doc for whatever you are building next and mark every line with its source.
The lines you cannot mark are your verification list for this week.
Related Articles
Here are a couple related articles to view, or return back to the main page.


Check out the BIZ/DEV podcast
Our weekly tech podcast focusing on AI, our industry, the founder's journey, and more.
