Guide
How to Design a Reliable AI-Moderated Interview Study (What the Research Actually Says)
The reliability of an AI-moderated interview study is mostly decided before the first interview runs. The 2025 and 2026 research converges on six design principles: anchor every question to an explicit research question, brief the AI moderator to probe vagueness and stop at sufficiency, give every question enough time and silence to be answered properly, pilot your interviewer prompt and freeze the version that works, structure the analysis so the AI cannot pad its way to false confidence, and benchmark against a small human baseline while reporting your own limitations. Platforms matter, but a well-designed study on an average platform beats a careless study on the best one.
The biggest of the six inverts standard practice. Most teams brief their AI moderator to elicit “rich, detailed responses” and judge success by how much participants said. The best current evidence says that is the wrong target entirely.
Reliability is a design decision, not a platform feature
By the time your first participant clicks the interview link, most of your data quality is already locked in. The research questions you did or did not define, the prompt you wrote, the session length you set, and the analysis structure you chose will shape every transcript. AI-moderated interviewing itself is now well validated: randomised comparisons show AI interviewers holding structure and probing credibly against human moderators (Wuttke et al., 2025), and multi-model evaluations show adaptive questioning working across standardised conditions (Panfilova et al., 2026, Nature Scientific Reports). The method works. Whether your study works depends on the six decisions below.
Principle 1: Start from research questions, not topic lists
Write down the two or three decisions this study must inform, turn them into explicit research questions, and anchor every interview question to one of them. This is not tidiness. It is the single strongest lever on data quality the evidence has found.
In 2026, researchers at Johns Hopkins ran the largest validation study of interview quality metrics to date: 343 real interview transcripts, nearly 17,000 participant responses, 14 research projects (Ivey, Field and Xiao, 2026). They tested ten proposed quality measures against whether responses actually contributed to study findings. Relevance to a key research question was the strongest predictor. Clarity and informativeness, the metrics most teams optimise for, were not predictive at all.
The practical consequence: a topic list (“explore onboarding, pricing, support”) produces conversation. Research questions (“what causes users to abandon onboarding in week one?”) produce evidence. Put the research questions verbatim in your moderator’s instructions and tell it that every probe should move the participant closer to one of them. And stop briefing for richness. A participant on autopilot can produce three articulate paragraphs that answer nothing you needed answered.
Principle 2: Brief the moderator to probe vagueness and stop at sufficiency
A good AI moderator does two things a survey cannot: it pushes past a vague answer, and it knows when to stop. Both behaviours have to be designed in, because the defaults get them wrong in opposite directions.
The research on interview probes shows that theory-based follow-ups meaningfully change response quality, and that different probe types suit different goals (Jacobsen et al., CHI 2025). The working brief: when an answer is vague or general, ask one short follow-up that grounds it in a specific, recent, personal example. When an answer has clearly addressed the research question, move on. Endless probing past sufficiency does not produce more insight; it produces fatigue, and fatigued participants start answering to finish rather than to be understood.
The fastest way to check whether your moderator does this is also the simplest: read your pilot transcripts and watch what happens after a vague answer. One good probe, then progress, is what you want to see.
Principle 3: Give every question room
Session length, question count, and silence handling are data quality parameters, not logistics. Three numbers worth stealing: budget at least 90 seconds of session time per question, keep at least 40 percent of your questions open-ended, and do not let your moderator jump on pauses shorter than about eight seconds.
The reasoning is behavioural. When participants run out of time relative to the question count, their answers compress: shorter, faster, thinner across the closing turns of the session. When closed questions dominate, you collect transcript rather than insight. And when the AI fills every silence, it interrupts exactly the pauses where reflection happens. A participant who goes quiet after a hard question is usually thinking. A moderator that cannot tolerate that silence trains them to stop thinking.
If your question list will not fit the time budget, cut questions. Ten questions answered reflectively beat twenty answered in completion mode.
Principle 4: Pilot the prompt, read the transcripts, freeze the version
Never trust an interviewer prompt on inspection alone, because LLM behaviour is fragile under trivial wording changes. Alterations as small as spacing, punctuation, or reordered examples can shift an LLM’s task accuracy by up to 76 percent (Lin, 2025). The prompt that reads perfectly in a document can behave badly in conversation, and the edit that looks harmless can change everything downstream.
The protocol is cheap: run at least three pilot sessions before launch, read every transcript in full, and judge each one on a single criterion drawn from Principle 1: did this conversation produce responses relevant to the research questions? If yes, freeze that prompt version and record it. If no, revise and pilot again. Then do not casually edit the prompt mid-campaign, because a mid-campaign edit quietly splits your study into two studies with different instruments.
Principle 5: Structure the analysis so the AI cannot pad
AI-assisted analysis fails in a predictable direction: toward confident, verbose over-generalisation. LLM judges systematically prefer longer outputs and can disagree with themselves across runs (Kim et al., 2025). Left unconstrained, a summarisation pass will inflate thin evidence into fluent findings.
The countermeasures are structural. Define a fixed output structure per session, organised by research question, so summaries are comparable across the study. Require an explicit “not addressed” option for every field, so the model is never forced to invent an answer for a question the session did not cover. Demand counted evidence at the synthesis stage: how many sessions support a finding, how many contradict it, how many are silent. And run extraction at low temperature with tightly specified formats, which measurably improves alignment between AI output and the underlying human responses (Zhang et al., 2025). A finding that appears in three of forty sessions is an emerging signal, not a conclusion, and your analysis structure should make that distinction impossible to blur.
Principle 6: Baseline it, and report your own limitations
For your first campaign on a new population, run a handful of human-moderated sessions alongside the AI ones and compare. The randomised evidence says AI moderators hold structure and consistency extremely well and occasionally miss emotional beats a skilled human catches (Wuttke et al., 2025). Your population may differ. A small parallel baseline, even five sessions, tells you where.
Then apply the same honesty to your own reporting that you should demand from vendors. State the sample size. Flag the sessions where engagement degraded. Name the research questions that went unanswered. Internal stakeholders trust qualified findings more, not less, and a limitations section is what separates research from advocacy. If you want the full version of that standard, we wrote a companion piece on how to tell whether an AI interview accuracy study is reliable; the tests you should apply to a vendor are the tests you should pass yourself.
Where Tadeus fits
Tadeus was built around these principles rather than retrofitted to them. Every session ships with structured per-session metrics scored from the transcript and the audio signal itself, including relevance to the campaign’s research objective, the one metric the current evidence says actually predicts quality, alongside voice-level comfort signals and a six-dimension emotional profile (where permitted under the EU AI Act and similar regulation). Templates separate the live interviewer instructions from the analysis prompts and a structured output schema, so Principle 5’s constraints are the default rather than a discipline you have to maintain by hand. The next layer, in development now, is a behavioural quality score derived from response latency, word count compression across turns, and silence patterns, designed to flag completion-mode sessions to the researcher rather than quietly hide them.
The design checklist
Before you launch, check:
- Two or three explicit research questions, written into the moderator’s instructions verbatim?
- Moderator briefed to probe vague answers once, with a specific example, and to move on at sufficiency?
- At least 90 seconds of session budget per question, at least 40 percent open questions, no nudging on pauses under eight seconds?
- At least three pilot sessions run, every transcript read, prompt version frozen?
- Analysis structured per research question, with “not addressed” options and counted evidence at synthesis?
- A small human baseline for new populations, and a limitations section in your own report?
A study that ticks all six will produce data you can defend in front of the most sceptical stakeholder in the room. A study that skips the first and fourth will produce fluent transcripts and thin findings, and no platform can rescue it.
Ready to try it? Design your study with this checklist, then run the pilot on Tadeus and read every transcript before you scale. The method is the product.
Building with the Tadeus MCP
If an AI agent is building your study for you, give it this context first. Tadeus exposes template and campaign creation over MCP, which means Claude, ChatGPT, or your own agents can design and launch a study end to end. An agent with the right guidelines builds a dramatically better study than one improvising from tool schemas alone. Copy the block below into your agent’s instructions alongside your Tadeus MCP connection, then describe your research goal in plain language.
TADEUS STUDY DESIGN GUIDELINES (for AI agents using the Tadeus MCP)
You are designing an interview study on Tadeus. Follow these rules. They
are grounded in the 2025-2026 research on interview quality: relevance to
a research question is the only validated quality signal; LLM behaviour
is fragile under prompt changes; LLM analysis inflates thin evidence
unless structurally constrained.
DESIGN ORDER (work backwards from the decision):
1. Ask the user what 2-3 decisions this study must inform. Turn these
into explicit research questions (RQs). Confirm them before building.
2. Design the output_schema from the RQs.
3. Write the summary_prompt.
4. Write the interview_prompt last.
Then create the template with create_template (status "draft"), and a
pilot campaign with create_campaign.
INTERVIEW_PROMPT rules:
- Under 800 words total. State the RQs verbatim inside the prompt and
instruct: every probe should move the respondent closer to one RQ.
- At least 40-60 percent open-ended questions. One question at a time,
each under 20 words.
- Probe a vague answer once, anchored to a specific recent personal
example, then move on. Stop probing an RQ once it is addressed.
- If the respondent hesitates or shortens answers, make the next
question shorter. Never re-explain a question. Never fill silence.
- Budget at least 90 seconds of session time per question. If the
question list does not fit, cut questions, not time.
OUTPUT_SCHEMA rules:
- Flat JSON object, under 12 fields, every field mapped to an RQ and a
decision. Prefer enums and numbers over free text.
- Every enum includes "not_addressed". Every extraction field is
nullable. Never force the model to invent a value.
- Walk the schema against the interview_prompt: any field with no
eliciting question will come back null or invented. Fix the prompt
or cut the field.
SUMMARY_PROMPT rules (sees the full transcript, feeds Insights):
- Fixed headings per RQ: position in 1-2 sentences, one short verbatim
quote, confidence (high/medium/low). Plus UNEXPECTED and RELIABILITY
sections. Write "not addressed" where true. Be terse; never pad.
INSIGHTS_PROMPT rules (sees only summaries, never transcripts):
- Answer each RQ with counted evidence: sessions supporting,
contradicting, not addressing. Patterns in fewer than 5 sessions are
"emerging signals", not findings. Require the strongest
counter-evidence for each finding. Separate FINDINGS from
RECOMMENDATIONS. Close with LIMITATIONS.
PILOT PROTOCOL (mandatory):
- Keep the template in "draft". Create the campaign with max_sessions
set to 3-5 and run pilot sessions.
- Tell the user to read every pilot transcript against one criterion:
did it produce responses relevant to the RQs? Also verify the
structured output against transcripts for invented values.
- Only after the user confirms: set the template status to "published"
via update_template, raise or remove max_sessions via
update_campaign, and do not edit any prompt mid-campaign. A
mid-campaign prompt edit splits the study into two instruments.
Always confirm the RQs, the schema, and the pilot results with the
user before scaling. Recommendation is not to skip the pilot in draft
stage before moving to production.
The block encodes this page’s six principles as executable constraints, mapped to the actual MCP fields (interview_prompt, summary_prompt, insights_prompt, output_schema) so the agent builds the study the way the evidence says studies should be built. It also defines the discipline agents most reliably skip: the pilot, and the freeze.
Frequently asked questions
What is an AI-moderated interview?
A qualitative research interview conducted by an AI moderator, usually by voice or chat, that asks your questions, adapts follow-ups to what the participant says, and produces transcripts and structured analysis at a scale human moderation cannot match. Randomised comparisons show AI moderators holding interview structure and probing credibly against human interviewers (Wuttke et al., 2025).
How many questions should an AI-moderated interview have?
Fewer than you think. Budget at least 90 seconds of session time per question and cut the list to fit, keeping at least 40 percent open-ended. Participants who run out of time compress their answers in the final turns, which quietly degrades exactly the data you launched the study to collect.
How do I know if my AI moderator is asking good follow-ups?
Read your pilot transcripts and watch what happens after a vague answer. A good moderator asks one short probe anchored to a specific, recent example, then moves on once the research question is addressed. Probe types measurably change response quality (Jacobsen et al., CHI 2025), so this behaviour is worth designing deliberately rather than leaving to defaults.
Can I trust AI-generated analysis of my interviews?
With structure, largely yes; unconstrained, no. LLM analysis drifts toward verbose over-generalisation and inconsistent judgments (Kim et al., 2025). Fix the output structure per session, require “not addressed” options, demand counted evidence at synthesis, and spot-check a sample of transcripts yourself.
How big should my pilot be?
Three to five sessions before launch, with every transcript read in full against one criterion: did this conversation produce responses relevant to the research questions? That criterion comes straight from the strongest evidence on interview quality (Ivey, Field and Xiao, 2026). Freeze the prompt version that passes and do not edit it mid-campaign.
References
- Ivey, J., Field, A., Xiao, Z. (2026). What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews. arXiv:2604.05163
- Lin, Z. (2025). From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology. arXiv:2506.16697
- Kim, E. et al. (2025). LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation. arXiv:2412.10424
- Wuttke, A. et al. (2025). AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers. arXiv:2410.01824
- Panfilova, A. et al. (2026). The AI Interviewer: Multi-Faceted Evaluation of Adaptive Questioning by Large Language Models. Scientific Reports
- Jacobsen, R.M. et al. (2025). Chatbots for Data Collection in Surveys: A Comparison of Four Theory-Based Interview Probes. CHI 2025
- Zhang et al. (2025). Leveraging Interview-Informed LLMs to Model Survey Responses. arXiv:2505.21997
Give every employee a voice
Tadeus holds real voice conversations with your whole workforce at once and returns structured records to the systems you already use.
Start for free