Multi-to-One Interview Paradigm
- Multi-to-One Interview Paradigm is a structured approach coordinating various questioning agents or modalities to elicit a comprehensive single-target assessment.
- It spans applications in clinical interviews, survey research, and NLP evaluations, ensuring depth, reduced bias, and improved feedback mechanisms.
- Empirical studies show that adaptive question sampling and multi-agent collaboration enhance diagnostic transparency, efficiency, and overall performance metrics.
Searching arXiv for the cited papers to ground the article in the primary sources. The Multi-to-One Interview Paradigm denotes a family of interview-centered research designs in which multiple sources of questioning, guidance, evidence, or evaluation are coordinated around a single interviewee or a single interview-level judgment. In current arXiv usage, the expression spans several non-identical operationalizations: multiple collaborating AI agents interviewing one human participant; multiple interviewer models evaluating one MLLM; multi-turn interviewer–interviewee interaction with feedback and follow-up questions; and aggregation of multiple responses and multiple modalities into one candidate-level assessment (Bi et al., 25 Apr 2025, Shen et al., 18 Sep 2025, Kim et al., 2024, Li et al., 30 Jul 2025). The paradigm therefore functions less as a single algorithm than as a general design pattern for structured elicitation, adaptive probing, and holistic assessment under constraints of rigor, efficiency, fairness, or scalability.
1. Conceptual scope and definitional variants
A narrow definition appears in psychiatric assessment, where the Multi-to-One Interview Paradigm is “a scenario where multiple collaborating AI agents conduct a structured, clinical-grade interview with a single human participant” (Bi et al., 25 Apr 2025). A social-science definition emphasizes the classical setting “where multiple interviewers collectively engage a single interviewee” in order to “maximize depth, minimize individual interviewer bias, and triangulate findings via multiple perspectives” (Wuttke et al., 2024). In MLLM evaluation, the paradigm is instantiated by “interacting with multiple ‘interviewer’ models” to assess one model (Shen et al., 18 Sep 2025). In multimodal interview assessment, it is operationalized as integrating “multiple interview responses” and “multiple data modalities” into “a comprehensive performance judgment” for each candidate (Li et al., 30 Jul 2025).
| Study | Interviewee or target | Operationalization |
|---|---|---|
| (Bi et al., 25 Apr 2025) | Single human participant | Four specialized agents conduct a structured clinical interview |
| (Wuttke et al., 2024) | Survey respondent | Multiple interviewers or AI substitutes/complements engage one respondent |
| (Shen et al., 18 Sep 2025) | One MLLM | Multiple interviewer models ask adaptively selected questions |
| (Kim et al., 2024) | One LLM | One interviewer LLM conducts multi-turn evaluation with feedback |
| (Li et al., 30 Jul 2025) | One candidate | Six responses and three modalities are aggregated into final scores |
| (Varshney et al., 2021) | One NLP system | Multi-stage interaction supplies clarifications, clues, and examples |
Taken together, these usages indicate that the “multi” in multi-to-one may refer to interviewer multiplicity, agent-function multiplicity, response multiplicity, or modality multiplicity. A plausible implication is that the literature treats the interview not merely as a conversational form, but as an adaptive measurement protocol.
2. Agent-decomposed structured interviewing in psychiatric assessment
The most explicit multi-agent formulation is MAGI: Multi-Agent Guided Interview for Psychiatric Assessment (Bi et al., 25 Apr 2025). MAGI transforms the Mini International Neuropsychiatric Interview into “automatic computational workflows through coordinated multi-agent collaboration” using four specialized agents: “an interview tree guided navigation agent,” “an adaptive question agent,” “a judgment agent,” and “a diagnosis Agent generating Psychometric Chain-of- Thought (PsyCoT) traces that explicitly map symptoms to clinical criteria.”
The Navigation Agent manages dialogue flow and interview progression, directs traversal of the MINI interview decision tree, enforces clinical logic, and refuses to let the conversation skip or prematurely exit required nodes. The Question Agent generates the actual questions, converts clinical items into adaptive and contextually tailored probes, and can switch between diagnostic probing, explanatory rephrasing, and empathetic support. The Judgment Agent evaluates incoming responses for clinical adequacy, maps natural-language responses to operationalized MINI criteria, and can trigger recursive clarifications or forced-choice questions when uncertainty persists. The Diagnosis Agent aggregates symptoms across the session, applies DSM-5 and MINI logic, and produces PsyCoT reasoning traces (Bi et al., 25 Apr 2025).
The workflow is explicitly sequential. The Navigation Agent selects the next node; the Question Agent formulates or reformulates a query; the participant answers; the Judgment Agent determines whether evidence is sufficient; and the Diagnosis Agent synthesizes the final output once state completion is reached. This architecture is designed to address “Risk of Protocol Drift,” “Adaptive Conversational Engagement,” “Diagnostic Transparency,” and “Ambiguity Handling” (Bi et al., 25 Apr 2025).
MAGI also includes a clinically interpretable rule representation. For depression, the diagnosis agent uses the explicit criterion:
Experimental results are reported on “1,002 real-world participants covering depression, generalized anxiety, social anxiety and suicide.” Human experts evaluated interviews for “relevance,” “accuracy,” “completeness,” and “guidance.” MAGI achieved 3.576 relevance, 3.503 accuracy, 3.497 completeness, and 3.460 guidance, compared with 3.506, 3.463, 3.350, and 3.294 for Direct LLM, and 3.519, 3.475, 3.369, and 3.313 for MINI. The paper further reports that MAGI’s PsyCoT improved difficult categories such as suicide risk, with “kappa increases from 0.363 to up to 0.942,” and that “Case F1 for depression rose by up to 41%” (Bi et al., 25 Apr 2025).
3. Conversational interviewing in survey research and human-subject elicitation
In survey research, the paradigm is tied to the longstanding trade-off between depth and scale. AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers examines whether LLMs can replace human interviewers in a controlled setting using “identical questionnaires on political topics” (Wuttke et al., 2024). University students were randomly assigned to an AI interviewer or human interviewers, interviews were monitored in real time, recorded for transcript-based evaluation, and followed by participant experience surveys.
The paper reports a mixed profile of strengths and weaknesses. LLM interviewers “followed structured protocols” and could be prompted for “active listening, paraphrasing, probing follow-ups” with high consistency. One quantitative result is that “AI committed only 6% of active-listening guideline violations, vs. 94% for humans.” At the same time, “AI interviewers: Avg. 72 violations per interview” and “Human interviewers: Avg. 64 violations per interview,” indicating that overall guideline adherence problems remained in both conditions (Wuttke et al., 2024).
On participant engagement, AI interviews elicited “longer responses (52 vs. 32 words on avg; +62% in AI).” Input mode mattered: “Spoken (audio) input yielded even longer AI-collected responses but at times less structured/elaborate; text input was more concise and deliberate,” with “Flesch scores: audio 48.32, text 77.66, human interviews 62.” Participant feedback was similar across AI and human interviews in clarity, understanding, and satisfaction, but AI interviews scored lower on “interestingness” and willingness to repeat, “likely tied to technical hitches” (Wuttke et al., 2024).
This work also reframes multi-to-one interviewing as a possible hybrid design. The paper states that LLMs can complement multi-interviewer panels by “acting as an ‘additional’ interviewer,” generating initial probing for later human elaboration, or conducting first-pass interviews that flag cases for deeper human-led exploration. It also states that “Full replicability of MOP’s organic, multi-expert probing and triangulation is not yet fully achieved,” and that “Human oversight, prompt iteration, and critical user interface choices remain crucial” (Wuttke et al., 2024). This suggests that, in human-subject interviewing, multi-to-one systems are being studied as both substitutes for and augmenters of traditional interviewer groups.
4. Dynamic interview-based evaluation of NLP systems and LLMs
A distinct research line uses the interview paradigm not to assess humans, but to evaluate models through staged or feedback-driven interaction. Interviewer-Candidate Role Play: Towards Developing Real-World NLP Systems presents a “multi-stage task that simulates a typical human-human questioner-responder interaction such as an interview,” explicitly allowing “seeking clarifications about the question, taking advantage of clues, abstaining in order to avoid incorrect answers” (Varshney et al., 2021). The stages are cumulative: raw input; “semantic-preserving simplifications”; “knowledge statements”; “similar labeled examples”; and “similar unlabeled examples.” Progression is confidence-driven, and the model may abstain after all support if confidence remains insufficient.
The reported gains are specific to out-of-domain performance. The paper states that the multi-stage formulation improves OOD generalization performance “up to 2.29% in Stage 1, 1.91% in Stage 2, 54.88% in Stage 3, and 72.02% in Stage 4 over the standard unguided prediction” (Varshney et al., 2021). The design treats the interview as incremental support rather than a fixed test.
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation generalizes this idea to multi-turn evaluation of LLMs (Kim et al., 2024). In that framework, “the LLM interviewer actively provides feedback on responses and poses follow-up questions to the evaluated LLM.” At the start of the interview, the interviewer “dynamically modifies datasets to generate initial questions, mitigating data contamination.” The process includes “Problem Solving,” “Feedback & Revising,” “Follow-up Questioning,” and “Iterative Rounds,” and ends with an “Interview Report” that aggregates performance scores, error types, transcript examples, and a “comprehensive analysis of the LLM’s strengths and weaknesses” (Kim et al., 2024).
The paper attributes several advantages to this interview format. It reports that the framework provides insight into “the quality of initial responses, adaptability to feedback, and ability to address follow-up queries like clarification or additional knowledge requests.” It also states that the approach addresses “verbosity bias,” “inconsistency across runs,” “self-enhancement bias,” and “data contamination” more effectively than conventional static methods (Kim et al., 2024).
5. Efficient MLLM benchmarking through multiple interviewer models
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation makes interviewer multiplicity explicit by evaluating an MLLM “through curated questions” posed by multiple interviewer models (Shen et al., 18 Sep 2025). The paper is motivated by the inefficiency of “conventional full-coverage Question-Answering evaluations,” which are said to suffer from “high redundancy and low efficiency.”
Its framework contains three components: a “two-stage interview strategy,” “dynamic adjustment of interviewer weights,” and “adaptive mechanism for question difficulty.” In the pre-interview phase, a small number of moderate-difficulty questions are sampled to determine an initial difficulty level using the thresholded rule
During the formal interview, interviewer selection is probabilistic and depends on dynamically updated weights. Each interviewer starts with , and the update rule is
with in experiments and clipping to . Difficulty is also updated adaptively:
If the level oscillates between two values for three times, a jump rule is applied, and if no questions exist at the current level, a separate fallback update is used (Shen et al., 18 Sep 2025).
The evaluation uses MMT-Bench, ScienceQA, and SEED-Bench, with SRCC, PLCC, and KRCC as ranking-correlation metrics. The paper reports that, on all three benchmarks and under question counts of 20, 30, 50, 80, and 100, the paradigm achieves higher SRCC, PLCC, and KRCC than random sampling. The reported average improvement is “up to 17.6% in PLCC and 16.7% in SRCC,” while reducing the number of required questions. A concrete example is “At 50 questions on SEED-Bench: SRCC of 0.7377 (vs. 0.5577 in random),” and another is “At 30 questions on ScienceQA: SRCC 0.6289 (vs. 0.3668 for random), a 71.5% relative improvement” (Shen et al., 18 Sep 2025). This formulation makes the interview an adaptive sampling mechanism for benchmarking rather than a fixed question list.
6. Multi-response, multimodal aggregation and adjacent interview mechanisms
In automated interview performance assessment, the multi-to-one paradigm can denote aggregation rather than multiple interviewer identities. Listening to the Unspoken: Exploring 365 Aspects of Multimodal Interview Performance Assessment defines the paradigm as “assessing each candidate by integrating information from multiple interview responses (six questions per candidate) and multiple data modalities (video, audio, text) to yield a comprehensive performance judgment across five evaluation dimensions” (Li et al., 30 Jul 2025). The five dimensions are “integrity, cooperation, social versatility, development orientation, overall hireability.”
The architecture processes six responses independently but identically, extracts modality-specific features, and fuses them through the Shared Compression Multilayer Perceptron (MSCMLP). Per response, 32 parallel regression heads predict the five target scores, and mean-pooling is applied both across heads and across responses. The final prediction is given by
where responses and regression heads. Training uses standard mean squared error. The paper states that this multi-level aggregation “reduces sensitivity to individual response anomalies or noisy data” and “facilitates fair, representative, and stable scoring.” The reported result is “a multi-dimensional average MSE of 0.1824,” and the framework “secured first place in the AVI Challenge 2025” (Li et al., 30 Jul 2025).
A neighboring but distinct line of work studies interview coordination in centralized markets rather than interview interaction. Interviewing Matching in Random Markets analyzes a random market in which candidates interview with potential employers before a clearinghouse produces a stable match (Allman et al., 2023). The paper studies “a novel many-to-many interview match mechanism” and reports that with a limit of 0 interviews per candidate and per position, “the fraction of positions that are unfilled vanishes quickly with 1,” “the ex post efficiency grows rapidly with 2,” and sincere pre-interview reporting is “an 3-Bayes Nash equilibrium.” This is not a multi-to-one interview paradigm in the conversational sense, but it places interview allocation within a broader theory of coordinated pre-match interaction (Allman et al., 2023).
Across these literatures, several recurring concerns appear. Clinical systems emphasize “protocol drift,” “ambiguity handling,” and “diagnostic transparency” (Bi et al., 25 Apr 2025). Survey systems emphasize “prompt sensitivity,” “interface design,” and “technical barriers (latency, transcription errors)” (Wuttke et al., 2024). Model-evaluation systems emphasize “verbosity bias,” “data contamination,” and stability across runs (Kim et al., 2024). Benchmarking systems emphasize interviewer weighting and difficulty adaptation for “fairness” and “coverage” (Shen et al., 18 Sep 2025). A common implication is that the multi-to-one interview paradigm is best understood as a structured response to the same underlying problem: how to probe one target deeply and efficiently without sacrificing rigor, comparability, or adaptability.