Skip to content
AI Atlas
PaperActive

Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training

arxiv.org/abs/2609.10939

quality89

Updated 2 h ago · first seen 12 Sept 2026

paper_01M29X34TBB2XEKAMPTQQX9HV3

Published
12 Sept 2026
T1 · 2 h ago
arXiv
2609.10939
T1 · 2 h ago
Category
cs.MA
T1 · 2 h ago

Abstract

Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.

Authors 7

Guanhua Chen, Haoxian Liu, Li Lu, Luming Yang, Rong Jia, Siqing Li, Yue Xiao

Specification

Official page

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Arxiv announce type
cross

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

arXiv id
2609.10939

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Categories
cs.AI, cs.HC, cs.MA

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

PDF

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Primary category
cs.MA

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Published
12 Sept 2026

Source:arXiv (Atom API + RSS)T1observed 2 h agohigh

Each value shows its source, tier and observation time. Conflicting claims are kept side by side and flagged — never averaged. How AI Atlas records facts →

Provenance

Attributed facts

9

Source tiers

T19

Freshest observation

2 h ago

Conflicts

None