First-author research · 2025
Multi-agent medical AI research
What happens when an AI diagnostic team receives outside advice just as its working diagnosis is changing? I studied whether correct, incorrect, and bias-shaped suggestions help or mislead the system.
Read the paper on arXivView the full poster
My role and the simulation
As co-first and corresponding author, I directed the study, built the base GPT-4.1 multi-agent simulation, and directed programming for the ablations. The work was published in CCIS and presented at NeurIPS workshops.
The simulation included a Patient who answered symptom questions, a Doctor who formed working diagnoses, a Specialist who reviewed the case, a Measurement Agent that returned case-grounded laboratory or imaging data, and a Priming Agent that inserted an outside suggestion. Each Doctor turn logged a working diagnosis so we could see how it changed.
I first ran 214 MedQA cases without intervention. I embedded consecutive diagnoses and measured cosine similarity. A score below the tenth-percentile threshold (0.5462) marked a diagnostic fault point; the lowest qualifying score selected the intervention turn. We then reran qualifying cases with one priming message at that point.
What the interventions changed
Correct guidance with reasoning raised Top-1 accuracy from 50% to 60% in the main comparison. An incorrect suggestion lowered Top-1 accuracy to 48%, even when the correct answer remained among the top five possibilities.
| Condition | Top-1 | Top-3 | Top-5 |
|---|---|---|---|
| No intervention | 50% | 70% | 74% |
| Correct subcategory | 56% | 72% | 80% |
| Incorrect subcategory | 48% | 72% | 78% |
| Correct + reasoning | 60% | 80% | 86% |
| Incorrect + reasoning | 48% | 76% | 80% |
For the bias experiments, the Priming Agent paired an incorrect diagnostic suggestion with one of nine reasoning patterns. Overconfidence produced the largest Top-1 decline in its 50-scenario comparison, from 50% to 44%. Other cues behaved differently: representative heuristics reached 52%, while anchoring and omission matched baseline. The study did not show that every bias reduced accuracy.
What I learned from the ablations
Timing and repetition mattered. In a 25-case experiment with qualifying faults in both phases, correct-subcategory guidance reached 60% Top-1 when inserted in either phase alone and 76% when inserted in both. In another 25-case set, repeating an incorrect subcategory suggestion at three fault points reduced Top-1 accuracy from 44.4% with one intervention to 36%.
Incorrect cues increased requested tests and Doctor–Specialist disagreement. Correct cues could also shorten deliberation or bypass specialist input. Those effects are why I looked beyond final accuracy at how the diagnosis was reached.
Limits of the experiment
All agents shared one model, the cases were structured, and prompt-injected biases simplify real cognitive biases. These results describe the simulation, not clinical behavior with real patients.
