All projects

First-author research · 2025

Multi-agent medical AI research

What happens when an AI diagnostic team receives outside advice just as its working diagnosis is changing? I studied whether correct, incorrect, and bias-shaped suggestions help or mislead the system.

Read the paper on arXivView the full poster

My role and the simulation

As co-first and corresponding author, I directed the study, built the base GPT-4.1 multi-agent simulation, and directed programming for the ablations. The work was published in CCIS and presented at NeurIPS workshops.

The simulation included a Patient who answered symptom questions, a Doctor who formed working diagnoses, a Specialist who reviewed the case, a Measurement Agent that returned case-grounded laboratory or imaging data, and a Priming Agent that inserted an outside suggestion. Each Doctor turn logged a working diagnosis so we could see how it changed.

I first ran 214 MedQA cases without intervention. I embedded consecutive diagnoses and measured cosine similarity. A score below the tenth-percentile threshold (0.5462) marked a diagnostic fault point; the lowest qualifying score selected the intervention turn. We then reran qualifying cases with one priming message at that point.

What the interventions changed

Correct guidance with reasoning raised Top-1 accuracy from 50% to 60% in the main comparison. An incorrect suggestion lowered Top-1 accuracy to 48%, even when the correct answer remained among the top five possibilities.

Main intervention comparison
ConditionTop-1Top-3Top-5
No intervention50%70%74%
Correct subcategory56%72%80%
Incorrect subcategory48%72%78%
Correct + reasoning60%80%86%
Incorrect + reasoning48%76%80%

For the bias experiments, the Priming Agent paired an incorrect diagnostic suggestion with one of nine reasoning patterns. Overconfidence produced the largest Top-1 decline in its 50-scenario comparison, from 50% to 44%. Other cues behaved differently: representative heuristics reached 52%, while anchoring and omission matched baseline. The study did not show that every bias reduced accuracy.

What I learned from the ablations

Timing and repetition mattered. In a 25-case experiment with qualifying faults in both phases, correct-subcategory guidance reached 60% Top-1 when inserted in either phase alone and 76% when inserted in both. In another 25-case set, repeating an incorrect subcategory suggestion at three fault points reduced Top-1 accuracy from 44.4% with one intervention to 36%.

Incorrect cues increased requested tests and Doctor–Specialist disagreement. Correct cues could also shorten deliberation or bypass specialist input. Those effects are why I looked beyond final accuracy at how the diagnosis was reached.

Limits of the experiment

All agents shared one model, the cases were structured, and prompt-injected biases simplify real cognitive biases. These results describe the simulation, not clinical behavior with real patients.

NeurIPS workshop poster summarizing the diagnostic simulation, intervention results, bias conditions, and ablations
Workshop poster. Open the image to zoom in.

See the full technical portfolio (PDF)