- Artificial intelligence (AI) systems, including large language models (LLMs) such as OpenAI’s GPT series and Meta’s Llama series, have gained currency in many fields, including anesthesiology.
- LLMs are trained on large data sets and can generate contextually appropriate answers, but their clinical use requires careful scrutiny to ensure they provide sound guidance.
- In this study, the authors selected 7 clinical scenarios from the Stanford Emergency Manual, incorporated them into a clinical vignette, entered them into 22 distinct LLMs, and then evaluated the performance of each LLM in managing each scenario.
- Scenarios included asystole, pulseless ventricular tachycardia, unstable bradycardia, anaphylaxis, airway fire, malignant hyperthermia, and local anesthetic systemic toxicity (LAST).
- The LLMs were instructed to respond to the vignette from the perspective of an anesthesiologist by identifying the emergency, providing time-critical management steps, and supplying additional actions in anticipation of the natural clinical course of the emergency.
- The authors created rubrics for each scenario, which were normalized to ensure that each scenario was weighed equally. The rubrics included penalties if the LLMs provided information that was deemed harmful.
- There was high inter-rater agreement between the three authors who applied the rubrics to each clinical scenario evaluated by the LLMs.
- LLM performance varied depending on the scenario assessed. Overall, LLMs performed best on asystole and pulseless ventricular tachycardia, and lower on LAST and airway fire.
- Of the 22 LLMs included, OpenAI’s o1-preview scored the highest across scenarios and Meta’s Llama 3.2 scored the lowest.
- LLM performance on clinical vignettes varied widely:
- Smaller models were noted to perform poorly, but parameter count alone (settings which control a LLM’s output and behavior) did not always predict outcomes.
- Medically fine-tuned models underperformed the larger models they were based upon.
- Vignettes that involved the management of emergencies that were not anesthesia-specific (i.e., asystole) yielded higher scores than those that were anesthesia-specific (i.e., malignant hyperthermia).
- Future areas of research could include expanding the scenarios/vignettes tested, incorporating simulation in the assessment of LLMs, and potentially the semi-autonomous use of LLMs to apply the rubrics and decrease the human effort required in scenario grading.
Summary of "Benchmarking Large Language Model Performance in Perioperative Crisis Responses"
Summary published August 31, 2026