Artificial Intelligence and Multidisciplinary Oncology Decision-Making: A Systematic Review of LLM Concordance with MDT Recommendations and Implications for Clinical Reasoning Education
Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine
Introduction Multidisciplinary team (MDT) decision-making is a cornerstone of oncology practice and an important context for developing clinical reasoning skills. Artificial intelligence (AI), particularly large language models (LLMs), has recently been explored as a potential tool to support clinical decision-making in oncology settings. intro Methods This systematic review, conducted according to PRISMA 2020 guidelines, synthesizes evidence on the concordance between LLM-generated recommendations and MDT decisions in oncology. Twenty-two studies were included, primarily comprising retrospective analyses comparing AI-generated outputs with expert MDT consensus across diagnostic, therapeutic, and workflow-related tasks. Results LLMs demonstrated generally moderate to high concordance with MDT decisions in guideline-constrained clinical scenarios. However, performance decreased in complex or less-structured cases. Variability in outcomes was associated with model architecture, prompting strategies, and retrieval-augmented approaches. Across studies, AI performance was most consistent when aligned with established clinical guidelines. Substantial heterogeneity existed among models and study designs, limiting direct comparability. Conclusion From a cognitive perspective, LLMs may be conceptualized as external information-processing tools that partially mirror structured aspects of MDT reasoning in specific contexts. However, evidence remains limited to retrospective concordance
Abstract
Introduction Multidisciplinary team (MDT) decision-making is a cornerstone of oncology practice and an important context for developing clinical reasoning skills. Artificial intelligence (AI), particularly large language models (LLMs), has recently been explored as a potential tool to support clinical decision-making in oncology settings. intro Methods This systematic review, conducted according to PRISMA 2020 guidelines, synthesizes evidence on the concordance between LLM-generated recommendations and MDT decisions in oncology. Twenty-two studies were included, primarily comprising retrospective analyses comparing AI-generated outputs with expert MDT consensus across diagnostic, therapeutic, and workflow-related tasks. Results LLMs demonstrated generally moderate to high concordance with MDT decisions in guideline-constrained clinical scenarios. However, performance decreased in complex or less-structured cases. Variability in outcomes was associated with model architecture, prompting strategies, and retrieval-augmented approaches. Across studies, AI performance was most consistent when aligned with established clinical guidelines. Substantial heterogeneity existed among models and study designs, limiting direct comparability. Conclusion From a cognitive perspective, LLMs may be conceptualized as external information-processing tools that partially mirror structured aspects of MDT reasoning in specific contexts. However, evidence remains limited to retrospective concordance analyses and does not demonstrate educational impact. Any implications for clinical reasoning education should be interpreted as theoretical and require prospective validation using established educational frameworks such as cognitive apprenticeship or shared mental models.
