TC3-VLM: A Vision–Language Model for Tactical Combat Casualty Care
Source: PubMed Central Open Access, NCBI / U.S. National Library of Medicine
Tactical Combat Casualty Care (TC3) defines the current standard of military trauma care in the pre-hospital and battlefield settings, where life-saving decisions are made while operating at the extremes of human performance, under high stress, and with limited information. Despite growing interest in AI-assisted battlefield medicine, visual reasoning remains underexplored in existing systems. Recent advances in vision–language models (VLMs) enable the integration of visual information with medical reasoning, but their application to combat casualty care remains limited. To address this gap, we present TC3-VLM, the first VLM explicitly tailored for combat casualty care. Using a dual question–answer (Q&A) generation framework, we create visual-grounded Q&A to capture spatial and situational reasoning alongside caption-based Q&A to encode protocol knowledge, resulting in a domain-specific dataset of 7,585 Q&A pairs derived from over 52 hours of TC3 videos. We then fine-tune multiple VLM backbones via Low-Rank Adaptation across varying model scales. Experiments demonstrate that fine-tuned models consistently outperform their baseline counterparts, with TC3-VLM surpassing a 72B-parameter general-purpose VLM despite being significantly smaller in scale. Ablation and a TC3-grounded error analysis indicate that visual grounding is the primary driver of these gains and that fine-tuning reduces safety-relevant errors, positioning TC3-VLM as a training and reference aid for combat casu
Abstract
Tactical Combat Casualty Care (TC3) defines the current standard of military trauma care in the pre-hospital and battlefield settings, where life-saving decisions are made while operating at the extremes of human performance, under high stress, and with limited information. Despite growing interest in AI-assisted battlefield medicine, visual reasoning remains underexplored in existing systems. Recent advances in vision–language models (VLMs) enable the integration of visual information with medical reasoning, but their application to combat casualty care remains limited. To address this gap, we present TC3-VLM, the first VLM explicitly tailored for combat casualty care. Using a dual question–answer (Q&A) generation framework, we create visual-grounded Q&A to capture spatial and situational reasoning alongside caption-based Q&A to encode protocol knowledge, resulting in a domain-specific dataset of 7,585 Q&A pairs derived from over 52 hours of TC3 videos. We then fine-tune multiple VLM backbones via Low-Rank Adaptation across varying model scales. Experiments demonstrate that fine-tuned models consistently outperform their baseline counterparts, with TC3-VLM surpassing a 72B-parameter general-purpose VLM despite being significantly smaller in scale. Ablation and a TC3-grounded error analysis indicate that visual grounding is the primary driver of these gains and that fine-tuning reduces safety-relevant errors, positioning TC3-VLM as a training and reference aid for combat casualty care. ABS1
