Jahrestagung der Gesellschaft für Medizinische Ausbildung (GMA)
Jahrestagung der Gesellschaft für Medizinische Ausbildung (GMA)
Large language model-based automated scoring of clinical summary statements in medical training
Text
Background: Clinical summary statements convert unstructured patient information into a concise format that supports initial diagnostic reasoning. In medical education, they are used to teach students how to organize, reduce, interpret, and synthesize clinical data [1]. Assessing these summaries to provide effective feedback remains challenging. Previous attempts to automate this process relied on hard-coded techniques [2], but these approaches require substantial development effort. Large Language Models (LLMs) have shown strong performance in medical education and represent a promising alternative for automating the assessment of clinical summary statements.
Methods: To assess the applicability of LLMs in scoring clinical summary statements, we conducted an experimental comparative study between four available LLMs (three large commercial models and a smaller local model). 122 summary statements from a dataset provided by Hege and colleagues [2] were scored in five distinct runs. We assessed test-retest reliability using Fleiss’ κ and Intraclass Correlation Coefficient (ICC) and concordance with human experts reported as inter-rater agreement using Gwet’s AC1/AC2 and exact match accuracy. A rubric first proposed by Smith and colleagues [3] and extended by Hege and colleagues was used as the basis for evaluation after annotating the rubric components for LLM use.
Results: Four different prompting configurations were examined: zero-shot, zero-shot with chain-of-thought (CoT), few-shot, and few-shot CoT. All models demonstrated high test-retest reliability. GPT-4o and mistral large exhibited very high consistency while mistral 7B (the smaller local model) showed high ICC values but only fair to moderate agreement. Few-shot CoT prompting yielded the highest concordance with human expert raters, as reported by inter-rater agreement and exact match accuracy peaking for mistral large. Some rare exceptions and occasional extreme deviations were found in the small local mistral 7B model. Binary rubric components (person, factual accuracy) showed consistently higher agreement than ternary ones (transformation, global) (see figure 1 [Fig. 1]).
Figure 1: Agreement values according to Gwet’s AC1/AC2 by rubric component, prompting configuration, and model
Binary components (person, factual accuracy) generally reached higher reliability than abstract ternary components (semantic qualifiers, transformation, narrowing, global). Agreement values are interpreted using the Landis and Koch (1977) scale, which ranges from slight to almost perfect agreement.
Discussion: LLMs can reliably rate clinical summary statements under appropriate prompting and with verbose and clear context data, such as annotated rubrics, enabling scalable automated feedback. Binary criteria showed higher agreement than ternary ones, indicating that LLMs excel at surface-level correctness and clearly defined features, while nuanced aspects of clinical reasoning remain a future opportunity. Human involvement is still recommended for oversight.
Take-home message: LLMs can provide reliable, scalable feedback on clinical summaries, especially for clearly defined rubrics and criteria.
References
[1] Feblowitz JC, Wright A, Singh H, Samal L, Sittig DF. Summarization of clinical information: A conceptual model. J Biomed Inform. 2011;44(4):688-699. DOI: 10.1016/j.jbi.2011.03.008[2] Hege I, Kiesewetter I, Adler M. Automatic analysis of summary statements in virtual patients - a pilot study evaluating a machine learning approach. BMC Med Educ. 2020;20(1):366. DOI: 10.1186/s12909-020-02297-w
[3] Smith S, Kogan JR, Berman NB, Dell MS, Brock DM, Robins LS. The Development and Preliminary Validation of a Rubric to Assess Medical Students’ Written Summary Statements in Virtual Patient Cases. Acad Med. 2016;91(1):94-100. DOI: 10.1097/ACM.0000000000000800



