Assessment of decision-making with locally run and web-based large language models versus human board recommendations in otorhinolaryngology, head and neck surgery

dc.contributor.authorBuhr, Christoph Raphael
dc.contributor.authorErnst, Benjamin Philipp
dc.contributor.authorBlaikie, Andrew
dc.contributor.authorSmith, Harry
dc.contributor.authorKelsey, Tom
dc.contributor.authorMatthias, Christoph
dc.contributor.authorFleischmann, Maximilian
dc.contributor.authorJungmann, Florian
dc.contributor.authorAlt, Jürgen
dc.contributor.authorBrandts, Christian
dc.contributor.authorKämmerer, Peer W.
dc.contributor.authorFoersch, Sebastian
dc.contributor.authorKuhn, Sebastian
dc.contributor.authorEckrich, Jonas
dc.date.accessioned2026-07-23T07:32:16Z
dc.date.issued2025
dc.description.abstractIntroduction Tumor boards are a cornerstone of modern cancer treatment. Given their advanced capabilities, the role of Large Language Models (LLMs) in generating tumor board decisions for otorhinolaryngology (ORL) head and neck surgery is gaining increasing attention. However, concerns over data protection and the use of confidential patient information in web-based LLMs have restricted their widespread adoption and hindered the exploration of their full potential. In this first study of its kind we compared standard human multidisciplinary tumor board recommendations (MDT) against a web-based LLM (ChatGPT-4o) and a locally run LLM (Llama 3) addressing data protection concerns. Material and methods Twenty-five simulated tumor board cases were presented to an MDT composed of specialists from otorhinolaryngology, craniomaxillofacial surgery, medical oncology, radiology, radiation oncology, and pathology. This multidisciplinary team provided a comprehensive analysis of the cases. The same cases were input into ChatGPT-4o and Llama 3 using structured prompts, and the concordance between the LLMs' and MDT’s recommendations was assessed. Four MDT members evaluated the LLMs' recommendations in terms of medical adequacy (using a six-point Likert scale) and whether the information provided could have influenced the MDT's original recommendations. Results ChatGPT-4o showed 84% concordance (21 out of 25 cases) and Llama 3 demonstrated 92% concordance (23 out of 25 cases) with the MDT in distinguishing between curative and palliative treatment strategies. In 64% of cases (16/25) ChatGPT-4o and in 60% of cases (15/25) Llama, identified all first-line therapy options considered by the MDT, though with varying priority. ChatGPT-4o presented all the MDT’s first-line therapies in 52% of cases (13/25), while Llama 3 offered a homologous treatment strategy in 48% of cases (12/25). Additionally, both models proposed at least one of the MDT's first-line therapies as their top recommendation in 28% of cases (7/25). The ratings for medical adequacy yielded a mean score of 4.7 (IQR: 4–6) for ChatGPT-4o and 4.3 (IQR: 3–5) for Llama 3. In 17% of the assessments (33/200), MDT members indicated that the LLM recommendations could potentially enhance the MDT's decisions. Discussion This study demonstrates the capability of both LLMs to provide viable therapeutic recommendations in ORL head and neck surgery. Llama 3, operating locally, bypasses many data protection issues and shows promise as a clinical tool to support MDT decisions. However at present, LLMs should augment rather than replace human decision-making.en_GB
dc.identifier.doihttps://doi.org/10.25358/openscience-14998
dc.identifier.urihttps://openscience.ub.uni-mainz.de/handle/20.500.12030/15019
dc.language.isoeng
dc.rightsCC-BY-4.0
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/
dc.subject.ddc610 Medizinde_DE
dc.subject.ddc610 Medical sciencesen_GB
dc.titleAssessment of decision-making with locally run and web-based large language models versus human board recommendations in otorhinolaryngology, head and neck surgeryen_GB
dc.typeZeitschriftenaufsatzde_DE
jgu.apc.netprice2453,73
jgu.apc.price2625,49
jgu.apc.taxrate7
jgu.apc.transformationcontractSpringer (DEAL)
jgu.dfg.year2025
jgu.identifier.uuid3ed96ad0-8752-40cf-a15c-ab279c75bb6b
jgu.journal.titleEuropean archives of oto-rhino-laryngology and head & neck
jgu.journal.volume282
jgu.nationalcurrency.eur2453,73
jgu.organisation.departmentFB 04 Medizinde_DE
jgu.organisation.nameJohannes Gutenberg-Universität Mainzde_DE
jgu.organisation.number2700
jgu.organisation.placeMainz
jgu.organisation.rorhttps://ror.org/023b0x485
jgu.pages.end1607
jgu.pages.start1593
jgu.publisher.doi10.1007/s00405-024-09153-3
jgu.publisher.eissn1434-4726
jgu.publisher.nameSpringer
jgu.publisher.placeBerlin, Heidelberg
jgu.publisher.year2025
jgu.rights.accessrightsopenAccessen_GB
jgu.subject.ddccode610
jgu.subject.dfgLebenswissenschaftende_DE
jgu.type.dinitypeArticleen_GB
jgu.type.resourceTexten_GB
jgu.type.versionPublished versionen_GB

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
assessment_of_decisionmaking_-20260723093216384149.pdf
Size:
1.17 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
5.14 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections