The potential of large language models to address patients’ preoperative questions before anterior cervical discectomy and fusion surgery
Highlight box
Key findings
• GPT-5 and GROK 4 provided highly rated responses to common preoperative questions regarding anterior cervical discectomy and fusion (ACDF), with high scores for accuracy, completeness, and comprehensibility.
What is known and what is new?
• Large language models (LLMs) have demonstrated the ability to provide generally accurate medical information across multiple specialties, but their performance in ACDF-specific patient education has not been evaluated.
• This study is the first to assess GPT-5 and GROK 4 responses to common ACDF preoperative questions. Both models performed well overall; GROK 4 demonstrated greater completeness, whereas GPT-5 demonstrated greater comprehensibility.
What is the implication, and what should change now?
• LLMs may serve as useful adjuncts for preoperative patient education by providing rapid and understandable information regarding ACDF.
• LLM-generated information should supplement, not replace, physician-directed counseling because current models lack individualized risk assessment and personalized clinical recommendations.
Introduction
Large language models (LLMs), such as ChatGPT, are increasingly used to obtain medical information because they can provide rapid, human-like responses to health-related questions (1-4). As patients increasingly seek information through online and AI-powered platforms, there is growing interest in whether LLMs can reliably address common preoperative concerns (5-7).
Anterior cervical discectomy and fusion (ACDF) is a commonly performed procedure for cervical spine disorders causing neural compression (8). Patients considering ACDF frequently have questions regarding surgical indications, risks, recovery, and expected outcomes. Previous studies in hepatology, hand surgery, neurosurgery, orthopedics, and other specialties have shown that LLMs generally provide accurate and understandable information, although limitations in completeness, clinical nuance, personalization, and readability remain (2-4,6-12). Despite increasing interest in the use of LLMs for patient education, their performance in answering ACDF-specific preoperative questions has not been evaluated. Therefore, this pilot study assessed the ability of LLMs to answer common preoperative questions regarding ACDF. We hypothesized that LLMs would provide generally accurate and comprehensible responses but would demonstrate limitations in completeness, specificity, and individualized guidance.
Despite promising early data, no study to date has specifically evaluated the utility of LLMs in addressing the unique preoperative informational needs of patients scheduled for ACDF surgery. Understanding the capabilities and limitations of LLMs in this setting is critical, as effective preoperative education is known to improve surgical outcomes, patient satisfaction, and adherence to postoperative instructions.
In this pilot study, we assessed the potential of a LLM to answer common preoperative questions posed by patients undergoing ACDF. We hypothesize that while LLMs can provide generally accurate and accessible responses, limitations in depth, specificity, and personalization will necessitate complementary discussion with healthcare providers.
Methods
Study design
This was a cross-sectional descriptive study designed to assess the potential of a LLM in answering common preoperative questions related to ACDF surgery. A curated list of frequently asked patient questions was developed based on clinical experience, prior literature on patient informational needs, and typical preoperative consultations.
Question list
The following 18 questions were used to evaluate the model’s responses.
Understanding the procedure:
- What is ACDF, and why do I need it?
- What are the alternatives to ACDF surgery?
- What is the success rate of ACDF?
- How long do I stay in the hospital after ACDF surgery?
Symptoms and outcomes:
- What symptoms is ACDF expected to improve?
- What limitations should I expect after ACDF surgery?
- Will I be able to return to all my usual activities after ACDF surgery?
Risks and complications:
- What are the risks or complications associated with ACDF?
- What is the likelihood I will need another surgery in the future after ACDF surgery?
Surgical details:
- Will the hardware (plates or screws) remain in my neck permanently after ACDF surgery?
- Will I have a visible scar after ACDF surgery?
- How much pain should I expect after ACDF surgery?
- Will I need to wear a neck brace after ACDF surgery? If so, for how long?
- Will I lose any neck mobility after ACDF surgery?
Recovery and lifestyle:
- How long is the typical recovery period after ACDF surgery?
- When can I return to work after ACDF surgery?
- When will I be able to drive again after ACDF surgery?
Preoperative and imaging concerns:
- Do X-rays or CT scans before and after ACDF increase my risk of cancer?
Data collection
Each question was independently entered into GPT-5 (OpenAI, publicly available web version accessed September 7, 2025) and GROK 4 (xAI, publicly available web version accessed September 7, 2025). Exact internal model version identifiers were not publicly disclosed by the providers. Each question was submitted once in a new chat session to minimize contextual carryover effects. No additional prompts, clarifications, or follow-up questions were provided after the initial query. The output was generated only once. Responses were not regenerated; the first complete response generated by each model was recorded in full without modification.
The models were accessed through their standard online platforms with default settings. Internet connectivity and access to publicly available online information were available during response generation. To ensure blinded assessment, all responses were anonymized by removing model identifiers and were randomized before evaluation. Neurosurgeon reviewers were blinded to model identity throughout the rating process. The responses are provided in Appendices 1,2 (available online: https://cdn.amegroups.cn/static/public/10.21037atm-2026-0085-1.pdf, https://cdn.amegroups.cn/static/public/10.21037atm-2026-0085-2.pdf).
Response evaluation
Three attending spine neurosurgeons and one PGY-2 neurosurgery resident independently evaluated the LLM responses (Table 1, Tables S1-S4, Figure S1). Responses were assessed based on the following criteria:
- Accuracy (correctness of information compared to current clinical guidelines and expert opinion);
- Completeness (degree to which the response addressed all aspects of the question);
- Comprehensibility (clarity and readability for a general patient audience).
Table 1
| Question number | Question | Evaluation of GPT-5 answer by neurosurgeon | Evaluation of GROK 4 answer by neurosurgeon | |||||
|---|---|---|---|---|---|---|---|---|
| Accuracy (correctness of information compared to current clinical guidelines and expert opinion): 1 (completely incorrect) to 5 (completely correct) | Completeness (degree to which the response addressed all aspects of the question): 1 (incomplete) to 5 (complete) | Comprehensibility (clarity and readability for a general patient audience): 1 (difficult to understand) to 5 (easy to understand) | Accuracy (correctness of information compared to current clinical guidelines and expert opinion): 1 (completely incorrect) to 5 (completely correct) | Completeness (degree to which the response addressed all aspects of the question): 1 (incomplete) to 5 (complete) | Comprehensibility (clarity and readability for a general patient audience): 1 (difficult to understand) to 5 (easy to understand) | |||
| Understanding the procedure | ||||||||
| 1 | What is ACDF, and why do I need it? | 4.33 | 3.67 | 4.67 | 4.67 | 4.33 | 3.67 | |
| 2 | What are the alternatives to ACDF surgery? | 4.33 | 4.33 | 4.00 | 4.33 | 5.00 | 4.33 | |
| 3 | What is the success rate of ACDF? | 4.33 | 4.00 | 4.67 | 4.33 | 4.33 | 4.33 | |
| 4 | How long do I stay in the hospital after ACDF surgery? | 4.33 | 4.33 | 4.33 | 3.33 | 3.00 | 4.33 | |
| Symptoms and outcomes | ||||||||
| 5 | What symptoms is ACDF expected to improve? | 4.33 | 4.00 | 5.00 | 4.33 | 4.33 | 4.33 | |
| 6 | What limitations should I expect after ACDF surgery? | 5.00 | 4.33 | 5.00 | 4.67 | 4.67 | 4.33 | |
| 7 | Will I be able to return to all my usual activities after ACDF surgery? | 4.67 | 4.00 | 5.00 | 5.00 | 5.00 | 4.33 | |
| Risks and complications | ||||||||
| 8 | What are the risks or complications associated with ACDF? | 4.67 | 4.00 | 5.00 | 4.67 | 4.67 | 4.33 | |
| 9 | What is the likelihood I will need another surgery in the future after ACDF surgery? | 4.00 | 3.67 | 4.67 | 4.67 | 4.67 | 4.33 | |
| Surgical details | ||||||||
| 10 | Will the hardware (plates or screws) remain in my neck permanently after ACDF surgery? | 5.00 | 4.67 | 5.00 | 5.00 | 5.00 | 4.67 | |
| 11 | Will I have a visible scar after ACDF surgery? | 5.00 | 5.00 | 5.00 | 5.00 | 5.00 | 5.00 | |
| 12 | How much pain should I expect after ACDF surgery? | 4.00 | 4.00 | 4.33 | 4.00 | 4.67 | 4.33 | |
| 13 | Will I need to wear a neck brace after ACDF surgery? If so, for how long? | 4.67 | 4.33 | 5.00 | 4.33 | 4.67 | 4.67 | |
| 14 | Will I lose any neck mobility after ACDF surgery? | 4.67 | 4.33 | 4.67 | 4.00 | 4.67 | 4.33 | |
| Recovery and lifestyle | ||||||||
| 15 | How long is the typical recovery period after ACDF surgery? | 4.33 | 4.33 | 4.67 | 4.67 | 4.33 | 4.67 | |
| 16 | When can I return to work after ACDF surgery? | 4.33 | 4.33 | 4.33 | 4.33 | 4.67 | 4.33 | |
| 17 | When will I be able to drive again after ACDF surgery? | 4.67 | 4.67 | 5.00 | 5.00 | 5.00 | 5.00 | |
| Preoperative and imaging concerns | ||||||||
| 18 | Do X-rays or CT scans before and after ACDF increase my risk of cancer? | 4.67 | 4.67 | 5.00 | 5.00 | 5.00 | 4.00 | |
| Average ± SD | 4.52±0.31 | 4.26±0.35 | 4.74±0.31 | 4.52±0.45 | 4.61±0.47 | 4.41±0.31 | ||
ACDF, anterior cervical discectomy and fusion; CT, computed tomography; SD, standard deviation.
Each criterion was scored using a Likert scale:
- Accuracy: 1 (completely incorrect) to 5 (completely correct);
- Completeness: 1 (incomplete) to 5 (complete);
- Comprehensibility: 1 (difficult to understand) to 5 (easy to understand).
Reviewers scored all responses independently.
Resident responses were collected to assess whether there were differences between board-certified specialized neurosurgeons and neurosurgical trainees.
Statistical analysis
Inter-rater reliability was evaluated separately for LLM No. 1 and LLM No. 2. Ratings were provided on a 5-point Likert scale by three attending neurosurgeons (K.M., J.S., J.K.H.) and one neurosurgery resident (Z.R.) across three evaluation domains (accuracy, completeness, and comprehensibility). Agreement was assessed both among attendings and between the attending mean score and the resident rater.
Intraclass correlation coefficients (ICC) were calculated using a two-way mixed-effects model with absolute agreement and single measurements [ICC (3,1)], treating raters as fixed effects. Agreement among attendings was calculated using the three attending ratings, while agreement between attendings and the resident was assessed by comparing the mean attending score with the resident rating. Ninety-five percent confidence intervals for ICC estimates were obtained using bootstrap resampling (2,000 iterations).
Because Likert-scale data are ordinal and may demonstrate ceiling effects, ordinal Krippendorff’s alpha was additionally computed to quantify agreement in ordinal ranking. Krippendorff’s α was calculated from the raw rater-level scores without prior averaging or aggregation using an ordinal distance metric. Because the evaluation scores were concentrated near the upper end of the 5-point Likert scale, we anticipated potential ceiling effects and restricted score variance. Under such conditions, ICC may underestimate agreement because it depends on between-subject variance, whereas ordinal Krippendorff’s α is less sensitive to score compression and better reflects consistency in ordinal ranking. Therefore, ordinal Krippendorff’s α was considered the preferred measure for ordinal Likert-scale data under restricted score variance; however, both α and ICC are presented because they assess different aspects of agreement.
All analyses were performed in Python (version 3.11) using pandas (version 2.x) and NumPy (version 1.26) for data handling, a custom implementation of ICC based on Shrout and Fleiss methodology, and the krippendorff Python package (version 0.6.0) for ordinal Krippendorff’s alpha.
The readability of the answers was assessed with the Flesch Kincaid grade level (FKGL), calculated using an online tool (Readability analyzer: https://datayze.com/readability-analyzer).
Ethical statement
This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The requirement for institutional ethics approval was waived because this study did not involve human participants, patient interventions, or the use of patient medical records or identifiable personal health information. The study consisted solely of the evaluation of responses generated by publicly available artificial intelligence models using standardized, non-patient-specific questions. Accordingly, informed consent was not required.
Results
Overall, both LLMs received high ratings for accuracy, completeness, and comprehensibility across the evaluated questions. There were some questions that had more consensus than others between reviewing spine surgeons. Question 11 received a perfect score from attending surgeons (regarding postoperative scars). On the other hand, Question 1, which focused on what the procedure entails and operative indications had more variety of opinions regarding AI accuracy and completeness and comprehensibility.
With attendings’ average score, accuracy was equal between GTP-5 and GROK 4, 4.52. Completeness was higher in GROK 4, 4.61, than in GTP-5, 4.26. Comprehensibility was higher in GTP-5, 4.74, than in GROK 4, 4.41 (Table 2).
Table 2
| LLM | Comparison | ICC(3,1) (95% CI) | Krippendorff’s α | Sidenote 1 | Sidenote 2 |
|---|---|---|---|---|---|
| GPT-5 | Attendings only | 0.12 (0.02, 0.22) | 0.987 | Per-domain ordinal α (attendings): accuracy: 0.957; completeness: 0.956; comprehensibility: 0.967 | Calculated from raw scores from: Tables S1-S3 (attending) |
| Attending average vs. resident | 0.52 (0.28, 0.69) | 0.993 | Per-domain ordinal α (attending average vs. resident): accuracy: 0.957; completeness: 0.979; comprehensibility: 0.982 | Calculated from mean attending score and raw resident score | |
| GROK 4 | Attendings only | 0.21 (−0.01, 0.40) | 0.987 | Per-domain ordinal α (attendings): accuracy: 0.964; completeness: 0.970; comprehensibility: 0.952 | Calculated from raw scores from: Tables S1-S3 (attending) |
| Attending average vs. resident | 0.51 (0.25, 0.69) | 0.995 | Per-domain ordinal α (attending average vs. resident): accuracy: 0.979; completeness: 0.989; comprehensibility: 0.985 | Calculated from mean attending score and raw resident score |
CI, confidence interval; ICC, intraclass correlation coefficient; LLM, large language model.
The variety between attendings ranged between 0 and 1.33, except for two questions. Question 14 GROK 4 accuracy score showed a variation of 3.0, and question 9 GTP-5 completeness showed 2.33. The variation score between the average of attendings vs. a resident was smaller, and it was between 0 and 0.89. The highest score was shown in question1 GTP-5 completeness, and question 9, GROK 4 comprehensibility.
Intraclass correlation coefficients (ICC) and ordinal Krippendorff’s α were both reported because these metrics capture different aspects of agreement. The distribution of ratings demonstrated a pronounced ceiling effect, with most scores clustered between 4 and 5. Under conditions of restricted score variance, ICC values can be artificially reduced because ICC depends on between-subject variability. In contrast, ordinal Krippendorff’s α is less affected by score compression and more directly reflects agreement in ordinal ranking. For this reason, ordinal Krippendorff’s α was considered the preferred measure of inter-rater reliability, while ICC was interpreted as a secondary measure of absolute agreement.
Inter-attending agreement was poor by ICC(3,1) for both GTP-5 (0.12, 95% CI: 0.02–0.22) and GROK 4 (0.21, 95% CI: −0.01, 0.40), but strong by ordinal Krippendorff’s alpha (α≈0.99), indicating consistent ordinal ranking despite limited absolute agreement. Agreement between the attending mean score and the resident rater was moderate by ICC(3,1) for GTP-5 (0.52, 95% CI: 0.28–0.69) and GROK 4 (0.51, 95% CI: 0.25–0.69), with strong ordinal agreement across all domains (α range, 0.95–0.99).
The Flesch-Kincaid Grade Level (FKGL) scores were 10.15 for GPT-5 and 10.24 for GROK 4. The readability was better in GTP-5, though both were over 10th grade level. This finding was consistent with the higher comprehensibility ratings assigned to GPT-5.
Discussion
This study evaluated the potential of a LLM to address common preoperative questions from patients undergoing ACDF surgery. Our findings suggest that LLMs are capable of providing generally accurate, and comprehensible responses to patient inquiries. However, several important limitations were noted regarding the depth, specificity, and clinical personalization of the information provided.
Consistent with prior studies evaluating LLM in other medical domains such as hepatology (2), hand surgery (3), neurosurgery (8), and orthopedics (9), the model demonstrated a strong ability to deliver basic knowledge and general explanations. In particular, the LLMs performed well in addressing questions regarding the purpose of ACDF, expected symptom relief, postoperative recovery timelines, and common surgical risks. However, despite their clarity, the FKGL exceeded 10 for both models (6), indicating that the reading level may be higher than recommended for patient-facing educational materials intended for the general public.
Despite these strengths, notable limitations emerged. First, while LLM accurately described common complications of ACDF, such as range of motion loss and dysphagia, it did not always clearly differentiate between rare and common risks, potentially influencing patient perception. This finding aligns with prior research demonstrating that LLMs tend to provide generalized rather than nuanced clinical information (4,9). Second, the responses occasionally omitted important qualifiers, such as the need for surgeon consultation when making definitive surgical decisions. This is well demonstrated in the answer of the first question “why do I need ACDF”.
A notable finding was the divergence between ICC and ordinal Krippendorff’s α. The low ICC values likely reflect restricted score variance due to ceiling effects, whereas ordinal Krippendorff’s α remained high because it is less sensitive to score compression and better captures consistency in ordinal ranking. For attendings only, ICC is 0.12 (95% CI: 0.02–0.22) for GPT-5 and 0.21 (95% CI: −0.01, 0.40) for GROK 4. The low ICC values likely reflect restricted score variance caused by ceiling effects, whereas ordinal Krippendorff’s α remained high because it is less sensitive to score compression and better captures consistency in ordinal rankings. Therefore, the discrepancy between these metrics likely reflects methodological differences rather than true disagreement among raters. These two statistics are measuring different things and their divergence at this magnitude almost always signals a ceiling or floor effect: when raters cluster scores at the top of the scale (4–5 out of 5), between-subject variance collapses and ICC, which depends on that variance in its denominator, becomes artificially deflated, while ordinal α can remain high because raters are consistently agreeing on rank order even when absolute dispersion is minimal. Because of the observed ceiling effect, mean ± SD values and score distributions were provided to facilitate interpretation of agreement metrics. Under restricted score variance, ordinal Krippendorff’s α may better reflect consistency in ranking than ICC. Accordingly, individual attending ratings should not be considered fully interchangeable, and the attending mean score should be interpreted cautiously as a reference standard.
An additional concern relates to the completeness of the answers. While LLM often provided a reasonable overview, responses to complex questions, such as the prognosis of spinal fusion or the management of nonunion, were occasionally superficial compared to what a practicing surgeon would provide (3,8). Moreover, for questions requiring numerical data (e.g., fusion success rates or risk percentages), LLM tended to offer broad statements without citing specific statistics or evidence-based guidelines, a limitation previously reported in other clinical settings (2,6).
Nevertheless, the practical utility of LLMs in augmenting preoperative patient education should not be understated. LLM’s ability to generate quick, understandable responses could help patients feel more informed and empowered prior to consultations, ultimately improving shared decision-making processes. However, it is crucial that patients recognize the limitations of AI-generated information and understand that it should supplement, not replace, direct communication with their healthcare team (7).
Limitations
This study has several limitations. Only two LLMs (GPT5 and GROK 4) were evaluated, and results may not be generalizable to other AI platforms. Furthermore, the questions selected were standardized and may not capture the full breadth of patient-specific or context-dependent inquiries seen in clinical practice. Additionally, while responses were graded by experienced neurosurgeons, inherent subjectivity in scoring accuracy, completeness, and comprehensibility remains a consideration. Finally, ChatGPT’s knowledge cutoff at September 30, 2024, may have limited its ability to reflect the most up-to-date surgical guidelines or advancements. The model outputs are non-deterministic and may change over time.
Future directions
Future studies should evaluate real-world patient studies assessing satisfaction, comprehension, and decision-making outcomes after exposure to AI-generated preoperative information are warranted. The development of supervised, specialty-specific AI tools integrated with evidence-based guidelines could further enhance the reliability and personalization of patient education materials.
Conclusions
LLMs demonstrated generally high expert-rated accuracy, completeness, and comprehensibility when answering common preoperative questions regarding ACDF. These findings suggest that LLMs can provide informative responses to commonly asked patient questions; however, future studies should evaluate whether incorporating LLM-generated information into preoperative education influences patient comprehension, satisfaction, engagement, shared decision-making, adherence, and surgical outcomes.
Acknowledgments
None.
Footnote
Data Sharing Statement: Available at https://atm.amegroups.com/article/view/10.21037/atm-2026-0085/dss
Peer Review File: Available at https://atm.amegroups.com/article/view/10.21037/atm-2026-0085/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://atm.amegroups.com/article/view/10.21037/atm-2026-0085/coif). J.D.L. received consulting fees from Orthofix and Medtronic. J.S. received consulting fees and royalties from Degen Medical and Surgical Theater. The other authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Wang M, Ma H, Piao M. Effectiveness of large language models in preoperative and discharge education: a systematic review based on an evaluation framework. NPJ Digit Med 2026;9:122. [Crossref] [PubMed]
- Yeo YH, Samaan JS, Ng WH, et al. Assessing the performance of ChatGPT in answering questions regarding cirrhosis and hepatocellular carcinoma. Clin Mol Hepatol 2023;29:721-32. [Crossref] [PubMed]
- Jagiella-Lodise O, Suh N, Zelenski NA. Can Patients Rely on ChatGPT to Answer Hand Pathology-Related Medical Questions? Hand (N Y) 2025;20:801-9. [Crossref] [PubMed]
- Pugliese N, Wai-Sun Wong V, Schattenberg JM, et al. Accuracy, Reliability, and Comprehensibility of ChatGPT-Generated Medical Responses for Patients With Nonalcoholic Fatty Liver Disease. Clin Gastroenterol Hepatol 2024;22:886-889.e5. [Crossref] [PubMed]
- Ma J, Zhang Y, Tang H, et al. Evaluating the quality of large language model-generated preoperative patient education material: a comparative study across models and surgery types. Front Med (Lausanne) 2025;12:1701344. [Crossref] [PubMed]
- Ghanem YK, Rouhi AD, Al-Houssan A, et al. Dr. Google to Dr. ChatGPT: assessing the content and quality of artificial intelligence-generated medical information on appendicitis. Surg Endosc 2024;38:2887-93.
- Sharma SC, Ramchandani JP, Thakker A, et al. ChatGPT in Plastic and Reconstructive Surgery. Indian J Plast Surg 2023;56:320-5. [Crossref] [PubMed]
- Gajjar AA, Kumar RP, Paliwoda ED, et al. Usefulness and Accuracy of Artificial Intelligence Chatbot Responses to Patient Questions for Neurosurgical Procedures. Neurosurgery 2024;95:171-8. [Crossref] [PubMed]
- Sparks CA, Fasulo SM, Windsor JT, et al. ChatGPT Is Moderately Accurate in Providing a General Overview of Orthopaedic Conditions. JB JS Open Access 2024;9:e23.00129.
- Gordon EB, Towbin AJ, Wingrove P, et al. Enhancing Patient Communication With Chat-GPT in Radiology: Evaluating the Efficacy and Readability of Answers to Common Imaging-Related Questions. J Am Coll Radiol 2024;21:353-9. [Crossref] [PubMed]
- Mika AP, Martin JR, Engstrom SM, et al. Assessing ChatGPT Responses to Common Patient Questions Regarding Total Hip Arthroplasty. J Bone Joint Surg Am 2023;105:1519-26. [Crossref] [PubMed]
- Hu X, Niemann M, Kienzle A, et al. Evaluating ChatGPT responses to frequently asked patient questions regarding periprosthetic joint infection after total hip and knee arthroplasty. Digit Health 2024;10:20552076241272620. [Crossref] [PubMed]

