Close Menu
AIToday7

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    How Much Does a Local LLM Actually Cost to Run? I Measured Every Watt on Apple Silicon

    July 28, 2026

    Discovering cryptographic weaknesses with Claude

    July 28, 2026

    Are you struggling to find a tech job on the West Coast?

    July 28, 2026
    Facebook X (Twitter) Instagram
    Trending
    • How Much Does a Local LLM Actually Cost to Run? I Measured Every Watt on Apple Silicon
    • Discovering cryptographic weaknesses with Claude
    • Are you struggling to find a tech job on the West Coast?
    • How AI Is Helping Teen Entrepreneurs Launch Startups
    • 12 keychain gadgets worth carrying every day (and why they’re worth it)
    • More than 30 Minnesota water systems targeted in cyberattack
    • Alaina Lamberson, Recognized by Influential Women, Serves as API Integration Specialist and Prompt Engineer at Portable
    • Elon Musk’s xAI sues to stop Minnesota law banning nudification technology
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AIToday7
    • Home
    • AI News
    • Tech News
    • AI Guides
    • Chatbots
    • Cybersecurity
    • Gadgets
    • More
      • Generative AI
      • Startups
    AIToday7
    Home»Generative AI»Tutorial: guidance on the use of large language models for medical research
    Generative AI

    Tutorial: guidance on the use of large language models for medical research

    aitoday7By aitoday7July 26, 2026No Comments17 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Tutorial: guidance on the use of large language models for medical research
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Abstract

    Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4 and DeepSeek-R1, represent a transformative class of artificial intelligence tools capable of revolutionizing various aspects of healthcare by generating human-like responses across diverse contexts and adapting to novel tasks following human instructions. Their potential application spans a broad range of medical tasks, such as clinical documentation, matching patients to clinical trials and answering medical questions. Here in this Tutorial, we discuss an actionable set of best practices to help healthcare professionals utilize LLMs more effectively and efficiently. The overall workflow follows sequential phases from formulating the task, choosing the most appropriate LLMs, engineering the prompts, fine-tuning the requests and through to model deployment. We discuss a set of critical considerations in identifying medical tasks that align with the core capabilities of LLMs and selecting models based on the required task, data, performance and model interface. We then review the strategies, such as prompt engineering and fine-tuning, to adapt standard LLMs to specialized medical tasks. We then cover deployment considerations, including regulatory compliance, ethical guidelines and continuous monitoring for fairness and bias. By providing a structured step-by-step methodology, this entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.

    Access through your institution
    Buy or subscribe

    This is a preview of subscription content, access

    Access options

    • Purchase on SpringerLink
    • Instant access to the full article PDF.

    Prices may be subject to local taxes which are calculated during checkout

    Fig. 1: Overview of the proposed systematic approach to utilizing LLMs in medicine.
    Fig. 2: An overview of five common task formulations enabled by LLMs in medical research, with a set of examples.
    Fig. 3: Considerations for choosing the LLMs.
    Fig. 4: An overview of prompt engineering and fine-tuning techniques.

    Code availability

    Tutorial scripts are provided at https://github.com/ncbi-nlp/LLM-Medicine-Primer.

    References

    1. GPT-5 system card. OpenAIhttps://cdn.openai.com/gpt-5-system-card.pdf (2025).

    2. Introducing Claude Opus 4.5. Anthropichttps://www.anthropic.com/news/claude-opus-4-5 (2025).

    3. A new era of intelligence with Gemini 3. Googlehttps://blog.google/products/gemini/gemini-3/ (2025).

    4. The Llama 4 herd: the beginning of a new era of natively multimodal AI innovation. Metahttps://ai.meta.com/blog/llama-4-multimodal-intelligence/ (2025).

    5. Guo, D. et al. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature645, 633–638 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    6. Brown, T. et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst.33, 1877–1901 (2020).

    7. Ouyang, L. et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst.35, 27730–27744 (2022).

    8. Tian, S. et al. Opportunities and challenges for ChatGPT and large language models in biomedicine and health. Brief. Bioinform.25, bbad493 (2024).

    9. Saab, K. et al. Capabilities of gemini models in medicine. Preprint at https://arxiv.org/abs/2404.18416 (2024).

    10. Liu, F. et al. Application of large language models in medicine. Nat. Rev. Bioeng.3, 445–464 (2025).

    11. Zhou, J. et al. Large language models in biomedicine and healthcare. npj Artif. Intell.1, 44 (2025).

    12. Lu, Z. et al. Large language models in biomedicine and health: current research landscape and future directions. J. Am. Med. Inform. Assoc.31, 1801–1811 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    13. Liévin, V., Hother, C. E., Motzfeldt, A. G. & Winther, O. Can large language models reason about medical questions? Patterns (2023).

    14. Nori, H. et al. Can generalist foundation models outcompete special-purpose tuning? Case study in medicine. Preprint at https://arxiv.org/abs/2311.16452 (2023).

    15. Singhal, K. et al. Large language models encode clinical knowledge. Nature620, 172–180 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    16. Hu, X. et al. Interpretable medical image visual question answering 103279 (2024)

      Article 
      PubMed 
      Google Scholar 

    17. Jin, Q. et al. Matching patients to clinical trials with large language models. Nat. Commun.15, 9074 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    18. Wong, C. et al. Scaling clinical trial matching using large language models: a case study in oncology. In Proc. 8th Machine Learning for Healthcare Conference Vol. 219 (eds Deshpande, K. et al.) 846–862 (PMLR, 2023).

    19. Wornow, M. et al. Zero-shot clinical trial patient matching with LLMs. NEJM AI2, AIcs2400360 (2025).

    20. Roberts, K. Large language models for reducing clinicians’ documentation burden. Nat. Med.30, 942–943 (2024).

      Article 
      PubMed 
      Google Scholar 

    21. Van Veen, D. et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med.30, 1134–1142 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    22. Patel, S. B. & Lam, K. ChatGPT: the future of discharge summaries? Lancet Digit. Health5, e107–e108 (2023).

      Article 
      PubMed 
      Google Scholar 

    23. Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature616, 259–265 (2023).

      Article 
      PubMed 
      Google Scholar 

    24. Acosta, J. N., Falcone, G. J., Rajpurkar, P. & Topol, E. J. Multimodal biomedical AI. Nat. Med.28, 1773–1784 (2022).

      Article 
      PubMed 
      Google Scholar 

    25. Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med.29, 1930–1940 (2023).

      Article 
      PubMed 
      Google Scholar 

    26. Omiye, J. A., Gui, H., Rezaei, S. J., Zou, J. & Daneshjou, R. Large language models in medicine: the potentials and pitfalls: a narrative review. Ann. Intern. Med.177, 210–220 (2024).

      Article 
      PubMed 
      Google Scholar 

    27. Zhou, H. et al. A survey of large language models in medicine: progress, application, and challenge. Preprint at https://arxiv.org/abs/2311.05112 (2023).

    28. Hu, Y. et al. Improving large language models for clinical named entity recognition(2024)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    29. Wang, L. et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. npj Digit. Med.7, 41 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    30. Jin, Q. et al. Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine. npj Digit. Med.7, 190 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    31. Nori, H., King, N., McKinney, S. M., Carignan, D. & Horvitz, E. Capabilities of GPT-4 on medical challenge problems. Preprint at https://arxiv.org/abs/2303.13375 (2023).

    32. Liévin, V., Hother, C. E. & Winther, O. Can large language models reason about medical questions? Preprint at https://arxiv.org/abs/2207.08143 (2022).

    33. Gilson, A. et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med. Educ.9, e45312 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    34. Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med.31, 943–950 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    35. Jin, Q., Leaman, R. & Lu, Z. Retrieve, summarize, and verify: how will ChatGPT affect information seeking from the medical literature? J. Am. Soc. Nephrol.34, 1302–1304 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    36. Liu, S. et al. Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J. Am. Med. Inform. Assoc.30, 1237–1245 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    37. Shaib, C. et al. Summarizing, simplifying, and synthesizing medical evidence using GPT-3 (with varying success). In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 2 (eds Rogers, A., Boyd-Graber, J. & Okazaki, N.) 1387–1407 (Association for Computational Linguistics, 2023).

    38. Tang, L. et al. Evaluating large language models on medical evidence summarization. npj Digit. Med.6, 158 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    39. Zhang, G. et al. Closing the gap between openmmarization. npj Digit. Med.7, 239 (2024)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    40. Papineni, K., Roukos, S., Ward, T. & Zhu, W.-J. BLEU: a method for automatic evaluation of machine translation. In Proc. 40th Annual Meeting of the Association for Computational Linguistics (eds Isabelle, P., Charniak, E. & Lin, D.) 311–318 (Association for Computational Linguistics, 2002).

    41. Lin, C.-Y. Rouge: a package for automatic evaluation of summaries. In Proc.Text Summarization Branches Out 74–81 (Association for Computational Linguistics, 2004).

    42. Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q. & Artzi, Y. BERTScore: evaluating text generation with BERT. In Proc. 8th International Conference on Learning Representations (ICLR, 2020).

    43. Wang, L. L. et al. Automated metrics for medical multi-document summarization disagree with human evaluations. In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds Rogers, A., Boyd-Graber, J. & Okazaki, N.) 9871–9889 (Association for Computational Linguistics, 2023).

    44. NLLB Team. Scaling neural machine translation to 200 languages. Nature630, 841–846 (2024).

    45. Preiksaitis, C. & Rose, C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: scoping review. JMIR Med. Educ.9, e48785 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    46. Abd-Alrazaq, A. et al. Large language models in medical education: opportunities, challenges, and future directions. JMIR Med. Educ.9, e48291 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    47. Mirza, F. N. et al. Using ChatGPT to facilitate truly informed medical consent. NEJM AI1, AIcs2300145 (2024).

    48. Wang, H., Gao, C., Dantona, C., Hull, B. & Sun, J. DRG-LLaMA: tuning LLaMA model to predict diagnosis-related group for hospitalized patients. npj Digit. Med.7, 16 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    49. Dagdelen, J. et al. Structured information extraction from scientific text with large language models. Nat. Commun.15, 1418 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    50. Topol, E. J. As artificial intelligence goes multimodal, medical applications multiply. Science381, adk6139 (2023).

      Article 
      PubMed 
      Google Scholar 

    51. Chen, P.-H. C., Liu, Y. & Peng, L. How to develop machine learning models for healthcare. Nat. Mater.18, 410–414 (2019).

      Article 
      PubMed 
      Google Scholar 

    52. Doshi, R. et al. Quantitative evaluation of large language models to streamline radiology report impressions: a multimodal retrospective analysis. Radiology310, e231593 (2024).

      Article 
      PubMed 
      Google Scholar 

    53. Jiang, A. Q. et al. Mixtral of experts. Preprint at https://arxiv.org/abs/2401.04088 (2024).

    54. Yang, A. et al. Qwen3 technical report. Preprint at https://arxiv.org/abs/2505.09388 (2025).

    55. Minaee, S. et al. Large language models: a survey. Preprint at https://arxiv.org/abs/2402.06196 (2024).

    56. Wu, C. et al. PMC-LLaMA: toward building open-1833–1843 (2024)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    57. Chen, Z. et al. Meditron-70b: scaling medical pretraining for large language models. Preprint at https://arxiv.org/abs/2311.16079 (2023).

    58. Nori, H. et al. From medprompt to o1: exploration of run-time strategies for medical challenge problems and beyond. Preprint at https://arxiv.org/abs/2411.03590 (2024).

    59. Tang, X. et al. Medagentsbench: benchmarking thinking models and agent frameworks for complex medical reasoning. Preprint at https://arxiv.org/abs/2503.07459 (2025).

    60. Health Insurance Portability and Accountability Act of 1996 (HIPAA). CDChttps://www.cdc.gov/phlp/php/re6-hipaa.html (2024)

    61. Cui, H. et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat. Methods21, 1470–1480 (2024).

      Article 
      PubMed 
      Google Scholar 

    62. Liu, N. F. et al. Lost in the middle: how language models use long contexts. Trans. Assoc. Comput. Linguist.11, 157–173 (2024).

    63. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Vol. 1 (eds Burstein, J., Doran, C. & Solorio, T.) 4171–4186 (Association for Computational Linguistics, 2019).

    64. Chen, Q. et al. Benchmarking large language models for biomedical natural language processing applications and recommendations. Nat. Commun.16, 3280 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    65. Jin, D. et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Appl. Sci.11, 6421 (2021).

    66. Jin, Q., Dhingra, B., Liu, Z., Cohen, W. & Lu, X. PubMedQA: a dataset for biomedical research question answering. In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (eds Inui, K., Jiang, J., Ng, V. & Wan, X.) 2567–2577 (Association for Computational Linguistics, 2019).

    67. Pal, A., Umapathi, L. K. & Sankarasubbu, M. MedMCQA: a large-scale multi-subject multi-choice dataset for medical domain question answering. In Proc. Conference on Health, Inference, and Learning Vol. 174 (eds Flores, G., Chen, G. H., Pollard, T., Ho, J. C. & Naumann, T.) 248–260 (PMLR, 2022).

    68. Griot, M., Hemptinne, C., Vanderdonckt, J. & Yuksel, D. Large language models lack essential metacognition for reliable medical reasoning. Nat. Commun.16, 642 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    69. Yang, Y. et al. Beyond multiple-choice accuracy: real-world challenges of implementing large language models in healthcare. Annu. Rev. Biomed. Data Sci.8, 305–316 (2025).

      Article 
      PubMed 
      Google Scholar 

    70. Arora, R. K. et al. Healthbench: evaluating large language models towards improved human health. Preprint at https://arxiv.org/abs/2505.08775 (2025).

    71. Chiang, W.-L. et al. Chatbot Arena: an open platform for evaluating LLMs by human preference. In Proc. 41st International Conference on Machine Learning Vol. 235 (eds Salakhutdinov, R. et al.) 8359–8388 (PMLR, 2024).

    72. Mukherjee, P., Hou, B., Lanfredi, R. B. & Summers, R. M. Feasibility of using the privacy-preserving large language model Vicuna for labeling radiology reports. Radiology309, e231147 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    73. Liu, P. et al. Pre-train, prompt, and predict: a systematic survey of prompting methods in natural language processing. ACM Comput. Surv.55, 1–35 (2023).

    74. Khattab, O. et al. DSPy: compiling declarative language model calls into state-of-the-art pipelines. In Proc. Twelfth International Conference on Learning Representations (ICLR, 2024).

    75. Yuksekgonul, M. et al. Optimizing generative AI by backpropagating language model feedback. Nature639, 609–616 (2025).

      Article 
      PubMed 
      Google Scholar 

    76. Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst.35, 24824–24837 (2022).

    77. Jin, Q. et al. Biomedical question answering: a survey of approaches and challenges. ACM Comput. Surv.55, 1–36 (2022).

    78. Ji, Z. et al. Survey of hallucination in natural language generation. ACM Comput. Surv.55, 1–38 (2023).

    79. Li, M. et al. Benchmarking retrieval-augmented large language models in biomedical NLP: aplication, robustness, and self-awareness. Sci. Adv.11, eadr1443 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    80. Wu, K. et al. An automated framework for assessing how well LLMs cite relevant medical references. Nat. Commun.16, 3615 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    81. Xiong, G., Jin, Q., Lu, Z. & Zhang, A. Benchmarking retrieval-augmented generation for medicine. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (eds Ku, L.-W., Martins, A. & Srikumar, V.) 6233–6251 (Association for Computational Linguistics, 2024).

    82. Xiong, G. et al. Improving retrieval-augmented generation in medicine with iterative follow-up questions. In Proc. Pacific Symposium on Biocomputing (eds Altman, R. B., Hunter, L., Ritchie, M. D. & Klein, T. E.) 199–214 (World Scientific, 2025).

    83. Fiorini, N., Leaman, R., Lipman, D. J. & Lu, Z. How user intelligence is improving PubMed. Nat. Biotechnol.36, 937–945 (2018).

    84. Jin, Q., Leaman, R. & Lu, Z. PubMed and beyond: biomedical literature search in the age of artificial intelligence. EBioMedicine100, 104988 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    85. Comeau, D. C., Wei, C.-H., Islamaj Doğan, R. & Lu, Z. PMC text mining subset in BioC: about three million full-text articles and growing. Bioinformatics35, 3533–3535 (2019).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    86. Jin, Q., Yang, Y., Chen, Q. & Lu, Z. GeneGPT: augmenting large language models with domain tools for improved access to biomedical information. Bioinformatics40, btae075 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    87. Wang, Z. et al. GeneAgent: self-verification language agent for gene-set analysis using domain databases. Nat. Methods22, 1677–1685 (2025).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    88. Khandekar, N. et al. Medcalc-bench: evaluating large language models for medical calculations. Adv. Neural Inf. Process. Syst.37, 84730–84745 (2024).

    89. Shi, W. et al. EHRAgent: code empowers large language models for few-shot complex tabular reasoning on electronic health records. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 22315–22339 (Association for Computational Linguistics, 2024).

    90. Wang, X. et al. Self-consistency improves chain of thought reasoning in language models. In Proc. Eleventh International Conference on Learning Representations (ICLR, 2023).

    91. Hou, W. & Ji, Z. Assessing GPT-4 for cell type annotation in single-cell RNA-seq analysis. Nat. Methods21, 1462–1465 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    92. Cosentino, J. et al. Towards a personal health large language model. Preprint at https://arxiv.org/abs/2406.06474 (2024).

    93. Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In Proc. Tenth International Conference on Learning Representations (ICLR, 2022).

    94. Dettmers, T., Pagnoni, A., Holtzman, A. & Zettlemoyer, L. Qlora: efficient finetuning of quantized llms. Adv. Neural Inf. Process. Syst.36, 10088–10115 (2023).

    95. Lialin, V., Deshpande, V. & Rumshisky, A. Scaling down to scale up: a guide to parameter-efficient fine-tuning. Preprint at https://arxiv.org/abs/2303.15647 (2023).

    96. Biderman, D. et al. Lora learns less and forgets less. Preprint at https://arxiv.org/abs/2405.09673 (2024).

    97. Chung, H. W. et al. Scaling instruction-finetuned language models. J. Mach. Learn. Res.25, 1–53 (2024).

    98. Anisuzzaman, D., Malins, J. G., Friedman, P. A. & Attia, Z. I. Fine-tuning large language models for specialized use cases. Mayo Clin. Proc. Digit. Health3, 100184 (2025).

      Article 
      PubMed 
      Google Scholar 

    99. Wu, E., Wu, K. & Zou, J. Limitations of learning new and updated medical knowledge with commercial fine-tuning large language models. NEJM AI2, AIcs2401155 (2025).

    100. Voigt, P. & Von dem Bussche, A. The EU General Data Protection Regulation (GDPR). A Practical Guide 1st edn (Springer, 2017).

    101. Yang, Y. et al. A survey of recent methods for addressing AI fairness and bias in biomedicine. J. Biomed. Inform.154, 104646 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    102. Omiye, J. A., Lester, J. C., Spichak, S., Rotemberg, V. & Daneshjou, R. Large language models propagate race-based medicine. npj Digit. Med.6, 195 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    103. Yang, Y., Liu, X., Jin, Q., Huang, F. & Lu, Z. Unmasking and quantifying racial bias of large language models in medical report generation. Commun. Med.4, 176 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    104. Huang, Y. et al. Trustllm: trustworthiness in large language models. Preprint at https://arxiv.org/abs/2401.05561 (2024).

    105. Wang, B. et al. DecodingTrust: a comprehensive assessment of trustworthiness in GPT models. In Proc. 37th Conference on Neural Information Processing Systems 31232–31339 (Association for Computing Machinery, 2023).

    106. Johnson, A. E. et al. MIMIC-III, a freely accessible critical care database. Sci. Data3, 160035 (2016).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    107. Meskó, B. & Topol, E. J. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digit. Med.6, 120 (2023).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    108. Clark, C. R. et al. Health care equity in the use of advanced analytics and artificial intelligence technologies in primary care. J. Gen. Intern. Med.36, 3188–3193 (2021).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    109. Wang, S., Hu, M., Li, Q., Safari, M. & Yang, X. Capabilities of GPT-5 on multimodal medical reasoning. Preprint at https://arxiv.org/abs/2508.08224 (2025).

    110. Vishwanath, K. et al. Generalist large language models outperform clinical tools on medical benchmarks. Preprint at https://arxiv.org/abs/2512.01191 (2025).

    111. Huang, Y. et al. MedReflect: teaching medical LLMs to self-improve0.03687 (2025)

    112. MedQA leaderboard. ValsAIhttps://www.vals.ai/benchmarks/medqa (2025).

    113. Li, A. et al. QuarkMed Medical Foundation Model Technical Report. Preprint at https://arxiv.org/abs/2508.11894 (2025).

    114. Shi, W. et al. MedAdapter: efficient test-time adaptation of large language models towards medical reasoning. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 22294–22314 (Association for Computational Linguistics, 2024).

    115. Pal, A. & Sankarasubbu, M. Gemini goes to med school: exploring the capabilities of multimodal large language models on medical challenge problems & hallucinations. In Proc. 6th Clinical Natural Language Processing Workshop (eds Naumann, T., Ben Abacha, A., Bethard, S., Roberts, K. & Bitterman, D.) 21–46 (Association for Computational Linguistics, 2024).

    116. Bran, A. M. et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell.6, 525–535 (2024).

    117. Zakka, C. et al. Almanac—retrieval-augmented language models for clinical medicine. NEJM AI1, AIoa2300068 (2024).

    118. Zhang, K. et al. A generalist vision–language foundation model for diverse biomedical tasks. Nat. Med.30, 3129–3141 (2024).

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    Acknowledgements

    This research was supported by the Intramural Research Program of the National Institutes of Health (NIH). The contributions of the NIH author(s) are considered works of the US Government. Q.J. was also supported by the NIH Pathway to Independence Award K99LM014903. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the US Department of Health and Human Services.

    Authors and Affiliations

    Ethics declarations

    Competing interests

    The authors declare no competing interests.

    Peer review

    Peer review information

    Nature Protocols thanks Daniel Alber and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.

    Additional information

    Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

    Supplementary information

    Rights and permissions

    About this article

    Cite this article

    Jin, Q., Wan, N., Leaman, R. et al. Tutorial: guidance on the use of large language models for medical research.
    Nat Protoc (2026). https://doi.org/10.1038/s41596-026-01408-z

    • Version of record:24 July 2026

    • DOI
      :https://doi.org/10.1038/s41596-026-01408-z

    guidance language Large models Tutorial
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleChatGPT now has a space for sharing medical records. Should you?
    Next Article Misleading AI-generated doctors pose ‘huge danger to public safety’
    aitoday7
    • Website

    Related Posts

    Generative AI

    How Much Does a Local LLM Actually Cost to Run? I Measured Every Watt on Apple Silicon

    July 28, 2026
    Generative AI

    Some people’s chats with Claude AI found publicly available online

    July 28, 2026
    Generative AI

    What Google has teased about Gemini 4

    July 27, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    How Much Does a Local LLM Actually Cost to Run? I Measured Every Watt on Apple Silicon

    July 28, 20260 Views

    Discovering cryptographic weaknesses with Claude

    July 28, 20260 Views

    Are you struggling to find a tech job on the West Coast?

    July 28, 20260 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    Chatbots

    OpenAI bets on families as ChatGPT goes deeper into households

    aitoday7July 11, 2026
    Generative AI

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    aitoday7July 11, 2026
    AI News

    Safe from AI: which jobs will help you thrive in the future?

    aitoday7July 11, 2026

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    How Much Does a Local LLM Actually Cost to Run? I Measured Every Watt on Apple Silicon

    July 28, 20260 Views

    Discovering cryptographic weaknesses with Claude

    July 28, 20260 Views

    Are you struggling to find a tech job on the West Coast?

    July 28, 20260 Views
    Our Picks

    OpenAI bets on families as ChatGPT goes deeper into households

    July 11, 2026

    MUSIC COMMUNITY INTRODUCES NEW LABELING PROGRAM TO DISTINGUISH GENERATIVE AI IN SOUND RECORDINGS

    July 11, 2026

    Safe from AI: which jobs will help you thrive in the future?

    July 11, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 AIToday7. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.