Abstract
Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4 and DeepSeek-R1, represent a transformative class of artificial intelligence tools capable of revolutionizing various aspects of healthcare by generating human-like responses across diverse contexts and adapting to novel tasks following human instructions. Their potential application spans a broad range of medical tasks, such as clinical documentation, matching patients to clinical trials and answering medical questions. Here in this Tutorial, we discuss an actionable set of best practices to help healthcare professionals utilize LLMs more effectively and efficiently. The overall workflow follows sequential phases from formulating the task, choosing the most appropriate LLMs, engineering the prompts, fine-tuning the requests and through to model deployment. We discuss a set of critical considerations in identifying medical tasks that align with the core capabilities of LLMs and selecting models based on the required task, data, performance and model interface. We then review the strategies, such as prompt engineering and fine-tuning, to adapt standard LLMs to specialized medical tasks. We then cover deployment considerations, including regulatory compliance, ethical guidelines and continuous monitoring for fairness and bias. By providing a structured step-by-step methodology, this entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
Access through your institution
Buy or subscribe
This is a preview of subscription content, access
Access options
- Purchase on SpringerLink
- Instant access to the full article PDF.
Prices may be subject to local taxes which are calculated during checkout
Code availability
Tutorial scripts are provided at https://github.com/ncbi-nlp/LLM-Medicine-Primer.
References
-
GPT-5 system card. OpenAIhttps://cdn.openai.com/gpt-5-system-card.pdf (2025).
-
Introducing Claude Opus 4.5. Anthropichttps://www.anthropic.com/news/claude-opus-4-5 (2025).
-
A new era of intelligence with Gemini 3. Googlehttps://blog.google/products/gemini/gemini-3/ (2025).
-
The Llama 4 herd: the beginning of a new era of natively multimodal AI innovation. Metahttps://ai.meta.com/blog/llama-4-multimodal-intelligence/ (2025).
-
Guo, D. et al. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature645, 633–638 (2025).
-
Brown, T. et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst.33, 1877–1901 (2020).
-
Ouyang, L. et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst.35, 27730–27744 (2022).
-
Tian, S. et al. Opportunities and challenges for ChatGPT and large language models in biomedicine and health. Brief. Bioinform.25, bbad493 (2024).
-
Saab, K. et al. Capabilities of gemini models in medicine. Preprint at https://arxiv.org/abs/2404.18416 (2024).
-
Liu, F. et al. Application of large language models in medicine. Nat. Rev. Bioeng.3, 445–464 (2025).
-
Zhou, J. et al. Large language models in biomedicine and healthcare. npj Artif. Intell.1, 44 (2025).
-
Lu, Z. et al. Large language models in biomedicine and health: current research landscape and future directions. J. Am. Med. Inform. Assoc.31, 1801–1811 (2024).
-
Liévin, V., Hother, C. E., Motzfeldt, A. G. & Winther, O. Can large language models reason about medical questions? Patterns (2023).
-
Nori, H. et al. Can generalist foundation models outcompete special-purpose tuning? Case study in medicine. Preprint at https://arxiv.org/abs/2311.16452 (2023).
-
Singhal, K. et al. Large language models encode clinical knowledge. Nature620, 172–180 (2023).
-
Hu, X. et al. Interpretable medical image visual question answering 103279 (2024)
-
Jin, Q. et al. Matching patients to clinical trials with large language models. Nat. Commun.15, 9074 (2024).
-
Wong, C. et al. Scaling clinical trial matching using large language models: a case study in oncology. In Proc. 8th Machine Learning for Healthcare Conference Vol. 219 (eds Deshpande, K. et al.) 846–862 (PMLR, 2023).
-
Wornow, M. et al. Zero-shot clinical trial patient matching with LLMs. NEJM AI2, AIcs2400360 (2025).
-
Roberts, K. Large language models for reducing clinicians’ documentation burden. Nat. Med.30, 942–943 (2024).
-
Van Veen, D. et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med.30, 1134–1142 (2024).
-
Patel, S. B. & Lam, K. ChatGPT: the future of discharge summaries? Lancet Digit. Health5, e107–e108 (2023).
-
Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature616, 259–265 (2023).
-
Acosta, J. N., Falcone, G. J., Rajpurkar, P. & Topol, E. J. Multimodal biomedical AI. Nat. Med.28, 1773–1784 (2022).
-
Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med.29, 1930–1940 (2023).
-
Omiye, J. A., Gui, H., Rezaei, S. J., Zou, J. & Daneshjou, R. Large language models in medicine: the potentials and pitfalls: a narrative review. Ann. Intern. Med.177, 210–220 (2024).
-
Zhou, H. et al. A survey of large language models in medicine: progress, application, and challenge. Preprint at https://arxiv.org/abs/2311.05112 (2023).
-
Hu, Y. et al. Improving large language models for clinical named entity recognition(2024)
-
Wang, L. et al. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. npj Digit. Med.7, 41 (2024).
-
Jin, Q. et al. Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine. npj Digit. Med.7, 190 (2024).
-
Nori, H., King, N., McKinney, S. M., Carignan, D. & Horvitz, E. Capabilities of GPT-4 on medical challenge problems. Preprint at https://arxiv.org/abs/2303.13375 (2023).
-
Liévin, V., Hother, C. E. & Winther, O. Can large language models reason about medical questions? Preprint at https://arxiv.org/abs/2207.08143 (2022).
-
Gilson, A. et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med. Educ.9, e45312 (2023).
-
Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med.31, 943–950 (2025).
-
Jin, Q., Leaman, R. & Lu, Z. Retrieve, summarize, and verify: how will ChatGPT affect information seeking from the medical literature? J. Am. Soc. Nephrol.34, 1302–1304 (2023).
-
Liu, S. et al. Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J. Am. Med. Inform. Assoc.30, 1237–1245 (2023).
-
Shaib, C. et al. Summarizing, simplifying, and synthesizing medical evidence using GPT-3 (with varying success). In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 2 (eds Rogers, A., Boyd-Graber, J. & Okazaki, N.) 1387–1407 (Association for Computational Linguistics, 2023).
-
Tang, L. et al. Evaluating large language models on medical evidence summarization. npj Digit. Med.6, 158 (2023).
-
Zhang, G. et al. Closing the gap between openmmarization. npj Digit. Med.7, 239 (2024)
-
Papineni, K., Roukos, S., Ward, T. & Zhu, W.-J. BLEU: a method for automatic evaluation of machine translation. In Proc. 40th Annual Meeting of the Association for Computational Linguistics (eds Isabelle, P., Charniak, E. & Lin, D.) 311–318 (Association for Computational Linguistics, 2002).
-
Lin, C.-Y. Rouge: a package for automatic evaluation of summaries. In Proc.Text Summarization Branches Out 74–81 (Association for Computational Linguistics, 2004).
-
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q. & Artzi, Y. BERTScore: evaluating text generation with BERT. In Proc. 8th International Conference on Learning Representations (ICLR, 2020).
-
Wang, L. L. et al. Automated metrics for medical multi-document summarization disagree with human evaluations. In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 1 (eds Rogers, A., Boyd-Graber, J. & Okazaki, N.) 9871–9889 (Association for Computational Linguistics, 2023).
-
NLLB Team. Scaling neural machine translation to 200 languages. Nature630, 841–846 (2024).
-
Preiksaitis, C. & Rose, C. Opportunities, challenges, and future directions of generative artificial intelligence in medical education: scoping review. JMIR Med. Educ.9, e48785 (2023).
-
Abd-Alrazaq, A. et al. Large language models in medical education: opportunities, challenges, and future directions. JMIR Med. Educ.9, e48291 (2023).
-
Mirza, F. N. et al. Using ChatGPT to facilitate truly informed medical consent. NEJM AI1, AIcs2300145 (2024).
-
Wang, H., Gao, C., Dantona, C., Hull, B. & Sun, J. DRG-LLaMA: tuning LLaMA model to predict diagnosis-related group for hospitalized patients. npj Digit. Med.7, 16 (2024).
-
Dagdelen, J. et al. Structured information extraction from scientific text with large language models. Nat. Commun.15, 1418 (2024).
-
Topol, E. J. As artificial intelligence goes multimodal, medical applications multiply. Science381, adk6139 (2023).
-
Chen, P.-H. C., Liu, Y. & Peng, L. How to develop machine learning models for healthcare. Nat. Mater.18, 410–414 (2019).
-
Doshi, R. et al. Quantitative evaluation of large language models to streamline radiology report impressions: a multimodal retrospective analysis. Radiology310, e231593 (2024).
-
Jiang, A. Q. et al. Mixtral of experts. Preprint at https://arxiv.org/abs/2401.04088 (2024).
-
Yang, A. et al. Qwen3 technical report. Preprint at https://arxiv.org/abs/2505.09388 (2025).
-
Minaee, S. et al. Large language models: a survey. Preprint at https://arxiv.org/abs/2402.06196 (2024).
-
Wu, C. et al. PMC-LLaMA: toward building open-1833–1843 (2024)
-
Chen, Z. et al. Meditron-70b: scaling medical pretraining for large language models. Preprint at https://arxiv.org/abs/2311.16079 (2023).
-
Nori, H. et al. From medprompt to o1: exploration of run-time strategies for medical challenge problems and beyond. Preprint at https://arxiv.org/abs/2411.03590 (2024).
-
Tang, X. et al. Medagentsbench: benchmarking thinking models and agent frameworks for complex medical reasoning. Preprint at https://arxiv.org/abs/2503.07459 (2025).
-
Health Insurance Portability and Accountability Act of 1996 (HIPAA). CDChttps://www.cdc.gov/phlp/php/re6-hipaa.html (2024)
-
Cui, H. et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat. Methods21, 1470–1480 (2024).
-
Liu, N. F. et al. Lost in the middle: how language models use long contexts. Trans. Assoc. Comput. Linguist.11, 157–173 (2024).
-
Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Vol. 1 (eds Burstein, J., Doran, C. & Solorio, T.) 4171–4186 (Association for Computational Linguistics, 2019).
-
Chen, Q. et al. Benchmarking large language models for biomedical natural language processing applications and recommendations. Nat. Commun.16, 3280 (2025).
-
Jin, D. et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Appl. Sci.11, 6421 (2021).
-
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. & Lu, X. PubMedQA: a dataset for biomedical research question answering. In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (eds Inui, K., Jiang, J., Ng, V. & Wan, X.) 2567–2577 (Association for Computational Linguistics, 2019).
-
Pal, A., Umapathi, L. K. & Sankarasubbu, M. MedMCQA: a large-scale multi-subject multi-choice dataset for medical domain question answering. In Proc. Conference on Health, Inference, and Learning Vol. 174 (eds Flores, G., Chen, G. H., Pollard, T., Ho, J. C. & Naumann, T.) 248–260 (PMLR, 2022).
-
Griot, M., Hemptinne, C., Vanderdonckt, J. & Yuksel, D. Large language models lack essential metacognition for reliable medical reasoning. Nat. Commun.16, 642 (2025).
-
Yang, Y. et al. Beyond multiple-choice accuracy: real-world challenges of implementing large language models in healthcare. Annu. Rev. Biomed. Data Sci.8, 305–316 (2025).
-
Arora, R. K. et al. Healthbench: evaluating large language models towards improved human health. Preprint at https://arxiv.org/abs/2505.08775 (2025).
-
Chiang, W.-L. et al. Chatbot Arena: an open platform for evaluating LLMs by human preference. In Proc. 41st International Conference on Machine Learning Vol. 235 (eds Salakhutdinov, R. et al.) 8359–8388 (PMLR, 2024).
-
Mukherjee, P., Hou, B., Lanfredi, R. B. & Summers, R. M. Feasibility of using the privacy-preserving large language model Vicuna for labeling radiology reports. Radiology309, e231147 (2023).
-
Liu, P. et al. Pre-train, prompt, and predict: a systematic survey of prompting methods in natural language processing. ACM Comput. Surv.55, 1–35 (2023).
-
Khattab, O. et al. DSPy: compiling declarative language model calls into state-of-the-art pipelines. In Proc. Twelfth International Conference on Learning Representations (ICLR, 2024).
-
Yuksekgonul, M. et al. Optimizing generative AI by backpropagating language model feedback. Nature639, 609–616 (2025).
-
Wei, J. et al. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst.35, 24824–24837 (2022).
-
Jin, Q. et al. Biomedical question answering: a survey of approaches and challenges. ACM Comput. Surv.55, 1–36 (2022).
-
Ji, Z. et al. Survey of hallucination in natural language generation. ACM Comput. Surv.55, 1–38 (2023).
-
Li, M. et al. Benchmarking retrieval-augmented large language models in biomedical NLP: aplication, robustness, and self-awareness. Sci. Adv.11, eadr1443 (2025).
-
Wu, K. et al. An automated framework for assessing how well LLMs cite relevant medical references. Nat. Commun.16, 3615 (2025).
-
Xiong, G., Jin, Q., Lu, Z. & Zhang, A. Benchmarking retrieval-augmented generation for medicine. In Proc. 62nd Annual Meeting of the Association for Computational Linguistics (eds Ku, L.-W., Martins, A. & Srikumar, V.) 6233–6251 (Association for Computational Linguistics, 2024).
-
Xiong, G. et al. Improving retrieval-augmented generation in medicine with iterative follow-up questions. In Proc. Pacific Symposium on Biocomputing (eds Altman, R. B., Hunter, L., Ritchie, M. D. & Klein, T. E.) 199–214 (World Scientific, 2025).
-
Fiorini, N., Leaman, R., Lipman, D. J. & Lu, Z. How user intelligence is improving PubMed. Nat. Biotechnol.36, 937–945 (2018).
-
Jin, Q., Leaman, R. & Lu, Z. PubMed and beyond: biomedical literature search in the age of artificial intelligence. EBioMedicine100, 104988 (2024).
-
Comeau, D. C., Wei, C.-H., Islamaj Doğan, R. & Lu, Z. PMC text mining subset in BioC: about three million full-text articles and growing. Bioinformatics35, 3533–3535 (2019).
-
Jin, Q., Yang, Y., Chen, Q. & Lu, Z. GeneGPT: augmenting large language models with domain tools for improved access to biomedical information. Bioinformatics40, btae075 (2024).
-
Wang, Z. et al. GeneAgent: self-verification language agent for gene-set analysis using domain databases. Nat. Methods22, 1677–1685 (2025).
-
Khandekar, N. et al. Medcalc-bench: evaluating large language models for medical calculations. Adv. Neural Inf. Process. Syst.37, 84730–84745 (2024).
-
Shi, W. et al. EHRAgent: code empowers large language models for few-shot complex tabular reasoning on electronic health records. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 22315–22339 (Association for Computational Linguistics, 2024).
-
Wang, X. et al. Self-consistency improves chain of thought reasoning in language models. In Proc. Eleventh International Conference on Learning Representations (ICLR, 2023).
-
Hou, W. & Ji, Z. Assessing GPT-4 for cell type annotation in single-cell RNA-seq analysis. Nat. Methods21, 1462–1465 (2024).
-
Cosentino, J. et al. Towards a personal health large language model. Preprint at https://arxiv.org/abs/2406.06474 (2024).
-
Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In Proc. Tenth International Conference on Learning Representations (ICLR, 2022).
-
Dettmers, T., Pagnoni, A., Holtzman, A. & Zettlemoyer, L. Qlora: efficient finetuning of quantized llms. Adv. Neural Inf. Process. Syst.36, 10088–10115 (2023).
-
Lialin, V., Deshpande, V. & Rumshisky, A. Scaling down to scale up: a guide to parameter-efficient fine-tuning. Preprint at https://arxiv.org/abs/2303.15647 (2023).
-
Biderman, D. et al. Lora learns less and forgets less. Preprint at https://arxiv.org/abs/2405.09673 (2024).
-
Chung, H. W. et al. Scaling instruction-finetuned language models. J. Mach. Learn. Res.25, 1–53 (2024).
-
Anisuzzaman, D., Malins, J. G., Friedman, P. A. & Attia, Z. I. Fine-tuning large language models for specialized use cases. Mayo Clin. Proc. Digit. Health3, 100184 (2025).
-
Wu, E., Wu, K. & Zou, J. Limitations of learning new and updated medical knowledge with commercial fine-tuning large language models. NEJM AI2, AIcs2401155 (2025).
-
Voigt, P. & Von dem Bussche, A. The EU General Data Protection Regulation (GDPR). A Practical Guide 1st edn (Springer, 2017).
-
Yang, Y. et al. A survey of recent methods for addressing AI fairness and bias in biomedicine. J. Biomed. Inform.154, 104646 (2024).
-
Omiye, J. A., Lester, J. C., Spichak, S., Rotemberg, V. & Daneshjou, R. Large language models propagate race-based medicine. npj Digit. Med.6, 195 (2023).
-
Yang, Y., Liu, X., Jin, Q., Huang, F. & Lu, Z. Unmasking and quantifying racial bias of large language models in medical report generation. Commun. Med.4, 176 (2024).
-
Huang, Y. et al. Trustllm: trustworthiness in large language models. Preprint at https://arxiv.org/abs/2401.05561 (2024).
-
Wang, B. et al. DecodingTrust: a comprehensive assessment of trustworthiness in GPT models. In Proc. 37th Conference on Neural Information Processing Systems 31232–31339 (Association for Computing Machinery, 2023).
-
Johnson, A. E. et al. MIMIC-III, a freely accessible critical care database. Sci. Data3, 160035 (2016).
-
Meskó, B. & Topol, E. J. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digit. Med.6, 120 (2023).
-
Clark, C. R. et al. Health care equity in the use of advanced analytics and artificial intelligence technologies in primary care. J. Gen. Intern. Med.36, 3188–3193 (2021).
-
Wang, S., Hu, M., Li, Q., Safari, M. & Yang, X. Capabilities of GPT-5 on multimodal medical reasoning. Preprint at https://arxiv.org/abs/2508.08224 (2025).
-
Vishwanath, K. et al. Generalist large language models outperform clinical tools on medical benchmarks. Preprint at https://arxiv.org/abs/2512.01191 (2025).
-
Huang, Y. et al. MedReflect: teaching medical LLMs to self-improve0.03687 (2025)
-
MedQA leaderboard. ValsAIhttps://www.vals.ai/benchmarks/medqa (2025).
-
Li, A. et al. QuarkMed Medical Foundation Model Technical Report. Preprint at https://arxiv.org/abs/2508.11894 (2025).
-
Shi, W. et al. MedAdapter: efficient test-time adaptation of large language models towards medical reasoning. In Proc. 2024 Conference on Empirical Methods in Natural Language Processing (eds Al-Onaizan, Y., Bansal, M. & Chen, Y.-N.) 22294–22314 (Association for Computational Linguistics, 2024).
-
Pal, A. & Sankarasubbu, M. Gemini goes to med school: exploring the capabilities of multimodal large language models on medical challenge problems & hallucinations. In Proc. 6th Clinical Natural Language Processing Workshop (eds Naumann, T., Ben Abacha, A., Bethard, S., Roberts, K. & Bitterman, D.) 21–46 (Association for Computational Linguistics, 2024).
-
Bran, A. M. et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell.6, 525–535 (2024).
-
Zakka, C. et al. Almanac—retrieval-augmented language models for clinical medicine. NEJM AI1, AIoa2300068 (2024).
-
Zhang, K. et al. A generalist vision–language foundation model for diverse biomedical tasks. Nat. Med.30, 3129–3141 (2024).
Acknowledgements
This research was supported by the Intramural Research Program of the National Institutes of Health (NIH). The contributions of the NIH author(s) are considered works of the US Government. Q.J. was also supported by the NIH Pathway to Independence Award K99LM014903. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the US Department of Health and Human Services.
Authors and Affiliations
Ethics declarations
Competing interests
The authors declare no competing interests.
Peer review
Peer review information
Nature Protocols thanks Daniel Alber and the other, anonymous, reviewer(s) for their contribution to the peer review of this work.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary information
Rights and permissions
About this article
Cite this article
Jin, Q., Wan, N., Leaman, R. et al. Tutorial: guidance on the use of large language models for medical research.
Nat Protoc (2026). https://doi.org/10.1038/s41596-026-01408-z
-
Version of record:24 July 2026
-
DOI
:https://doi.org/10.1038/s41596-026-01408-z
