An Investigation of Memorization Risk in Healthcare Foundation Models
Paper in proceeding, 2025

Foundation models trained on large-scale de-identified electronic health records (EHRs) hold promise for clinical applications. However, their capacity to memorize patient information raises important privacy concerns. In this work, we introduce a suite of black-box evaluation tests to assess privacy-related memorization risks in foundation models trained on structured EHR data. Our framework includes methods for probing memorization at both the embedding and generative levels, and aims to distinguish between model generalization and harmful memorization in clinically relevant settings. We contextualize memorization in terms of its potential to compromise patient privacy, particularly for vulnerable subgroups. We validate our approach on a publicly available EHR foundation model and release an open-source toolkit to facilitate reproducible and collaborative privacy assessments in healthcare AI.

Author

Sana Tonekaboni

Broad Institute

Lena Stempfle

Chalmers, Computer Science and Engineering (Chalmers), Data Science and AI

Adibvafa Fallahpour

University Health Network

Walter Gerych

Worcester Polytechnic Institute

Marzyeh Ghassemi

Massachusetts Institute of Technology (MIT)

Advances in Neural Information Processing Systems

10495258 (ISSN)

Vol. 38 11709-11734
9798331338275 (ISBN)

39th Conference on Neural Information Processing Systems, NeurIPS 2025
San Diego, USA,

Subject Categories (SSIF 2025)

Computer Sciences

Computer Systems

More information

Latest update

9/4/2026 7