Electronic Health Records (EHRs) represent an unprecedented resource for clinical research and AI development. Yet the use of EHR data for training machine learning models raises profound ethical questions that the research community must address. A recent mixed-methods study identified four central ethical challenges—and proposes pathways forward.
The Promise and the Problem
EHR data offers extraordinary potential:
Large-scale, real-world patient dataLongitudinal records spanning years or decadesDiverse populations including underrepresented groupsCost-effective compared to prospective trialsBut this data was collected for clinical care, not research. This fundamental mismatch creates ethical tensions that cannot be ignored.
Four Key Ethical Challenges
1. Informed Consent in the Era of AI
Traditional informed consent was designed for specific research purposes. AI development often involves:
Secondary use of data for purposes not envisioned at collectionTransfer learning where models trained on one dataset are applied to anotherContinuous learning systems that evolve with new dataThe Challenge: How can patients consent to uses that don't yet exist?
Proposed Solutions:
Broad consent for categories of researchDynamic consent with ongoing choiceOpt-out registries with transparent withdrawal processesPatient engagement in governance decisions2. Algorithmic Fairness and Health Equity
AI systems trained on EHR data can perpetuate and amplify existing biases:
Historical disparities in healthcare access reflected in training dataUnderrepresentation of minority populations in datasetsMeasurement bias where diagnostic codes reflect systemic inequitiesThe Challenge: Models may perform well overall but poorly for marginalized groups.
Proposed Solutions:
Mandatory disaggregated performance reporting by demographic groupFairness metrics as regulatory requirementsCommunity engagement in AI developmentDiverse representation in research teams3. Privacy in the Age of Genomic Medicine
EHRs increasingly contain genomic data, creating novel privacy risks:
Re-identification risks from partial genomic informationFamily implications where one individual's data reveals information about relativesLong-term risks as genomic data is inherently identifyingThe Challenge: Traditional de-identification may be insufficient for genomic data.
Proposed Solutions:
Federated learning to keep data decentralizedDifferential privacy techniques for query resultsConsent for genomic data sharing separate from general EHR consentRight to be informed about genomic data use4. Transparency and Explainability
Machine learning models often function as "black boxes":
Clinical validation requires understanding why predictions are madeTrust depends on interpretable reasoningAccountability necessitates explainable decisionsThe Challenge: The most accurate models may be the least interpretable.
Proposed Solutions:
Explainable AI (XAI) as a regulatory requirementHybrid models combining performance with interpretabilityPost-hoc explanation methods validated for clinical useClinician involvement in model developmentRegulatory and Governance Frameworks
GDPR and Secondary Use
The European General Data Protection Regulation provides:
Purpose limitation: Data use must be compatible with original collectionData minimization: Only necessary data should be processedRights of data subjects: Access, rectification, erasureFDA Guidance on AI/ML
The FDA's emerging framework for AI/ML-based software:
Good Machine Learning Practice (GMLP) principlesTotal Product Lifecycle approach to oversightPredetermined change control plans for adaptive algorithmsRecommendations for Research Institutions
Governance Structures
**Multi-stakeholder ethics boards** including patients, clinicians, ethicists, and data scientists**Data use agreements** specifying permitted and prohibited uses**Audit trails** tracking how data is used and for what purposes**Public reporting** of AI system performance and limitationsTechnical Best Practices
**Privacy-preserving ML** including federated learning and differential privacy**Fairness testing** across demographic groups before deployment**Explainability layers** even when using complex models**Continuous monitoring** for performance drift and bias emergenceThe Path Forward
The ethical challenges of using EHR data for AI development are not insurmountable—but they require deliberate, proactive engagement. The research community must:
Move beyond compliance-as-usual to genuine ethical reflectionInclude patient and community voices in governanceInvest in technical solutions that preserve privacy while enabling innovationDevelop new consent frameworks for the AI eraThe promise of AI in healthcare is immense. Realizing that promise responsibly requires navigating these ethical challenges with care, transparency, and genuine commitment to equity.
---
Key Points
EHR data enables large-scale clinical research but was collected for care, not AI trainingFour key challenges: informed consent, algorithmic fairness, genomic privacy, and explainabilitySolutions require both technical innovation and governance reformPatient and community engagement is essential, not optionalRegulatory frameworks are evolving but need strengtheningReferences
Mixed-Methods Study on EHR Data Ethics in AI Development. Journal of Medical Ethics. March 2026.FDA. Good Machine Learning Practice (GMLP) Principles. 2021.European Commission. GDPR and Secondary Use of Health Data. 2024.World Health Organization. Ethics and Governance of AI for Health. 2021.