← Back to Blog

Ethics of EHR Data for AI Development: New Pathways for Responsible Clinical Research

by 21Stable Team

Electronic Health Records (EHRs) represent an unprecedented resource for clinical research and AI development. Yet the use of EHR data for training machine learning models raises profound ethical questions that the research community must address. A recent mixed-methods study identified four central ethical challenges—and proposes pathways forward.

The Promise and the Problem

EHR data offers extraordinary potential:

  • Large-scale, real-world patient data
  • Longitudinal records spanning years or decades
  • Diverse populations including underrepresented groups
  • Cost-effective compared to prospective trials
  • But this data was collected for clinical care, not research. This fundamental mismatch creates ethical tensions that cannot be ignored.

    Four Key Ethical Challenges

    1. Informed Consent in the Era of AI

    Traditional informed consent was designed for specific research purposes. AI development often involves:

  • Secondary use of data for purposes not envisioned at collection
  • Transfer learning where models trained on one dataset are applied to another
  • Continuous learning systems that evolve with new data
  • The Challenge: How can patients consent to uses that don't yet exist?

    Proposed Solutions:

  • Broad consent for categories of research
  • Dynamic consent with ongoing choice
  • Opt-out registries with transparent withdrawal processes
  • Patient engagement in governance decisions
  • 2. Algorithmic Fairness and Health Equity

    AI systems trained on EHR data can perpetuate and amplify existing biases:

  • Historical disparities in healthcare access reflected in training data
  • Underrepresentation of minority populations in datasets
  • Measurement bias where diagnostic codes reflect systemic inequities
  • The Challenge: Models may perform well overall but poorly for marginalized groups.

    Proposed Solutions:

  • Mandatory disaggregated performance reporting by demographic group
  • Fairness metrics as regulatory requirements
  • Community engagement in AI development
  • Diverse representation in research teams
  • 3. Privacy in the Age of Genomic Medicine

    EHRs increasingly contain genomic data, creating novel privacy risks:

  • Re-identification risks from partial genomic information
  • Family implications where one individual's data reveals information about relatives
  • Long-term risks as genomic data is inherently identifying
  • The Challenge: Traditional de-identification may be insufficient for genomic data.

    Proposed Solutions:

  • Federated learning to keep data decentralized
  • Differential privacy techniques for query results
  • Consent for genomic data sharing separate from general EHR consent
  • Right to be informed about genomic data use
  • 4. Transparency and Explainability

    Machine learning models often function as "black boxes":

  • Clinical validation requires understanding why predictions are made
  • Trust depends on interpretable reasoning
  • Accountability necessitates explainable decisions
  • The Challenge: The most accurate models may be the least interpretable.

    Proposed Solutions:

  • Explainable AI (XAI) as a regulatory requirement
  • Hybrid models combining performance with interpretability
  • Post-hoc explanation methods validated for clinical use
  • Clinician involvement in model development
  • Regulatory and Governance Frameworks

    GDPR and Secondary Use

    The European General Data Protection Regulation provides:

  • Purpose limitation: Data use must be compatible with original collection
  • Data minimization: Only necessary data should be processed
  • Rights of data subjects: Access, rectification, erasure
  • FDA Guidance on AI/ML

    The FDA's emerging framework for AI/ML-based software:

  • Good Machine Learning Practice (GMLP) principles
  • Total Product Lifecycle approach to oversight
  • Predetermined change control plans for adaptive algorithms
  • Recommendations for Research Institutions

    Governance Structures

  • **Multi-stakeholder ethics boards** including patients, clinicians, ethicists, and data scientists
  • **Data use agreements** specifying permitted and prohibited uses
  • **Audit trails** tracking how data is used and for what purposes
  • **Public reporting** of AI system performance and limitations
  • Technical Best Practices

  • **Privacy-preserving ML** including federated learning and differential privacy
  • **Fairness testing** across demographic groups before deployment
  • **Explainability layers** even when using complex models
  • **Continuous monitoring** for performance drift and bias emergence
  • The Path Forward

    The ethical challenges of using EHR data for AI development are not insurmountable—but they require deliberate, proactive engagement. The research community must:

  • Move beyond compliance-as-usual to genuine ethical reflection
  • Include patient and community voices in governance
  • Invest in technical solutions that preserve privacy while enabling innovation
  • Develop new consent frameworks for the AI era
  • The promise of AI in healthcare is immense. Realizing that promise responsibly requires navigating these ethical challenges with care, transparency, and genuine commitment to equity.

    ---

    Key Points

  • EHR data enables large-scale clinical research but was collected for care, not AI training
  • Four key challenges: informed consent, algorithmic fairness, genomic privacy, and explainability
  • Solutions require both technical innovation and governance reform
  • Patient and community engagement is essential, not optional
  • Regulatory frameworks are evolving but need strengthening
  • References

  • Mixed-Methods Study on EHR Data Ethics in AI Development. Journal of Medical Ethics. March 2026.
  • FDA. Good Machine Learning Practice (GMLP) Principles. 2021.
  • European Commission. GDPR and Secondary Use of Health Data. 2024.
  • World Health Organization. Ethics and Governance of AI for Health. 2021.