Who Gets Left Out of the Data: Clinical Trials' Representation Crisis and the Patients Paying the Price
The drug approval process in the United States is predicated on a fundamental assumption: that the patients enrolled in clinical trials are a reasonable proxy for the patients who will eventually receive the therapy. That assumption, it turns out, has never been particularly well-founded — and the scientific costs are now impossible to ignore.
Decades of enrollment data reveal a persistent, structural imbalance in who participates in clinical research. Black, Hispanic, Asian, and Native American patients are chronically underrepresented relative to their share of the US population. Rural communities, where nearly 20 percent of Americans live, contribute a disproportionately small fraction of trial participants. Elderly patients — among the heaviest consumers of prescription medicine — are routinely excluded by eligibility criteria designed to minimize medical complexity rather than reflect clinical reality. And low-income individuals, for whom the logistical and financial demands of trial participation can be prohibitive, are largely absent from the datasets that regulators use to approve new therapies.
The result is a pharmacological record that is, in meaningful ways, a portrait of a patient who does not represent most of the people who will eventually take these drugs.
The Scientific Stakes of a Skewed Sample
The diversity deficit in clinical trials is frequently discussed as an ethical failure, and it is. But framing it exclusively as a matter of fairness obscures an equally serious problem: the data produced by homogeneous trial populations is scientifically compromised in ways that have concrete downstream consequences.
Pharmacokinetics — how the body absorbs, distributes, metabolizes, and eliminates a drug — varies meaningfully across populations. Genetic variants that influence drug metabolism, such as polymorphisms in the CYP450 enzyme family, are distributed differently across ancestral groups. A therapy calibrated on a predominantly white European-ancestry cohort may be systematically under- or over-dosed for patients whose metabolic profiles differ. These are not hypothetical concerns. Documented disparities in drug response across populations have been identified for anticoagulants, antidepressants, antihypertensives, and oncology agents, among others.
Beyond pharmacokinetics, disease biology itself can vary across populations. Certain cancers present with different mutational landscapes depending on ancestry. Cardiovascular disease follows distinct epidemiological patterns across racial and ethnic groups. When trials enroll narrow populations, they generate efficacy and safety signals that may not generalize — and regulators, clinicians, and patients have limited means to know where those signals break down until real-world use reveals the gaps.
How Eligibility Criteria Quietly Narrow the Field
One of the less visible mechanisms driving enrollment disparities is the design of eligibility criteria themselves. Trial protocols are constructed to maximize the signal-to-noise ratio of experimental results — a scientifically defensible goal that, in practice, systematically filters out complexity. Patients with comorbidities are excluded. Those on concomitant medications are screened out. Kidney and liver function thresholds eliminate patients whose organ function reflects the normal variation of aging or chronic disease.
These exclusions do not occur randomly. Comorbidity burden is higher among lower-income populations and communities of color, in part because of differential access to preventive care. The patients most likely to be excluded by strict eligibility criteria are, disproportionately, the same patients already underrepresented in clinical research.
The cumulative effect is a trial population that is healthier, wealthier, whiter, and more urban than the patient population that will eventually receive the approved therapy. The drug that emerges from this process has been tested, in a meaningful sense, on a cohort that does not reflect the full range of people who will take it.
Geography as a Barrier
Site concentration is another structural driver of enrollment inequality that receives insufficient attention. The majority of clinical trial activity in the United States is clustered in academic medical centers located in major metropolitan areas. Patients in rural counties — who already face elevated rates of chronic disease and reduced access to specialty care — must frequently travel hundreds of miles to participate in trials, absorb associated costs, and arrange time away from work. For many, participation is simply not feasible.
This geographic concentration also reflects a feedback loop in research infrastructure. Institutions with the resources to run trials attract the patients who can reach them, which reinforces the patterns of participation that have characterized the field for generations. Breaking that loop requires not just goodwill but deliberate structural intervention.
Emerging Strategies to Scramble Traditional Recruitment
There is genuine movement within the research community to address these entrenched patterns, though the scale of the response remains uneven relative to the scope of the problem.
Decentralized clinical trials — a model accelerated by necessity during the COVID-19 pandemic — represent one of the more promising structural shifts. By enabling remote participation through telemedicine visits, home-based sample collection, and local laboratory partnerships, decentralized designs can meaningfully reduce the geographic and logistical barriers that exclude rural and low-income patients. Several large sponsors have expanded decentralized components in ongoing studies, and the FDA has issued guidance supporting the approach.
Community-based participatory research offers a complementary strategy, one that engages historically underrepresented communities as partners in trial design rather than targets of recruitment campaigns. This model acknowledges that trust deficits — rooted in documented historical abuses of minority research participants — cannot be overcome through marketing alone. Building durable relationships with community health organizations, faith institutions, and local clinicians takes time and sustained investment, but the evidence suggests it produces more equitable and more stable enrollment.
Regulatory pressure is also intensifying. The FDA Omnibus Reform Act of 2022 included provisions requiring sponsors to submit diversity action plans for Phase III trials and certain Phase II studies, outlining how they intend to enroll participants reflective of the disease population. Whether this mandate produces substantive change or generates compliance theater will depend heavily on how rigorously the agency enforces it.
Some biotechs are embedding diversity metrics into trial design from the outset, working with biostatisticians to ensure that enrollment targets are powered to detect differential effects across subgroups. This approach treats population heterogeneity not as a confound to be minimized but as a scientific variable to be measured — a reframing that has implications for how trial results are interpreted and how labels are written.
The Cost of Getting This Wrong
When a therapy reaches the market without adequate representation in its trial data, the consequences do not distribute evenly. Patients whose biology, pharmacogenomics, or disease context differs from the trial population may experience attenuated efficacy or elevated adverse event rates. Physicians treating those patients have limited evidence to guide dose adjustments or therapeutic alternatives. Post-market surveillance may eventually surface the problem — but that process takes years, and in the interim, real patients are receiving therapies calibrated for someone else.
For gene therapies and genomic medicines in particular, where treatment is often irreversible and the biological mechanisms are closely tied to individual genetic architecture, the stakes of population-level data gaps are especially high. A gene therapy validated predominantly in one ancestry group carries uncharacterized risk when administered to patients whose genomic context differs in ways that were never studied.
The diversity problem in clinical trials is, at its core, a data integrity problem. A drug approval built on a narrow sample is a scientific conclusion drawn from incomplete evidence. Correcting this is not merely an act of equity — it is a prerequisite for the kind of rigorous, generalizable science that medicine depends on. The field has known this for a long time. The question now is whether the structural reforms underway are moving fast enough to matter.