The Code We Cannot Read

The Code We Cannot Read

When Maya sat in the bright, sterile waiting room of a prominent university clinic, she held a folder containing a single piece of paper. On it was a score. It was a Polygenic Risk Score, the latest triumph of modern medicine, a number calculated by looking at millions of tiny variations across her DNA to predict her lifetime risk of developing breast cancer.

The number on the page said her risk was average. Safe. Meanwhile, you can explore related developments here: Why tracking down the cyclospora outbreak turned into a total mess.

Her doctor smiled, offering the kind of reassurance that only data-driven certainty is supposed to provide. But Maya felt a cold knot in her stomach. She looked at her own hands, then thought of her mother and her grandmother, both of whom had developed aggressive tumors before the age of fifty.

What the doctor did not say—what the algorithm itself could not articulate—was a quiet, systemic truth. Maya is Black. The vast, sweeping data sets used to build that predictive tool, the genetic libraries that taught the software how to spot the early warning signs of disease, were built almost entirely from the blood of people of European descent. To see the full picture, we recommend the recent analysis by Psychology Today.

Maya was being measured against a ruler that wasn't made for her.


The Mirage of Perfect Data

We have been told a beautiful story about the future of medicine. It goes like this: one day, you will walk into a clinic, spit into a tube, and a computer will map your destiny. It will catch the heart attack a decade before your arteries begin to stiffen. It will flag the diabetes before your blood sugar ever spikes.

This is the promise of genomic prediction. It bypasses the old, blunt instruments of family history and lifestyle questionnaires, looking instead at the hardwired blueprint of life itself.

But science has a representation problem, and it is warping the glass through which we view human health.

To understand why these tools stumble, we have to look at how they learn. A polygenic risk score does not look for a single, catastrophic mutation like the BRCA gene. Instead, it scans the entire genome, tracking hundreds of thousands of microscopic variations called single-nucleotide polymorphisms. Think of them as typographical errors in a massive, multi-volume encyclopedia. Individual typos mean very little. But if you accumulate thousands of specific typos in specific chapters, the plot of the book changes. You get sick.

Scientists find these patterns through Genome-Wide Association Studies. They compare the DNA of ninety thousand people who have a disease against ninety thousand people who do not. The computer spots the typos unique to the sick group, builds a mathematical model, and applies it to the rest of us.

It is brilliant. It is elegant. It is also deeply flawed.

Nearly eighty percent of the individuals included in these massive genetic research databases are of European ancestry. Yet, people of strictly European descent make up a fraction of the global population. We have built a cutting-edge telescope, pointed it at a single corner of the night sky, and declared that we understand the entire universe.


When Math Inherits Human History

DNA is a historical record. It carries the map of where our ancestors walked, where they bottlenecked, and how they survived. Because humanity began in Africa, populations of African descent possess the highest genetic diversity on earth. European populations, descended from a much smaller group of individuals who migrated northward millennia ago, share a much more uniform genetic baseline.

When an algorithm trained on European genomes tries to read the DNA of a person of African, Hispanic, or Asian descent, it gets confused.

It sees variations it does not recognize and misinterprets them as dangerous flags. Or, far worse, it misses the genuine warning signs entirely because those specific genetic markers were never present in the European training data.

Consider a hypothetical patient named David, a second-generation line cook of Japanese descent. David goes to a clinic requesting a genetic screen for coronary artery disease, a condition that took his father too soon. The tool analyzes his saliva and returns a low-risk profile. David breathes a sigh of relief. He continues his life, unaware that the specific genetic variants that trigger heart disease in East Asian populations are fundamentally different from those that trigger it in someone from Edinburgh or Munich. The tool looked for the errors it knew how to find. It found nothing. It left David exposed.

This is not a theoretical glitch. It is a measurable statistical drift.

Research shows that the predictive accuracy of polygenic risk scores drops precipitously when applied to non-European populations. For individuals of African descent, the accuracy of these scores can be up to five times lower than for their white peers.

We are inadvertently building a two-tiered system of preventive medicine. One tier receives hyper-personalized, predictive warnings that save lives. The other tier receives a lottery ticket wrapped in a statistic.


The Weight of the Unknown

The danger here is not just bad math; it is the false sense of absolute certainty that numbers provide.

When a test comes back with a percentage, we believe it. Doctors believe it. Insurance companies believe it. If a genetic risk tool decides a patient is at low risk for colorectal cancer based on flawed, biased data, that patient might be told they can delay their routine colonoscopy. The algorithm becomes a gatekeeper, shutting out the very people who need intervention the most.

It feels lonely to sit in that gap. To know that the most advanced medical technology on earth views your biology as an anomaly, an outlier to be adjusted for with a statistical correction factor.

The researchers building these tools are not malicious. They are running against the clock, using the data that is most readily available to them. The historical repositories of genetic information are overwhelmingly Western because that is where the initial funding, the major universities, and the biobanks were concentrated.

But awareness of a bias does not automatically cure it.

Fixing this requires more than just a software patch. It requires a profound, expensive, and logistically daunting restructuring of how genetic material is gathered. It means building trust in communities that have historically been exploited by medical research. It means going out into the world and collecting diverse blood samples with the same urgency and capital that was used to map the first human genome.

Until that happens, the data remains a mirror reflecting only a portion of humanity.


Reading the Whole Book

We cannot abandon the promise of genetic prediction. The potential to stop chronic illness before it takes root is too profound to discard. But we must learn to hold these numbers with a degree of humility.

A polygenic risk score is a single sentence in a biography that is still being written. It cannot replace the lived reality of a patient’s family history, the air they breathe, or the systemic pressures of their environment.

Maya left the clinic that afternoon with her folder tucked firmly under her arm. She didn't throw the paper away, but she didn't let it lull her into a false complacency either. She knew her body, her history, and her lineage better than the code could ever hope to.

We will eventually map the variations of every population on this planet, filling in the blank pages of our collective biological manual. But until that day arrives, we must remember that a tool is only as wise as the data we feed it, and right now, our finest instruments are still profoundly, dangerously blind to the true scale of human diversity.

EH

Ella Hughes

A dedicated content strategist and editor, Ella Hughes brings clarity and depth to complex topics. Committed to informing readers with accuracy and insight.