At a life sciences conference earlier this year, a pharma analytics lead pulled me aside with a question that, on its face, seemed mundane: “Is there an ICD-10 code for the specific disease state we’re studying?”
The answer was a definitive no.
This is not an isolated incident; it is a pervasive, structural reality in modern medicine. For thousands of clinical conditions—particularly rare diseases, nuanced molecular subtypes, and complex severity gradients—there is no precise, standardized index. When a condition cannot be captured or indexed cleanly within the data architecture, every AI-enabled analysis or decision built upon that foundation begins to wobble.
The bottleneck in AI-enabled real-world data (RWD) analysis is not a lack of computation power, architectural sophistication, or the volume of training data. It is the semantic layer—the precise clinical meaning—that the AI is attempting to reason over.
The Invisible Patients: The Core Failure of Coding Standards
Most current AI systems for life sciences are built on the bedrock of standard code sets: ICD-10 (International Classification of Diseases), SNOMED CT, and LOINC. While these standards are the backbone of hospital billing, administrative documentation, and cross-system interoperability, they were never intended to serve as the engine for precision research.
The misalignment between billing-era standards and modern clinical requirements manifests in three distinct failure points that compromise the integrity of AI outputs.
1. The Rare Disease "Dark Matter"
Patients suffering from rare conditions are frequently invisible in datasets simply because they cannot be classified. A 2024 study in the Orphanet Journal of Rare Diseases revealed a sobering statistic: only 34% of 454 rare diseases surveyed could be specifically coded in the ICD-10-GM system. For the remaining 66%, the data is effectively a void. When an AI model is trained to identify a cohort, if the primary identifier does not exist, the patient is either omitted or lumped into a broad, inaccurate "catch-all" category, leading to statistically significant bias.
2. The Erosion of Severity Gradients
Clinical decision-making is defined by nuance: Is a patient’s condition early-stage or advanced? Is the presentation mild, moderate, or severe? In clinical practice, these distinctions dictate treatment pathways. However, in the standardized codesets used for research, these crucial differences often collapse into a single, monolithic category. When an AI model tries to predict patient outcomes, it lacks the resolution to differentiate between a patient who is thriving and one who is deteriorating because the input data has been stripped of its clinical fidelity.
3. The Great Divide: Structured vs. Unstructured Data
Precision medicine demands molecular and clinical specificity that billing systems were never designed to carry. Consequently, the most valuable clinical insights often remain trapped in the free-text narratives of electronic health records (EHRs). A 2025 study in JMIR Medical Informatics highlighted a massive disconnect: only 13% of concepts extracted from patient records showed overlap between structured codes and free-text notes. This indicates that the vast majority of clinical information exists in a silo, unavailable to models that only read structured, billable fields.
The "Billable" Distortion: A Chronology of Data Degradation
To understand why our current AI models are underperforming, we must examine how data flows from the point of care to the research dataset.
- The Point of Care (High Fidelity): A clinician observes a patient and documents a nuanced diagnosis, such as "psoriasis, moderate severity, with comorbid rheumatoid arthritis."
- The Billing Layer (The Compression): The documentation is processed by administrative staff. To satisfy insurance requirements, the complex, multi-faceted description is reduced to a single code, such as "L40: Psoriasis." The severity and comorbidity—the very data needed for trial eligibility—are discarded.
- The Data Pipeline (The Normalization): This "flattened" code is ingested into a data warehouse. Because it is the most standardized format available, it becomes the primary input for AI training sets.
- The Research Outcome (The Hallucination of Precision): Researchers use the AI to identify a study cohort. The model, trained on the "flattened" data, produces a cohort that is riddled with noise—patients who do not actually fit the clinical criteria because the underlying data never preserved the distinctions that the physician originally made.
A study in Frontiers in Digital Health provided a vivid illustration of this "flattening" effect. When researchers used Natural Language Processing (NLP) to analyze unstructured patient notes, they discovered a 16.8% increase in identified vaccine administrations compared to relying on structured EHR codes alone. This demonstrates that the "ground truth" is not in the codes; it is in the clinical narrative.

Implications for the Life Sciences Industry
The current industry reflex is to chase "more": more data volume, more parameter counts in LLMs, and more aggressive training cycles. However, this is a distraction. If the clinical meaning is diluted at the source, no amount of downstream modeling sophistication can recover it.
The Trust Deficit
The erosion of trust in healthcare AI is an existential threat. When models look sophisticated in laboratory demos but fail in the real-world clinical environment, stakeholders—regulators, clinicians, and patients—lose faith. This failure is often incorrectly attributed to the AI’s "logic," when in fact, the fault lies in the impoverished semantic infrastructure feeding the model.
The Regulatory and Research Cost
For life sciences organizations, the price of this data degradation is high. Clinical trial recruitment, evidence synthesis, and real-world evidence (RWE) submissions to regulatory bodies all depend on the ability to distinguish between patient subgroups. If an AI system cannot reliably identify a specific subtype, the resulting clinical trial may be doomed by poor recruitment or misinterpreted results.
Defining the Path Forward: A Call for Semantic Fidelity
To move beyond the current impasse, the life sciences industry must shift its focus from "model-centric" AI to "data-centric" AI. The goal is to build a terminology layer that AI can reliably reason over.
Such a layer must possess three defining characteristics:
- Clinical Granularity: It must support the precise, multi-dimensional descriptors used by clinicians in daily practice.
- Contextual Provenance: It must preserve the history of how a term was used, ensuring that researchers can trace a cohort definition back to the original clinical documentation.
- Interoperability with Unstructured Data: It must bridge the gap between structured billing codes and the wealth of information hidden in physician notes.
A Roadmap for Leaders
For research leaders, the mandate is clear. It is time to stop viewing AI as a "plug-and-play" solution and start treating the data layer as a strategic asset:
- Audit the Foundation: Examine the terminology and clinical content your AI is currently consuming. Where is the specificity being lost?
- Demand Provenance: Ensure that every cohort definition can be audited and traced back to the clinical source material.
- Prioritize Fidelity over Scale: Before investing in the next iteration of a model, evaluate whether your existing terminology system preserves the clinical meaning required for your specific research question.
The marginal value of an incrementally better model, when fed poor-quality, flattened data, is negligible. Conversely, the marginal value of injecting high-fidelity, clinically meaningful terminology into a competent model is immense.
Conclusion: The Quest for Meaning
The researcher who pulled me aside at that conference was not asking for a miracle; they were asking for the basic language required to conduct valid science. With precision terminology, the complex questions of modern medicine become answerable. Without it, entire categories of human disease remain obscured, buried under the weight of administrative shorthand.
The organizations that will lead the next decade of drug discovery, commercialization, and patient access will not necessarily be those with the largest models. They will be the ones that have the courage to fix the "clinical meaning" beneath their data. AI in healthcare will not fail because the models are insufficiently smart; it will struggle until we stop forcing those systems to learn from data that cannot express what the clinician actually meant.
In the final analysis, the pursuit of better AI is, fundamentally, a pursuit of better clinical truth. It is time for the life sciences industry to stop settling for the billable, and start reaching for the meaningful.
