Home HealthBetter Models Won’t Fix Pharma’s AI Problem — Better Terminology Will

Better Models Won’t Fix Pharma’s AI Problem — Better Terminology Will

by Staff Reporter
0 comments

At a life sciences conference earlier this year, a pharma analytics lead pulled me aside with what sounded like a simple question: “Is there an ICD-10 code for the disease we’re studying?” 

There wasn’t. As many researchers know, the reality is that this is true for thousands of medical conditions and disease states. When a condition itself cannot be captured and indexed cleanly in the data, every AI-enabled analysis or decision built on top of that data starts to wobble. 

The bottleneck in AI-enabled real-world data analysis is not computation, model architecture, or training data volume. It is the semantic layer over which the AI is trying to reason – the place where precise clinical meaning lives.

The questions today’s data still can’t ask

Most AI systems in real-world data analysis are built on standard code sets like ICD-10, SNOMED CT, and LOINC. Those standards create consistency across billing, documentation, and interoperability, but they were not designed for the kinds of research questions modern pharma teams are trying to answer. This challenge can be illustrated through three frequent failure points.

Rare disease cohorts: In many cases, patients with rare diseases are effectively invisible in datasets because there is no precise code to identify them. For example, a 2024 study in Orphanet Journal of Rare Diseases found that just 34% of 454 rare diseases could be specifically coded in ICD-10-GM. 

Severity gradients and disease activity: Clinicians and patients care deeply about distinctions between early- and late-stage disease, or mild versus moderate versus severe presentation. These clinically meaningful differences often collapse into a single standardized category in codesets used for research and patient cohorting, making it more difficult to isolate patient subgroups of interest.

Phenotyping and subtyping: Precision medicine increasingly depends on molecular and clinical specificity that billing-era terminology systems were never designed to carry. That specificity tends to hide in the free-text section of electronic health records (EHRs). A study in JMIR Medical Informatics showed that only 13% of extracted concepts from patient records (and 7% from visits) showed any overlap between structured codes and free-text notes, indicating that the vast majority of clinical information exists in one form or the other, not both. 

I spend much of my time talking with pharma leaders wrestling with these issues – trying to use the best AI tools to enhance their work, but hobbled by data far from fit-for-purpose. What makes the challenge striking is its scale and ubiquity. Critical clinical nuance routinely disappears before it ever reaches an AI model. The result is predictable. Models that look sophisticated in demos fail at the questions clinicians and researchers actually need answered.

The cohort gets defined by what was billable

What is happening on the ground is relatively straightforward. Most AI models used for life sciences analytics are trained on data normalized to billable code sets because that is the data the industry has available. Claims-derived terminology dominates many real-world data pipelines, and even when enhanced with clinical data from EHR systems, clinical fidelity is frequently lost.

This dynamic creates a subtle but consequential distortion: The cohort gets defined by what was billable, not what was clinically true.

Consider a common study design problem. A researcher wants to identify patients with mild, moderate, or severe disease. The standardized code set may simply say “L40: psoriasis”. But the clinician may have documented “psoriasis, moderate severity, with comorbid rheumatoid arthritis.” 

That specificity is exactly what determines trial eligibility, treatment response, and downstream outcomes. The nuance exists at the point of care, only to be flattened the moment it enters the claims layer.

A study of vaccine administration published in Frontiers in Digital Health illustrates the effects of this “flattening.” Using natural language processing (NLP) to analyze unstructured data in patient records, researchers found that NLP led to a 16.8% increase in the identification of vaccine administrations compared with using structured data alone, highlighting the limitations of relying on structured EHR data alone. 

The same pattern appears in oncology and rare disease research. Molecular subtypes, progression markers, and clinically meaningful modifiers often disappear into broader administrative categories that were never intended to support precision research.

AI reasoning models for research then inherit this flatness. Researchers wonder why cohort selection yields patients who are not truly study-eligible, when the underlying issue is that the source terminology never preserved the distinctions that clinicians made in the first place.

This is a precision problem, not a scale problem

The industry response to AI limitations has largely been to chase larger models and larger training sets. However, every clinically meaningful AI decision depends on something more fundamental: whether the data beneath the model can accurately express what a physician meant when documenting a patient encounter.

That fidelity affects patient identification for clinical trials, evidence synthesis from literature, real-world cohorting, and regulatory submissions backed by real-world evidence. If the clinical meaning is diluted upstream, no amount of downstream modeling sophistication can fully recover it.

This is why the conversation around trustworthy healthcare AI is increasingly less about the model itself and more about the clinical intelligence beneath it.

A terminology layer that AI can reliably reason over has several defining characteristics: 

  • First, it must be curated and clinically validated by physicians, terminologists, and subject matter experts, not assembled passively or machine-learned by brute force from billing data alone. 
  • Next, it must be provenanced. Researchers should be able to trace concepts back to trusted clinical sources and understand why a patient was included in a cohort. 
  • Finally, it must function as a connected, evolving knowledge graph rather than a static lookup table, allowing researchers to reason consistently across EHR data, registries, claims, and literature without losing fidelity.

That is the infrastructure layer that many healthcare AI deployments are still missing. This is not fundamentally a model problem. It is a precision problem. Pretending otherwise risks eroding trust in healthcare AI at exactly the moment the industry can least afford it.

For life science research leaders, the implications are immediate. Audit the data layer underneath your AI systems. Ask what terminology and clinical content the model is reasoning over, how much specificity survives end-to-end, and where that specificity gets lost. 

Demand provenance so every cohort definition and evidence synthesis can be traced back to clinical source material. And before investing in another model upgrade, evaluate whether the terminology underlying your existing systems preserves the clinical meaning you need.

The marginal value of a more advanced model, given inadequate terminology, is relatively small. The marginal value of better terminology underneath a competent model is enormous.

The quest for clinically meaningful data 

I keep thinking back to that conversation at the conference. The researcher’s question was not complicated. They simply wanted to know whether the language existed in real-world data and clinical practice to identify the patients they were studying. With precision terminology, that question becomes answerable. Without it, entire categories of research remain obscured to analysis, even when enhanced with the best AI tools.

AI in healthcare will not fail because the models are insufficiently “smart.” It will remain limited because we trained these systems on data that could not fully express what the clinician meant.

The organizations that lead the next decade of drug discovery, commercialization, and patient access will be the ones that ensure fidelity in the clinical meaning underneath their data before iterating on the next version of their AI tools.

Photo: claudenakagawa, Getty Images


Joseph Zabinski, PhD, MEM, joined IMO Health in 2025 as Senior Vice President of Product Management. He leads the Life Sciences product portfolio and go-to-market strategy. Dr. Zabinski is a published authority and thought leader in AI-driven personalization and outcome prediction using healthcare data. Previously, he oversaw the AI & Personalized Medicine business unit at OM1, a real-world evidence and technology provider. Dr. Zabinski began his career as a consultant at McKinsey, advising pharmaceutical clients on AI strategy and implementation. He holds a PhD from UNC Chapel Hill’s Gillings School of Global Public Health, where his research focused on Bayesian graphical modeling applied to large healthcare datasets. He also earned an MEM in engineering management from Dartmouth, a BS in physics and German studies from Boston College, and a Fulbright Fellowship to Austria.

This post appears through the MedCity Influencers program. Anyone can publish their perspective on business and innovation in healthcare on MedCity News through MedCity Influencers. Click here to find out how.

You may also like

Leave a Comment

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More