The pharmaceutical industry has never had more data, but AI-powered drug development will not advance on data volume alone. Clinical trials, electronic health records, imaging archives, pathology reports, genomic repositories, laboratory systems and patient registries generate petabytes of information every year. Add the growing volume of patient-generated and real-world data, and the industry has more than enough raw material to fuel the next generation of AI-driven discovery.
Yet AI continues to struggle to connect observations into meaningful evidence. The reason is simple: drug development is not suffering from a shortage of data. It is suffering from a shortage of usable knowledge.
For decades, healthcare has optimized for data collection, storage and exchange. We have become remarkably good at capturing observations, but far less effective at preserving the relationships that give those observations meaning. A patient’s health journey is continuous, yet the data describing that journey is often captured as fragmented snapshots rather than a complete clinical narrative.
A clinical trial database captures one chapter of the journey. An electronic health record captures another. A radiology archive stores images. A pathology system records diagnoses. A genomics platform identifies mutations. A claims database documents utilization and reimbursement.
Each system faithfully records its own perspective. The challenge lies in stitching those perspectives together into a complete patient journey.
Too often, AI is expected to reconstruct that journey from disconnected tables and unstructured text, written in different clinical languages and stripped of the context that originally made them meaningful. Outside pilot programs, the results are inconsistent, difficult to reproduce and challenging to translate into regulatory-grade evidence.
Researchers increasingly recognize that improving data quality and interoperability may be more valuable than simply increasing data volume. But one additional attribute deserves equal attention: preserving meaning.
Without that connective tissue, AI may find patterns, but it cannot reliably generate evidence that researchers, clinicians or regulators can trust.
From datasets to evidence networks
Instead of thinking about training healthcare AI on collections of datasets, we should start thinking about building evidence networks.
An evidence network is a connected representation of a patient’s health journey, where every observation retains its meaning, context and relationship to every other relevant observation.
This matters because AI does not reason over isolated data points. It reasons over evidence. And evidence is created only when observations remain connected.
Three elements are essential to building a reusable evidence network.
1. Semantic harmonization: speaking the same clinical language – Records from different sources within the same hospital may use different terminology. One dataset may record a diagnosis as “heart attack.” Another may say “myocardial infarction.” A third may record “acute MI,” while a fourth stores an ICD-10 code.
A clinician immediately understands that these are different expressions of the same clinical concept. An AI model may not.
Semantic harmonization provides a common clinical language that allows the model to understand that different representations refer to the same underlying event while preserving clinically important nuances.
2. Context: turning numbers into evidence – Every clinician knows that no laboratory value should be interpreted in isolation.For example, a hemoglobin level of 10.2 g/dL may be concerning for one patient but an expected treatment effect for a patient undergoing chemotherapy. If the patient is recovering from major surgery, the same value may reflect a normal course of recovery when viewed against prior blood tests.
Context includes everything surrounding a clinical observation, such as disease stage, prior treatments, patient characteristics, care setting, timing and concurrent events. Without this information, AI sees numbers. With it, AI can interpret those numbers as clinical evidence.
Temporality is equally important. Elevated liver enzymes before treatment tell a very different story from elevated liver enzymes two weeks after treatment begins. Medical practice is fundamentally a sequence of events, and preserving that sequence is essential to building trustworthy AI models.
3. Multi-modal relationships: connecting the patient journey – Imagine a patient enrolled in an oncology study. A CT scan identifies a lung lesion. A biopsy confirms adenocarcinoma. Genomic sequencing detects an EGFR mutation. The patient receives targeted therapy. Follow-up imaging demonstrates tumor shrinkage.
The value of these observations does not lie in the individual records alone. It lies in the fact that they describe the same patient, the same disease episode and the same treatment journey.
A multi-modal relationship is the thread that connects the radiology image to the biopsy that confirmed it, the genomic mutation that explained it, the therapy that targeted it and the outcome that ultimately determined whether it mattered.
If this thread breaks during data extraction, harmonization or aggregation, AI ends up interpreting disconnected facts instead of mapping a patient’s journey.
The promise of AI-powered drug development
AI models are increasingly expected to identify biomarkers, optimize clinical trial design, generate external comparators, discover safety signals and support regulatory decision-making. For these outputs to carry regulatory weight, the underlying evidence must be traceable, clinically interpretable and reproducible across datasets and settings. Regulatory agencies are also exploring AI-enabled clinical trial processes to shorten development timelines.
Larger models and greater computing power will undoubtedly improve these capabilities. But no amount of computational sophistication can reliably recreate information that was lost during the data lifecycle.
The next breakthrough in AI-powered drug development is therefore unlikely to come solely from a new algorithm or a larger foundation model. It will come from specialized, multilayered AI solutions that connect clinical observations through semantic harmonization, clinical context, temporal relationships and multi-modal links.
Representativeness, privacy and data residency will remain critical considerations as these evidence ecosystems expand across regions and populations. But the foundation must come first: evidence networks that preserve the clinical context needed for robust scientific and regulatory decisions.
For years, healthcare has focused on collecting more data. The next era should focus on making that data more meaningful. AI does not create evidence from data alone. It reasons over evidence that humans have carefully connected, contextualized and preserved. That context-preserving architecture may be the missing ingredient that allows AI-powered drug development to move from promise to practice.
Photo: ClaudioVentrella, Getty Images
Narasimha Kumar, Global Head of Technology and Data Services, BC Platforms, brings 25+ years of experience in technology consulting, AI product management, and compliance across life sciences and healthcare. Previously Chief Product and Commercial Officer at Datafoundry AI and senior leader at PAREXEL International, Narashimha specialises in AI-driven platforms and cloud transformation.
This post appears through the MedCity Influencers program. Anyone can publish their perspective on business and innovation in healthcare on MedCity News through MedCity Influencers. Click here to find out how.
