STAT recently reported that despite the National Institutes of Health’s All of Us Research Program persuading 98% of its nearly 750,000 participants to share their electronic health records, more than 300,000 of those participants have no EHR data in the program’s database at all. The program’s latest data release attempts to close that gap by piggybacking on the clinical data-sharing networks that hospitals already use to move records between health systems.
It’s a sound fix, and it’s worth pausing on why it’s necessary in the first place, because the explanation says something important about where American health data infrastructure actually stands in 2026, and it isn’t where most people assume.
For a decade, the working theory of interoperability was essentially a plumbing problem: connect the pipes, and the data would flow. That theory was largely correct, and the plumbing has, in fact, been built. The 21st Century Cures Act’s information-blocking rules established a patient’s right to their own data. Certified EHR vendors were required to expose standardized FHIR APIs. National networks now let hospitals query one another’s records with something close to the ease of a phone lookup. The people who did this work, inside federal health IT policy offices, health systems, and the EHR vendor community, accomplished something genuinely difficult, over real institutional resistance, and it shows in numbers like All of Us’s 98% consent rate. Patients are willing. The wires are, for the most part, connected.
So why do 300,000 records still not show up?
Because retrieving a record and having usable data from it are two different problems, and the industry solved the first one before it finished the second. A hospital can now, in principle, transmit a patient’s full chart to a research database in seconds. What arrives, though, is rarely a clean, structured dataset. It’s a mixture: some discrete FHIR resources, some scanned referral letters, some faxed prior-authorization forms, some PDF exports of visit notes with fields that vary hospital to hospital, EHR to EHR, decade to decade. A patient’s chart may span several different systems, several different document formats, and years of care delivered before anyone thought to make records machine-readable in the first place. The pipe works. The contents of what flows through it often don’t arrive in a form any research database can use.
That’s the part of interoperability that rarely comes up in policy conversations, because it isn’t a policy problem. It’s a data-engineering problem; and it happens to be one that’s considerably closer to being solved than most health system leaders, researchers, or even technologists tend to assume.
The tools for converting genuinely messy, inconsistent, real-world documents into structured, validated, research-ready data have matured substantially over the past two years. What makes this possible isn’t a single model breakthrough. It’s the pairing of probabilistic extraction (AI systems that can read a scanned, handwritten, or inconsistently formatted document roughly the way a trained person would) with deterministic validation layers that check every extracted field against defined rules before it is ever accepted into a dataset. That pairing matters because probabilistic AI alone isn’t reliable enough at the accuracy thresholds research and clinical use demand, and rules-based systems alone can’t handle the sheer format diversity of real-world records. Together, they can. And the gap between the two, once measured in years of custom engineering per health system, is now closing in weeks.
That matters for three reasons that go directly to what All of Us, and every research program, health system, and payer facing a version of the same 300,000-record problem, actually needs.
It’s faster than the timelines most institutions have priced in. Turning a backlog of unstructured records into structured, usable data no longer requires a multi-year integration project or a rip-and-replace of legacy systems. It can run alongside the systems that already exist, ingesting whatever format a record arrives in (fax, scan, PDF, or structured export) without asking a health system to change how or where it stores anything.
It’s more accurate than the manual alternative, not less. The instinct in research and regulatory settings is to treat automation as a tradeoff against accuracy, with manual chart abstraction as the safe default. In practice, deterministic validation layers can flag every field that falls outside expected bounds for human review, producing a documented audit trail showing exactly what was extracted, from what source, and at what confidence level, demonstrating a more rigorous evidentiary record than most manual abstraction processes generate today.
And it’s compatible with the security and provenance standards a program like All of Us has to meet. None of this requires records to leave controlled environments, exposing data to consumer AI tools, or relaxing the compliance postures, such as SOC 2 and HIPAA, that regulated research programs already operate under. The validation layer is, in fact, what gives auditors and institutional review boards the traceability they need to trust automated extraction in the first place.
None of this is a criticism of All of Us’s approach. Piggybacking on existing data-sharing networks to retrieve more records is the right move, and it will close part of the gap. But it will eventually run into the same wall that every large-scale EHR retrieval effort hits: some meaningful share of what comes back will not be clean, structured data. It will be the messy remainder, the scanned document, the legacy export, or the record from a system that never fully adopted FHIR, that retrieval alone cannot fix.
The honest lesson from the 300,000-record gap isn’t that interoperability failed. It’s that interoperability, as the field has largely defined it (can a record move from system A to system B?) has substantially succeeded, and the work that remains is a different, more tractable problem than the one the field spent the last decade solving. Turning what arrives into what researchers can actually use is no longer the multi-year undertaking it was treated as five years ago. It is closer to a solved engineering problem than most people building health data infrastructure realize. The research programs willing to treat it that way will close gaps like this one far faster than the last decade of policy work did.
Photo: Bigstock
Adam Cohen is the co-founder of Morph Services, Inc., a healthcare technology company that transforms unstructured clinical documents into trusted, structured data for health systems, payers, and other regulated organizations. With more than 15 years of experience spanning healthcare operations, interoperability, AI-enabled data processing, credentialing, and enterprise software, he has helped organizations modernize complex workflows while maintaining the accuracy, governance, and compliance required in heavily regulated environments. Prior to co-founding Morph, Cohen held executive leadership roles in healthcare technology companies focused on telemedicine, credentialing, and digital transformation, supporting many of the nation’s largest health systems and healthcare organizations. He is a frequent writer and speaker on healthcare interoperability, AI, and the practical challenges of converting fragmented clinical information into usable data that improves care, research, and operational performance.
This post appears through the MedCity Influencers program. Anyone can publish their perspective on business and innovation in healthcare on MedCity News through MedCity Influencers. Click here to find out how.
