Real Data Better AI Is That Really So

A Road to Bilbao reflection by Dr Selena Baset
During the first Real Data, Better AI Foundation Circle, held on 30 June, we began exploring a deceptively simple proposition: does better real-world data necessarily lead to better artificial intelligence?
Dr Selena Baset, a member of the Health Data Expert Hub and a participant in that conversation, takes the question back to first principles. Her reflection is not intended as a final position, but as an invitation to examine what we really mean by "real data", what should qualify as "better AI", and what must happen between the two.
As we continue on the Road to Bilbao, we invite our community to engage with Selena's argument and help turn the Real Data, Better AI proposition into a practical implementation agenda.
Real Data, Better AI: Is That Really So
I go back and forth on it.
If real data were enough on its own to produce better AI, you would not be reading this piece. The name gets one thing right: AI is only as good as the real data it feeds on. Period.
The name does not spell out the mechanics. It tells us that the second thing follows from the first, while leaving out how — or under what conditions — that promise holds. That gap is the hard, unspoken part, and the raison d'être of this initiative.
Before answering how, we need to agree on the definitions. What is real data, and what makes AI good — or better? What follows is my working view. I hold it loosely and welcome other readings.
Starting with real data, the term that first comes to mind in healthcare is real-world data as we know it. RWD is a large part of what we mean, but stopping there misses the point. I prefer to think more broadly about data as it is first captured from real people, real clinical environments and real care processes — together with the context and provenance needed to understand how it came into being.
What about better AI? This is the harder question. It looks obvious until we try to pin it down.
As our use of AI matures, different perspectives, patterns of use and a hard-won awareness of its limitations are beginning to converge into a kind of collective wisdom, even if we continue to describe it in different words.
I recently came across an acronym that captures something important, though I haven't been able to trace its original source. For AI outputs in healthcare to be trusted, they should be TRUE: Tracked, Reproducible, Understandable and Ethical.
This is best regarded as a useful test rather than an established framework. But it gives us a way to inspect the gap: what does it take for real data to produce a TRUE AI experience?
The ethical dimension depends on public policy, governance, clearly assigned responsibilities and meaningful accountability. Tracking, reproducibility and understandability depend on more than the data alone. They also require documented provenance, appropriate standards, transparent transformations and systems designed to preserve meaning throughout the data lifecycle.
The good news is that we do not need to reinvent the wheel. Many of the foundational pieces are already in place. What we need is disciplined, systematic adoption.
The first piece is FAIR: data and metadata that are Findable, Accessible, Interoperable and Reusable. FAIR is a well-established set of guiding principles, with increasingly mature implementation approaches in the life sciences. Yet it is still too often mistaken for an additional layer that can simply be applied at the end.
It cannot. FAIRness must be designed into how data is captured, described, governed, and maintained. An ecosystem determines whether its data can be discovered, interpreted and responsibly reused beyond the purpose for which it was initially collected.
The second piece is standards. Standards help data collected in one setting retain its structure and meaning when exchanged, transformed, or analysed elsewhere.
CDISC supports the organisation and regulatory use of clinical research data. FHIR provides a standard for exchanging healthcare information electronically. OMOP offers a common data model and standardised vocabularies for observational health data and reproducible analytics. These approaches perform different but complementary roles. Together, they illustrate the infrastructure required to move from fragmented source data towards evidence that can be understood and reused with confidence.
CDISC and HL7 have also developed mappings intended to streamline the movement of information from electronic health records into submission-ready clinical research datasets. This work matters because provenance and meaning must survive the journey from routine care to research, regulatory evaluation and, increasingly, AI development.
None of this is glamorous work. It is an upfront investment: the necessary homework before we seek the rewards of AI-enabled automation.
It is the difference between Real Data, Better AI as an outcome and Real Data, Better AI as a slogan.
Dr Selena Baset
Join the Road to Bilbao conversation
Selena's reflection leaves us with questions that reach well beyond technology.
What should count as "real data" when healthcare information has already passed through multiple systems and transformations? Are FAIR data and common standards sufficient, or must "better AI" also depend on representativeness, clinical relevance and demonstrable benefit? Who should ultimately decide whether an AI system is better — developers, health professionals, regulators, patients, or all of them together?
We invite members of the Health Data Forum community to share their responses, practical examples and challenges. These contributions will help inform the conversation we take to Bilbao.
