News

AI and data linking - the challenges and opportunities 

May 27, 2025

The term ‘artificial intelligence (AI)’ was first coined by John McCarthy in 1956, but it’s only in recent years that AI has exploded onto the health scene, first with radiology and pathology powered by machine learning, and more recently broad adoption of generative AI (GenAI) such as AI scribes in medical practices.  Health practitioners and researchers around the world are grappling with the implications of using AI in their practices, while patients are, for better or worse, already using public systems such as GenAI systems such as ChatGPT for unregulated “medical advice”.

It’s clear that there are both exciting opportunities and significant challenges to be considered when using AI with health data. A significant challenge lies in the fact that data are gathered from individuals and requires the highest levels of security given the potential impacts on those individuals if the data is stolen or inappropriately used.

One promising approach is to use the capabilities of the emerging AI systems to assist: AI is already revolutionising the way data are collected, connected and interpreted, with particular anticipation around bringing together large quantities of diverse data types to enable new ways of predicting, diagnosing and treating disease.

During BioGrid’s 20th Anniversary Celebration on 26 November 2024, we asked a panel of experts in the health field to share their thoughts about AI and the impact it will have on how we use health data, particularly in clinical research.

Panelists

L-R: Dr Ting Dang, Senior Lecturer at the University of Melbourne and a Visiting Fellow at the University of New South Wales

Dr Alex Duong, Deputy Chief Medical Informatics Officer, Monash Health

Mark Nevin, CEO of Cancer Council Australia

A/Prof Tam Nguyen, Deputy Director of Research at St Vincent’s Hospital Melbourne

Prof Jeannie Paterson, Professor of Consumer Protection and Technology Law, Director of the Centre for AI and Digital Ethics, The University of Melbourne

Affiliate A/Prof Paul Cooper (Convenor), Affiliate A/Prof. Unit Chair, Faculty of Health and the Deakin University School of Medicine

There was general agreement that ultimately, the use of AI has a clear objective - equity of access to healthcare and improved outcomes for patients. AI can potentially help achieve this through improving the ease and increasing the efficiency of data sharing, while maintaining good governance and security.


Most patients would probably be surprised at how siloed their data is; a clinician in one venue of care has no idea what's happened in another venue of care for the purposes of clinical care, let alone research. AI has great potential for breaking down data silos for more effective patient care, and the step beyond that is breaking down data silos for the purpose of research. This of course requires the resolution of complex issues such as data standardisation and interoperability.

Four main themes emerged from the experts’ discussion about identifying and defining how AI might change or enhance the use of health data. 

  1. Security and regulation are essential for the appropriate use of AI as a tool in healthcare and research, and to ensure trust in the tools 
  2. The importance of data quality is even more pronounced when using AI, because it has the potential to significantly magnify data quality issues 
  3. The human element is critical for ensuring AI is used appropriately, and to ensure accountability 
  4. Digital literacy for both practitioners and consumers needs a significant level up before AI use gets any more widespread

1. Security and regulation are essential 

How to regulate the use of AI is a subject that governments around the world are grappling with. Consumers are generally supportive in theory of the use of their data, and the linking and sharing of that data, but any hesitation usually stems from a lack of trust in the governing bodies.

Panelists commented on the relatively high degree of trust European people have in their governments in the areas of data protection and security, and so people are confident about sharing and using data. Whereas in Australia, consumers are more doubtful. This is despite Australia having robust consumer protection law that applies to AI use in general.

The Australian government has recently released the Voluntary AI Safety Standard (VAISS), which gives practical guidance to Australian organisations on how to safely and responsibly use and innovate with AI. The core of this standard is a set of 10 guardrails that help organisations create a foundation for safe and responsible AI use.

These types of guidelines acknowledge that there is mutuality in establishing trust – while consumers may provide consent to the use of their data with AI tools, it’s also incumbent on the organisation using the AI tools to ensure that it’s being done in a fair and reasonable way. 

Panelists also discussed the importance of defining the intended use case for AI, as this will clarify the risks involved. For example, use of AI within a clinical care context is high risk as people's lives are at stake in that environment, and there's a whole range of different guardrails that need to be considered for implementation. Use case definition also helps guide the ethical considerations - just because AI can do something, doesn't mean that it should be done.

In addition, defining use cases helps the development of appropriate practice standards to guide hospitals and providers in the deployment of AI. If suitable regulations are in place, when an AI tool comes on the market practitioners know it's at least met certain thresholds and standards. This provides a baseline for comparing the tool against the use cases so that deployment can be appropriate, and additional protections put in place if necessary.

It was noted that there are some use cases where the sensitivity and complexity of the information in some areas of healthcare and research might require a lot more work in terms of regulation, such as genetic and genomic testing. These complex use cases highlight where a regulatory framework for AI use might not work if it’s a ‘one size fits all’ model. There are already examples of tiered consent models for complex areas, where for example patients might be presented with a menu of options they can opt into for how their data would be used by AI, which also incorporates the possibility that a patient might change their mind overtime.

There is work already happening around interoperability in healthcare, so that a patient’s data follows them along the healthcare system and empowers patients to make choices about their providers and actively consent in relation to how their data is being used. This foundational work, and the protections that are put in place for it, may help with the implementation of specific regulations for AI use. In terms of individual organisations, some Australian healthcare institutions have already set up centralised data governance and AI committees with representation from clinical, legal, and consumer experts, accountable at a high level for any decision about the use of AI. The existence of layers of accountability will help to instil trust in the use of AI tools.

2. AI will magnify data quality issues 

When it came to considering AI and data quality, the term ‘garbage in, garbage out’ inevitably arose in the panel’s discussion.

As we enter this era of big data, advanced analytics and AI where people want to use more and more data, we’re requiring people in organisations to become digital workers where paper-based systems have been used for many decades. They are required to understand how data they enter is going to be used, and why that data has to be entered in a specific manner. Digital literacy across those staff has not been elevated to an appropriate level yet, but we're still requiring them to capture data that we expect to be usable.

This presents a question about the quality of the data being captured. As we import large quantities of variable quality data for AI to interpret, how much confidence can we have in the conclusions that are being formed from those data? The best way to deal with this is through standardisation of data capture, and interoperability standards.

It was noted that generative AI shows promise in helping to improve data quality through synthetic data generation. For example, if there is missing information in an electronic health record, generative AI data could be used to improve the completeness. If there are underrepresented groups in the data, generative AI could help balance the data set and also remove the ‘noise’ to improve the data quality.

However the challenge then is how to evaluate the data quality of the synthetic data or the generated data. While there is no solution to this at the moment, there is still a lot of potential for using generative AI for improving data quality, which also will lead to better AI modelling capabilities or effectiveness. 

Data quality is also important when training AI models – if the data being used isn’t clean, then the predicted outcomes generated by AI won’t make sense. In addition, AI models need to be trained on all modalities to provide the whole picture of a patient’s condition. This is where interoperability and standards across the whole health network, between silos and between practices, are essential for quality outcomes.

3. AI use requires human involvement

The discussion among panel members kept returning to the importance of human intervention in the control of AI use. They noted that while consumers might be hesitant to hand over consent for data use, there is still a remarkable level of trust in digital technology, and once consent is given it’s important not to lose vigilance.

Generative AI in particular is currently considered a ‘black box’ technology, in that we don’t fully understand how it actually generates the outputs and this creates ethical and governance challenges.

Many people talk about large language models (LLM) having a ‘reasoning capability’ but currently it is unclear whether it's reasoning or not in the human sense of reasoning, despite giving a strong appearance of doing so. This is a field of active research at present and there are less powerful “explainable AI” approaches that go some way towards mitigating the challenges of LLM approaches and their inherent lack of transparency and repeatability. There is currently a long way to go before generative AI outputs can be safely used in clinical situations in healthcare.

Consumers need to understand what they are consenting to, healthcare staff need to understand what is being collected and why, and clinicians should still question what's coming out of AI. Ultimately, the humans in the system are still accountable. 

4. Digital literacy needs to improve

Digital literacy is a key element in the success of AI, as it is a foundation for the data used in any AI tool in the healthcare system. As previously noted, data input across healthcare, particularly in clinical settings, has traditionally been paper-based, and multiple people could touch the data of a single patient, opening the possibility of errors, differences in terminology, omissions and other data problems.

Requiring the same staff to suddenly change their way of working to digital platforms, requiring specific data standards, means a period of upskilling is required, which is starting to occur. This however takes time.

In addition, the digital literacy of consumers may need to improve so they are able to understand the complexities of what is being asked of them in the consent process, and what the potential implications are. Legally, no-one ‘owns’ their own data, and it’s in everyone’s interests to understand how data can be used and how it is protected. Even more importantly, the potential good that can be done for the general health of the population through the use of large collections of data needs to be understood by everyone, and this is a long-term educational process. Once again though, it is possible to use the powers of generative AI to make such complexities more understandable so there is the prospect of using AI to assist with digital literacy.

Acknowledgement

Thank you to Dr Paul Cooper, Convenor of the AI panel, for his contribution to this article

Useful links

AI’s Ascendance in Medicine: A Timeline (Cedars Sinai), 20 April 2023. 

Artificial Intelligence and Healthcare: A Journey through History, Present Innovations, and Future Possibilities - Rahim Hirani, Kaleb Noruzi, Hassan Khuram, Anum S Hussaini, Esewi Iyobosa Aifuwa, Kencie E Ely, Joshua M Lewis, Ahmed E Gabr, Abbas Smiley, Raj K Tiwari, Mill Etienne, 26 April 2024. 

Australian Voluntary AI Safety Standard (Federal), 5 September 2024. The Voluntary AI Safety Standard (VAISS) helps organisations develop and deploy AI systems in Australia safely and reliably. 

Digital Empires: The Global Battle to Regulate Technology, by Anu Bradford, 2023, Oxford University Press. 

European Union AI Act: The European Union’s priority was to make sure that AI systems used in the EU are safe, transparent, traceable, non-discriminatory and environmentally friendly. AI systems should be overseen by people, rather than by automation, to prevent harmful outcomes. They also wanted to establish a technology-neutral, uniform definition for AI that could be applied to future AI systems. 

FHIR: Fast Healthcare Interoperability Resources (FHIR) is a standard for health care data exchange, published by HL7. 

HL7 International: Founded in 1987, Health Level Seven International (HL7) is a not-for-profit, ANSI-accredited standards developing organisation dedicated to providing a comprehensive framework and related standards for the exchange, integration, sharing and retrieval of electronic health information that supports clinical practice and the management, delivery and evaluation of health services. 

Discover the impactful projects BioGrid supports and stay updated with the latest news.

Connecting health information

Sign up to keep informed
Sign up to keep informed