News  

Data curation – unlocking the research value in clinical data sets

January 27, 2025

Using clinical data for research comes with its own special challenges, mainly due to the multi-faceted nature of the records that contain these data.

Clinical data is collected primarily for the purposes of treating patients. However, if there is already a consideration about using these data at a later date for research, clinicians involved in establishing the data collection systems may include specific data fields for collection at the time of treatment that are not part of the treatment process, but provide important information to support future research.

As a result, an integral part of BioGrid’s work with our collaborators is data curation, a process that ensures that patient data collected during routine clinical care can be prepared for future analyses.

“The majority of the datasets available in BioGrid are ingested in their raw format/structure from the source system. This provides flexibility for analyses but requires curation to create a data set for analysis,” says Javier Haurat, Head, Data Science and Registries for BioGrid. “No data set is extracted from a source system ready to go in terms of the format/structure required to facilitate the required analyses.”

The data sets BioGrid work with are from many different data repositories, such as electronic medical record (EMR) systems, bespoke databases, open source systems or even spreadsheets. These data in these repositories are usually in their raw format, as many data custodians do not have the time or resources to perform data curation and analyses.

To make things more complicated, many clinical records include free text fields, containing information about the same data elements but in different formats.

This is fine for clinicians who review individual records to record and track patients’ treatment progress. However, a researcher who needs to analyse and compare hundreds or thousands of patient records, often completed by different health professionals, is faced with a massive task even before any analyses begins.

BioGrid’s Data Engineer Sai Whiley says “There’s a high level of complexity in designing the rules that apply to data in free text fields, especially if you've got a large volume of data. You could try to enforce the collection of data in a different way, but that is best done at the point when the original data repository is being designed. BioGrid recommends this when we’re developing a registry for a project partner.”

“There’s a high level of complexity in designing the rules that apply to data in free text fields, especially if you've got a large volume of data. You could try to enforce the collection of data in a different way, but that is best done at the point when the original data repository is being designed. BioGrid recommends this when we’re developing a registry for a project partner.”

Sai Whiley, BioGrid's Data Engineer

When BioGrid designs and creates registries such as iTestis, data fields are included that have known drop down lists to select from, so these data are curated as these data are entered.

“But if a researcher wants to analyse an existing data set, you’ve got to negotiate a trade-off between what the system requires [in terms of format], what the user clinicians want in terms of routine patient clinical care, and what is useful for analyses,” says Sai. “I think that's where BioGrid adds value through the data curation process – we enable researchers to effectively use a source system that was originally designed to be a clinical system.”

A bespoke registry system that is designed for research typically won’t allow free text fields, whereas electronic medical records, which are a wealth of patient data, have data elements that are not captured in a controlled, structured way.

For example, microbiology test results may contain ‘semi structured’ data that need to be broken down into individual components. Progress notes in a medical record, which contain clinicians’ notes, require a high level of curation to make the information accessible and analysable.

BioGrid has performed data curation for projects using hospital admitted episode data, a challenging data set to understand and work with. A recent project looking at haematological malignancies required data curation which included filtering out variables and creating a cohort to be linked to the National Death Index ready for analysis by researchers.

EINSTEIN – next level complexity in data curation

"For the EINSTEIN project, we’re venturing into the unknown as it's the first time BioGrid or the client has worked with these unstructured data. It is challenging based on the nature of the data, which is from multiple sources and contains unstructured pathology report data such as microbiology and histology,” says Javier.

The EINSTEIN project’s complexity is compounded by the fact that multiple hospitals are contributing electronic medical records created in different systems, resulting in varying data formats.

“These data need to be ingested into a usable format, so that other members of the project team can analyse and interpret them. EINSTEIN will use algorithms and natural language processing to analyse these report data or the data that's been curated by BioGrid to improve the detection of invasive fungal infections.”

To achieve this with such a complex project, the BioGrid team has regular meetings with the EINSTEIN project team to work through the research requirements and ensure that the data set is curated to achieve those objectives. Prior to working with BioGrid, the EINSTEIN team was able to access and analyse these data to create the curation requirements rules. BioGrid’s role is now to execute those rules and refine the process with the project team, as well as independently assess and advise how to remediate any issues arising from how the rules interact with the source data.

BioGrid’s data sharing platform – where the magic happens

Data curation occurs within the BioGrid platform. It’s not possible to put a timeframe on how long data curation takes, as it depends entirely on how the original data sets are collected and structured and how close these data are to being usable for the specific project requirements.

BioGrid has extensive experience with assessing data sets and working with project teams to make sure data sets align with research objectives, ensuring that the data curation process is an efficient and effective part of the research project cycle.

Discover the impactful projects BioGrid supports and stay updated with the latest news.

Connecting health information

Sign up to keep informed
Sign up to keep informed