News

Data rules – essential to great data outcomes 

May 26, 2025

In the database development ecosystem, it’s hard to imagine anything drier and more mysterious than data rules. Yet good data rules are an important tool in clinical research, because if they’re not correct or not used, data that is collected may not be usable, and clinical data is not something that can easily be replicated. Researchers risk wasting time, money and effort if data is collected without data rules to guide the process.

However, this is easier said than done. The development of data rules is a niche skill that requires a lot of discussion and a deep understanding of what a research program is trying to achieve. It’s also an important part of BioGrid’s software development process for data collection applications like registries.

To understand this important element of clinical database development, we need to go back to the fundamentals with BioGrid’s in-house expert on data rules.

What are data rules?

Also called logic rules, data rules are logical explanations of how data elements are recorded, and how they relate to other data elements. Data rules are captured through the manner in which the database is coded, which guides how data entry occurs.    

At the simplest level, data elements need to be collected in a consistent manner so they can be compared and analysed with confidence. How do you achieve this when the data is being collected or entered by many different people? Data rules can specify and even restrict data collection to force data to be collected in particular formats.

Most people would be familiar with data rules in applications like Microsoft Excel – you can set limited data rules for specific cells in a spreadsheet by selecting the format that data must be entered in, such as ‘date’ or ‘number’, and more granular specifications such as what the date format should look like (eg. ‘DD/MM/YYYY’ or ‘Day Month YYYY’), or how many decimal points a number can have.

Clinical data rules codify researcher knowledge

For clinical data collection, data rules can range from simple rules like the format data should be entered (number, text) and what units should be used (ml, kg, and so on), to more sophisticated relationships between data items.

A very simple data rule is setting a field as mandatory – data entry cannot continue if that field is not completed.

An example of a relationship rule is smoking history. If ‘yes’ is entered for whether the patient has smoked in the past, it would make sense to record whether the patient is a current smoker. Data rules can ensure this subsequent question is asked.

Data rules will also stipulate the relationship between ‘parent’ and ‘child’ variables, and indicate fields which must have a variable entered and not be left blank. In some cases a study will require certain fields to be filled out only for a subset of patients, or have other special conditions.

With a limited opportunity to collect clinical data, it’s imperative that all of these are captured in the data rules to ensure that data captured is suitable for the research objectives.  

Why are data rules so important?

When data rules are not developed properly, there’s a risk the dataset will contain “dirty” data that cannot be trusted, requiring data cleaning to occur prior to analysis.

For example, if adverse events are captured and indicated to be minor, but in the next data field it is recorded that the patient was returned to theatre, you can't trust either of those pieces of data without data cleaning.

Cleaning can be a very time and resource-intensive process, with data validation checks being run on the data and sent to data entry personnel. This isn’t always possible if the study has been completed and there is no one available to check the data.

BioGrid works with researchers to develop the right data rules

When developing a database and collection application for a clinical study, BioGrid’s expert team works closely with researchers during the design process to ensure that appropriate data rules are in place. This can take some time but it helps both BioGrid’s software developers and the researchers clarify what is required based on what the research outputs should be.

Researcher input is critical – while there are some obvious correct and incorrect data inputs, there are also more complex relationships that need to be considered in the data, based on what is clinically possible. Data entered might be in the correct format, but for example, is it clinically appropriate that if the patient was diagnosed with condition A that they would have received treatment B? Is it likely that the patient would be experiencing symptom C?

As well as being coded into the database and collection application, data rules are captured in a Data Definitions manual and explained for each data variable. This provides critical guidance for any person entering the data.   

The top 3 questions researchers should ask themselves when developing a clinical data collection tool


  1. When should the data field be entered?
    In other words, what are the relationships between the data fields? If the patient has had procedure X but not Y, the data collector only need to complete data fields regarding procedure X and shouldn’t be prompted for more information about procedure Y. If the patient is in a particular demographic group, the data collector may have additional specific data fields to complete about conditions of interest experienced by that group.  
  2. What are the options for data for a specific data field?
    Can a letter be included in a Medicare number? There may be a selection list of options but what if migrated data has a rogue value that doesn’t fit within accepted values?  
  3. What are the real-world scenarios of what is possible, and does the data reflect what has happened?
    A patient may have been listed as having chronic obstructive pulmonary disease but a normal FEV1 level, which clinically cannot be true. They may have indicated a level CTCAE of 1, a grading on the severity of adverse events, that indicates very mild adverse events, but data entry also indicated the patient has had a stroke. Appropriate data rules can pick up these issues as the data is entered, for immediate checking and correction.

How are data rules affected when a database needs to change because new types of data need to be added or old types no longer collected?


The addition, removal or changing of data fields should be carefully considered in relation to the whole dataset by asking some key questions.  

If you’re adding a new variable:  

  • when should this be entered?   
  • does this mean an existing variable is no longer required?   
  • Is this related to another variable, and if so should a data validation check be introduced?   
  • Is there a way to backfill old data, or do you want it backfilled by data entry personnel?   

If you’re removing a variable:  

  • Should this data still be available for analysis, or is it considered incorrect?

Discover the impactful projects BioGrid supports and stay updated with the latest news.

Connecting health information

Sign up to keep informed
Sign up to keep informed