With synthetic data towards open and reproducible research

The search for a balance between openness and privacy

Anyone who wants to scrutinize scientific research needs access to the underlying data. But what if that data cannot be shared because of privacy concerns? TU/e is exploring whether synthetic data can offer a solution: artificial datasets that resemble real research data but do not contain sensitive information about actual individuals.

by
photo Iurii_Motov / iStock

Researchers are increasingly encouraged to share their research data. Open science is important: when other scientists have access to the data behind a study, they can verify its results, reproduce the research, and ask new questions. But what if that data cannot simply be shared?

That is the paradox many researchers face. Research funders such as NWO and the European Research Council (ERC) expect data to be made as widely available as possible to make research reproducible. 

At the same time, datasets often contain personal information or commercially sensitive data. Making such data public can therefore conflict with privacy laws, agreements with companies, or other legal and ethical requirements.

Open where possible

The principle within open science is therefore not that all research data should always be fully public. Rather, the aim is to make data as accessible as possible while taking its sensitivity into account: open where possible, closed where necessary

The TU/e Data Steward team helps researchers navigate this balance. “We work with researchers to find ways to make their data as accessible as possible while at the same time protecting sensitive information,” says Liz Guzman-Ramirez, head of the TU/e Data Steward team.

There are various ways to make sensitive data less identifiable. In interviews, for example, names and other identifying information can be removed manually. With large quantities of quantitative data, however, that process can be complex and time-consuming. Removing information can also affect the value of the dataset for further research.

Synthetic data offer another possibility.

A dataset that looks real, but isn't

Synthetic data are generated by training a model on an existing dataset. The model analyzes, among other things, the statistical relationships between the different variables and then uses that information to generate a new dataset.

The new data are not copies of the original dataset. Instead, new, fictional records are created. At the same time, the model tries to preserve the statistical properties of the original as accurately as possible.

Here is a simple example: suppose researchers have conducted a large survey in which participants answered questions about their age, education, and preferences. The synthetic dataset no longer contains any actual participants. The model does, however, attempt to preserve relationships between age, education, and the answers to the survey questions.

The result is a dataset that allows others to perform analyses without gaining access to the original personal data.

“You can still do what you want to do with the data and make the data publicly available, but you don't put the participants’ privacy at risk and you preserve their rights,” says Guzman-Ramirez.

This can also be useful when dealing with intellectual property issues or commercially sensitive information. If the original data are covered by a confidentiality agreement, for example, it may be impossible for a student to work with them. With synthetic data, that student can still practice using the dataset and perform analyses.

Not simply a matter of pressing a button

Synthetic data may sound like a simple solution, but the process requires considerable care. The model needs enough information to recognize the statistical patterns in the original dataset. This can work well, particularly with quantitative datasets containing relatively large amounts of data. The method is much less suitable for small datasets or qualitative research, such as a limited number of interviews.

The original dataset also needs to be prepared properly. Researchers may need to indicate what type of data each column contains and which variables could potentially be traced back to an individual. The system automatically recognizes certain types of personal information, such as names, addresses, and email addresses, but researchers can also identify additional variables that may be sensitive.

Reproducibility

According to Guzman, the greatest value of synthetic data is not necessarily in conducting entirely new research on the synthetic dataset. The biggest benefit lies in being able to reproduce existing analyses.

Suppose a researcher describes an interesting correlation in a scientific publication. Another researcher wants to verify whether the analysis is correct, but cannot access the original dataset because it contains personal information. Normally, that second researcher would have to collect the entire dataset again.

With a synthetic dataset, another researcher can rerun the analysis described in the publication and check whether it produces the same results.

Caveat

There is an important caveat, however. Synthetic data can help determine whether a described analysis can be reproduced, but new conclusions cannot automatically be drawn from them. Suppose a researcher discovers an interesting new correlation in the synthetic dataset. There is no guarantee that the same correlation exists in the original dataset.

Guzman therefore sees synthetic data primarily as a way to prepare analyses, test research methods, allow students to practice, and reproduce existing analyses. If a researcher makes a new scientific finding based on a synthetic dataset, that finding should subsequently be tested against real data.

Synthetic and real data are therefore not competing alternatives; they can complement each other.

Almost as good as real data

TU/e is currently exploring these possibilities in a pilot project. To generate synthetic data, the university is working with BlueGenAI, a startup that produces synthetic data based on real datasets, and anDREa, a secure cloud environment for research data.

One of the researchers participating in the pilot is Freek Relouw. He is in the final year of his PhD research at the Department of Biomedical Engineering. In collaboration with Radboudumc, he is using real medical data to develop a model that can predict whether a patient will develop kidney failure after surgery.

At one point, he was asked whether he would be interested in participating in a pilot to test the use of synthetic data. “I was happy to take part,” says Relouw.

BlueGenAI created synthetic datasets based on real patient data, which Relouw then used to train his model. “I already had a working model, so that wouldn't be the reason it failed,” says Relouw. This made it possible to test whether the model could also produce good results using synthetic data.

The result? “It turned out to work almost as well as it did with real data,” says Relouw. At the same time, synthetic data are not perfect: their quality depends on the model used to generate them. Nevertheless, Relouw sees synthetic data as “a useful new tool in the toolbox.”

Less red tape

The PhD candidate also found that working with synthetic data offered several advantages. “With real data, you work in a secure environment — some kind of cloud. But you're using a computer that you have to pay for by the minute. 

Because BlueGenAI makes the data privacy-safe, you can simply bring them to TU/e. I can download the data to my own computer and work with them anywhere for free, or share them with others.”

I want to focus on my research, not on all the legal and bureaucratic red tape around it

Freek Relouw
PhD researcher

Synthetic data also make the process much simpler. “Normally, you have to draw up a contract with a hospital specifying how the data will be protected. The data often then also need to be manually anonymized before you can share them. That takes a huge amount of time,” he says. “I want to focus on my research, not on all the legal and bureaucratic red tape around it.”

Continuing the pilot

Guzman hopes the university can continue the pilot so that it can further explore the possibilities of synthetic data. “We welcome researchers to use these services and explore their added value.” 

“My ultimate goal would be that when a researcher reads a publication and thinks, ‘I haven’t seen this correlation before; I want to check whether it holds up,’ the author can say: ‘My data are sensitive, but here is synthetic data that you can use to reproduce what I did,’” she concludes.

This article was translated using AI-assisted tools and reviewed by an editor.

Share this article