Join us in this Open Science Lunch to hear about synthetic data - what it is, how the use of artificial intelligence affects the generation of synthetic data and what are the implications for open science.
Abstract
Access to research data is essential for transparency, reproducibility, and scientific collaboration, yet sharing data can be difficult when datasets contain sensitive or personal information. Synthetic data have increasingly been proposed as a way to address this tension: instead of simply releasing the original records or anonymization, researchers can use statistical and AI-based methods to generate artificial data that reproduce useful characteristics of the original dataset.
In this talk, Qinyi Liu will introduce what synthetic data are, how recent advances in artificial intelligence have changed the way they are generated, and where they may support more open and reusable research data. Using concrete research examples, she will illustrate how synthetic data can be incorporated into research workflows, for instance to enable data exploration, method development, collaboration, teaching, and selected forms of data sharing without providing direct access to the original sensitive records. Qinyi will also discuss how researchers can assess whether synthetic data are suitable for these purposes by evaluating their statistical fidelity, analytical utility, privacy risk, and fairness. At the same time, synthetic data have important limitations: they are not automatically anonymous, may reproduce biases or sensitive patterns present in the source data, and may not preserve all of the information required for valid scientific conclusions. The talk will therefore consider both the opportunities synthetic data offer for privacy-conscious open science and the practical questions researchers should ask before generating, sharing, or reusing them.
Information about the speaker
Qinyi Liu is an Assistant Professor at the City University of Macau. Her research focuses on trustworthy AI and educational data, with particular interests in synthetic data, privacy-preserving learning analytics, and responsible AI for education. She completed her PhD at the Centre for the Science of Learning & Technology (SLATE), University of Bergen, where her doctoral research focused on synthetic data and privacy-enhancing technologies for learning analytics. Her work on synthetic educational data has examined its potential for data sharing as well as the associated challenges of privacy, utility, and fairness.