<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3704</article-id>
<article-id pub-id-type="publisher-id">11:5:3704</article-id>
<article-id pub-id-type="pii">S2399490821037046</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Guiding principles for the provision of synthetic data</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Oliver</surname><given-names initials="E">Emily</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>ADR UK, London, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3704</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3704">This article is available from the IJPDS website at: https://ijpds.org/article/view/3704</self-uri>
<abstract>
<p>Synthetic data is an emerging tool for those learning to use and/or planning to use secure datasets, and for collaborators on secure data projects. However, there is a lack of consistency in its provision, and this affects both trust in its use and appetite to routinely produce it as a resource for researchers. There are important considerations for data owners and providers to ensure they provide synthetic data that meets the needs of users, remains compliant, and has overall positive impact. As the synthetic data field is still relatively new, norms and precedents are yet to fully emerge and develop, and standards and consistent approaches are lacking. We collaborated with experts across a range of sectors and drew on recent research with stakeholders to unpack the key facets of synthetic data provision, including worries and realities. Starting with a focus on synthetic data with low disclosure risk (typically low fidelity), we created a structured workflow that involves planning, creating, validating and releasing a synthetic dataset. We then used it to develop guiding principles for providers framed around a ‘Plan, Do, Check, Act’ model. Each stage includes specific steps to ensure clarity, consistency, quality and effective risk control. By sharing the principles and the steps to support implementation, we anticipate that this approach can be helpful in informing organisational policies on synthetic data provision. We also hope that further discussion, debate and testing will lead to increased engagement, trust and a more consistent approach to creating and sharing synthetic data.</p>
</abstract>
</article-meta>
</front>
</article>