<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
  xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v8i2.2301</article-id>
      <article-id pub-id-type="publisher-id">8:3:087</article-id>
      <title-group>
        <article-title>Methods to control disclosure risk of synthetic data created by National Statistics Agencies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Raab</surname>
            <given-names initials="G">Gillian</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>University of Edinburgh, Edinburgh, United Kingdom</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>14</day>
        <month>09</month>
        <year>2023</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2023</year>
      </pub-date>
      <volume>8</volume>
      <issue>3</issue>
      <elocation-id>2301</elocation-id>
      <permissions>
        <license license-type="open-access"
          xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2301">This article is available from the IJPDS website at: https://ijpds.org/article/view/2301</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>With the recent explosion of interest in using synthetic data (SD) for disclosure control many NSAs are releasing, or considering releasing. synthetic versions of their administrative data. This presentation will review the methods that NSAs can use to limit the disclosure risk of any planned release of synthetic data.</p>
    </sec>
    <sec>
      <title>Methods</title>
      <p>This paper will review the ways in which methods of creating can be adapted to control the disclosure risk that could arise by the release of such data either to trusted researchers or to a wider group. Methods that will be evaluated will include:</p>
      <list list-type="order">
        <list-item>
          <p>The use of Statistical Disclosure Control (SDC) methods on the synthetic data before its release</p>
        </list-item>
        <list-item>
          <p>Selecting methods producing low fidelity synthetic data</p>
        </list-item>
        <list-item>
          <p>Adapting the synthesis method until it satisfies measures of disclosure risk</p>
        </list-item>
        <list-item>
          <p>Incoporating differential privacy (DP) into the method of creating synthetic data</p>
        </list-item>
      </list>
    </sec>
    <sec>
      <title>Results</title>
      <p>NSAs can use different methods to create SD based on real data (RD); see e.g. <uri>https://unece.org/info/publications/pub/373531</uri>. Tthe disclosure risk of SD depends on the context of its release, to whom, in what environment etc. Even if the planned method of release ensures low disclosure risk, NSAs will want to know what the disclosure risk might be if the SD got into the wrong hands.</p>
      <p>The SD can reveal that an identified person is in the RD (identity disclosure) or can disclose information about other measures for an individual that are part of the RD. Measures of identity disclosure and attribute disclosure are described. Results will be presented on the disclosure risk of examples of SD created for real examples by the methods 1 to 4.</p>
    </sec>
    <sec>
      <title>Conclusion</title>
      <p>Each of the methods 1 to 4 have strengths and weaknesses. Methods 2 and 4 will be ruled out for many applications because of poor fidelity to the RD. A practical way forward is suggested by combining methods 1 and 3.</p>
    </sec>
  </body>
</article>