<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.2765</article-id>
      <article-id pub-id-type="publisher-id">9:5:274</article-id>
      <title-group>
        <article-title>Probabilistic Record Linkage for Families (PRLF): A Discussion of the Development and Validation of this Open-Source Linkage Tool</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Prindle</surname>
            <given-names initials="J">John</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Suthar</surname>
            <given-names initials="H">Himal</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Putnam-Hornstein</surname>
            <given-names initials="E">Emily</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Foust</surname>
            <given-names initials="R">Regan</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>USC Suzanne Dworak-Peck School of Social Work</institution></aff>
      <aff id="affil-2"><label>2</label><institution>UNC School of Social Work</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>2765</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2765">This article is available from the IJPDS website at: https://ijpds.org/article/view/2765</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>Synthetic data (SD) promises to unlock health data for training, research, and innovation. However, where utility evaluation is performed, it is applied ad-hoc for a single task of interest. We produce an initial design for a robust benchmark across a range of tasks.</p>
    </sec>
    <sec>
      <title>Approach</title>
      <p>We undertook several projects as a prototyping experiment to gather requirements. These projects replicate previous studies performed on the Medical Information Mart for Intensive Care — a dataset used in more than 4,000 studies. We refine definitions, identify personas, draft a user statement, and collect requirements.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>Definitions: We define utility as an extrinsic measure of SD on a larger system, most often through comparison to system performance on real data. This contrasts with fidelity, which measures the accuracy of SD through direct comparison to real data.</p>
      <p>Personas: Data custodian, User of SD, SD researcher.</p>
      <p>User statement: As a technical stakeholder, I need a reliable way to measure the utility of datasets and a benchmark to compare generation techniques.</p>
      <p>Requirements: SD researchers can focus on generation not evaluation; Supports comparison and leaderboards; Based on relevant and applications; Comprehensive across study types and applications; Future proof for population research requiring linking.</p>
    </sec>
    <sec>
      <title>Conclusion</title>
      <p>We propose the following design:</p>
      <list list-type="bullet">
        <list-item>
          <p>Data pipelines follow an extract-generate-evaluate workflow.</p>
        </list-item>
        <list-item>
          <p>Study types include cross-sectional and longitudinal.</p>
        </list-item>
        <list-item>
          <p>Applications include predictive modelling and clinical research.</p>
        </list-item>
      </list>
      <p>This results in a comprehensive utility benchmarking suite that complements current frameworks for fidelity and privacy of SD.</p>
    </sec>
  </body>
</article>