<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
  xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v10i3.3204</article-id>
      <article-id pub-id-type="publisher-id">10:3:170</article-id>
      <title-group>
        <article-title>Gold Standard Identifiers within Linkage</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Edwards</surname>
            <given-names initials="M">Michael</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Zamani</surname>
            <given-names initials="A">Anahita</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Thompson</surname>
            <given-names initials="S">Simon</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>SeRP, Swansea, United Kingdom</institution></aff>
      <aff id="affil-2"><label>2</label><institution>Swansea University, Swansea, United Kingdom</institution></aff>
      <pub-date>
        <day>01</day>
        <month>06</month>
        <year>2025</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2025</year>
      </pub-date>
      <volume>8</volume>
      <issue>4</issue>
      <elocation-id>3204</elocation-id>
      <permissions>
        <license license-type="open-access"
          xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International
            License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/3204">This article is available from the
        IJPDS website at: https://ijpds.org/article/view/3204</self-uri>
    </article-meta>
  </front>
  <body>
    <p>Explore utilisation of “gold standard” personal identifiers within linkage processes,
      focusing on quality of NHS numbers within a representative sample of the Welsh population. We
      provide insight into reliability and impact on metrics of accuracy for linkages produced,
      including the need to consider merging or splitting such gold standard IDs.</p>
    <p>A longitudinal, population-scale sample of Welsh individuals was internally linked using
      deterministic models, with an NHS number presumed as a “gold standard” personal identifier.
      18,906,276 records were analysed, with 5,654,212 distinct “ground truth” individuals. Initial
      linkage was performed using only the gold standard ID, whilst comparison models required exact
      matching on name, date of birth, address, and postcode to produce novel links not identified
      by the gold standard, indicating proposed ID merger based on tight overlap of record-level
      data. Performance metrics were compared against the presumed ground truth, with cluster
      quality metrics identifying instances where merging of clusters should occur.</p>
    <p>An initial deterministic model built on name, date of birth, and postcode exact match
      produces 5,651,983 clusters containing a single distinct ID, 1,113 clusters having 2 distinct
      IDs, and a single cluster containing 3 distinct IDs. Tightening rules to only consider those
      also matching on the first line of address results in 5,652,125 clusters containing a single
      distinct ID, 1,042 clusters having 2 distinct IDs, and a single cluster containing 3 distinct
      IDs. Cases of being allocated secondary IDs, where multiple gold standard IDs should be merged
      to a singular ID, can be easily rectified by the linkage process. Conversely, individuals
      being erroneously allocated onto another’s ID, causing linkage to produce strongly weighted
      bridging links, is more problematic to split even with probabilistic linkage methods.</p>
    <p>“Gold standard” IDs still fall foul of classic data quality issues and their impact should be
      carefully considered depending on the application. Lower metrics against a ground truth often
      reflect not just model performance but also underlying data collection errors or systemic
      discrepancies, suggesting areas for further refinement and research.</p>
  </body>
</article>