<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
  xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v10i3.3205</article-id>
      <article-id pub-id-type="publisher-id">10:3:171</article-id>
      <title-group>
        <article-title>Gold Standard Identifiers within Linkage</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Edwards</surname>
            <given-names initials="M">Michael</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Zamani</surname>
            <given-names initials="A">Anahita</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Thompson</surname>
            <given-names initials="S">Simon</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>SeRP, Swansea, United Kingdom</institution></aff>
      <aff id="affil-2"><label>2</label><institution>Swansea University, Swansea, United Kingdom</institution></aff>
      <pub-date>
        <day>01</day>
        <month>06</month>
        <year>2025</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2025</year>
      </pub-date>
      <volume>8</volume>
      <issue>4</issue>
      <elocation-id>3205</elocation-id>
      <permissions>
        <license license-type="open-access"
          xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International
            License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/3205">This article is available from the
        IJPDS website at: https://ijpds.org/article/view/3205</self-uri>
    </article-meta>
  </front>
  <body>
    <p>An abstracted multi-phase, multi-model linkage strategy is presented via an operational
      workflow applied within a real-world multi-dataset environment, with population-scale datasets
      spanning multiple domains. The approach provides a consistent, repeatable playbook for data
      linkage, leveraging the pipeline in producing quality linked data for research in a robust and
      scalable manner.</p>
    <p>At the abstract level, we define building blocks to produce linked cohorts from component
      models, merging identified sub-graphs into larger cohorts in a systematic pipeline from data
      acquisition to cohort provisioning. First, each dataset is intra-linked and quality checked,
      leveraging dataset-specific information to produce high quality within-dataset links. Next,
      inter-dataset links are produced, utilising common identifiers between sets. Both intra- and
      inter-set linking phases make use of deterministic and probabilistic linkage methods to
      produce a comprehensive set of edges. In the final phase, Intra- and inter-set edges are
      integrated into a single multi-set linked cohort before versioning and release.</p>
    <p>Producing a larger cohort from dataset-specific sub-graphs allowed the exploitation of
      focused models at each stage of the pipeline to improve quality, giving finer control and
      inspection of potential assumptions and biases during the linkage process. One example was in
      linkage of a population-scale survey, in which a logical negation excluded possible false
      links of individuals within a household, who commonly share data features. The approach also
      allows for a modular, controlled nature to revisions of linkages in improving quality, with
      modifications being confined to within specific phases of the workflow, with sub-graphs
      versioned and revised in isolation of the wider linkage. The multi-phase pipeline produces
      relatively more complex linkages on the whole, however the proposed workflow allows for a
      systematic approach which supports delivery.</p>
    <p>From an operational standpoint, this approach improves linkage turnaround by standardising
      workflows for linkage operators, abstracts to a wide range of real-world linkage applications,
      and provides granular control over a cohort’s construction at the expense of producing a more
      complex composition of links through the variety of edge types produced.</p>
  </body>
</article>