<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2"  article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v7i3.2067</article-id>
      <article-id pub-id-type="publisher-id">7:03:290</article-id>
      <title-group>
        <article-title>Linkage of national clinical datasets without patient identifiers using probabilistic methods.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Blake</surname>
            <given-names initials="H">Helen</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Sharples</surname>
            <given-names initials="L">Linda</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Harron</surname>
            <given-names initials="K">Katie</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>van der Meulen</surname>
            <given-names initials="J">Jan</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Walker</surname>
            <given-names initials="K">Kate</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label>
        <institution>London School of Hygiene and Tropical Medicine</institution>
      </aff>
      <aff id="affil-2"><label>2</label>
        <institution>Royal College of Surgeons of England</institution>
      </aff>
      <aff id="affil-3"><label>3</label>
        <institution>University College London (UCL) Great Ormond Street Institute of Child Health</institution>
      </aff>
      <pub-date date-type="pub" publication-format="electronic"><day></day><month>09</month><year>2022</year></pub-date>
      <pub-date date-type="collection" publication-format="electronic"><year>2022</year></pub-date>
      <volume>7</volume>
      <issue>3</issue>
      <elocation-id>2067</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2067">This article is available from the IJPDS website at: https://ijpds.org/article/view/2067</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>To develop a step-by-step process for probabilistic linkage of national clinical and administrative datasets without personal information, providing guidance on selecting variables for linkage, estimating match weights, and choosing the probabilistic linkage threshold. To validate this process against deterministic linkage using patient identifiers.</p>
    </sec>
    <sec>
      <title>Approach</title>
      <p>We undertook probabilistic linkage without personal information using electronic health records from the National Bowel Cancer Audit (NBOCA) and Hospital Episode Statistics (HES) databases for bowel cancer patients undergoing emergency surgery in England. We selected linkage variables based on completeness, and ability to discriminate between matches and non-matches, assessed using a novel score derived from m-probabilities and u-probabilities. Taking deterministic linkage using patient identifiers as the reference-standard, we calculated sensitivity and specificity of probabilistic linkage, plotted a Receiver Operating Characteristic curve across alternative thresholds of match weights, and compared patient characteristics and estimates from fitted regression models between linkage methods.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>When considering the ability to discriminate between matches and non-matches, patient and administrative variables tended to discriminate better than clinical variables. 81.4% of NBOCA records were linked to HES using probabilistic linkage, versus 82.8% using deterministic linkage. Most NBOCA records were linked to HES using both methods (8,427/10,566). Probabilistic linkage had over 96% sensitivity and 90% specificity compared to deterministic linkage using patient identifiers.  Patients that linked deterministically, but not probabilistically, were younger and more likely to have emergency admission, but otherwise had similar characteristics. Regression models for mortality and length of hospital stay according to patient and tumour characteristics were not sensitive to the linkage approach.</p>
    </sec>
    <sec>
      <title>Conclusion/Implications</title>
      <p>Probabilistic linkage without personal information can be used as an alternative to deterministic linkage using patient identifiers, or as a method for enhancing deterministic linkage. It allows analysts outside highly secure data environments to undertake linkage while minimising costs and delays, protecting data security, and maintaining linkage quality.</p>
    </sec>
  </body>
</article>