<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v8i2.2219</article-id>
      <article-id pub-id-type="publisher-id">8:3:038</article-id>
      <title-group>
        <article-title>A generalisable linkage pipeline (GLADIS) to facilitate research for the public good</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Pratibha</surname>
            <given-names initials="V">Vellanki</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Cleaton</surname>
            <given-names initials="M">Mary</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>University of Manchester, Manchester, United Kingdom</institution></aff>
      <aff id="affil-2"><label>2</label><institution>Office for National Statistics, Newport, United Kingdom</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>14</day>
        <month>09</month>
        <year>2023</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2023</year>
      </pub-date>
      <volume>8</volume>
      <issue>3</issue>
      <elocation-id>2219</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2219">This article is available from the IJPDS website at: https://ijpds.org/article/view/2219</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>The Integrated Data Service (IDS) is a new cross-government service that facilitates research for the public good. Key to its success are Integrated Data Assets (IDAs): de-identified, grouped datasets that are joinable on an artificial ID and themed on a given topic. The Demographic Index (DI) comprises five linked administrative datasets. We are developing a generalisable method that will link administrative and survey datasets to the DI via a customisable, reproducible pipeline, to produce IDAs.</p>
    </sec>
    <sec>
      <title>Method</title>
      <p>The method focuses on the traditional methodologies of deterministic and probabilistic data linkage and uses the Splink implementation of the Fellegi-Sunter method for probabilistic matching. The pipeline will include a tool for quality-assurance (QA) via clerical review.</p>
      <p>We are researching a generalisable implementation of Splink, deriving the method’s control parameters using the results of the deterministic matching. Additionally, we are researching application of Locality Sensitive Hashing (LSH), a dimensionality-reduction method suggested to improve computational efficiency, for blocking. This is especially important due to the large size of the datasets involved.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>We plan to produce linked datasets with three quality levels – prioritising precision, balancing precision and recall and prioritising recall. As the datasets are always linked to the DI, the DI’s artificial ID can be used as a ‘spine’ to bring them together as assets (IDAs).</p>
      <p>Initially, the method will be used on the 2021 England and Wales Census. Despite not including clerical matching in the method (except for quality-assurance), we anticipate a high precision and recall due to the quality of the Census and the number of linkage variables available. Thereafter, we plan for user testing with other datasets, including the Labour Market Survey.</p>
    </sec>
    <sec>
      <title>Conclusion</title>
      <p>Our generalisable linkage pipeline for the DI will, through its IDA outputs, facilitate research for the public good. This research will directly impact government policy and responses to national health emergencies, including Covid-19, and support government priorities such as Levelling Up and the transition towards Net Zero.</p>
    </sec>
  </body>
</article>