<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.1651</article-id>
      <article-id pub-id-type="publisher-id">9:5:167</article-id>
      <title-group>
        <article-title>Kids’ Environment and Health Cohort</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Maizey</surname>
            <given-names initials="L">Leah</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Platcha</surname>
            <given-names initials="J">Josie</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Gammon</surname>
            <given-names initials="T">Tim</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Wray</surname>
            <given-names initials="M">Matt</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Thompson</surname>
            <given-names initials="G">Gavin</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Antal</surname>
            <given-names initials="L">Laszlo</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Archer</surname>
            <given-names initials="R">Rosaland</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>Office for National Statistics</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>1650</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/1650">This article is available from the IJPDS website at: https://ijpds.org/article/view/1650</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>Data linkage is a vital process in the creation of many national statistics, but understanding the quality of linked data is currently highly inefficient. To find errors, data must be reviewed by humans which is costly and lengthy. Sampling is used to reduce the clerical burden. This research aims to develop a method for stratifying links to create representative samples while reducing the number reviewed. The final method will enable nuanced stratification of data for review whilst optimising resource efficiency.</p>
      <p>The objectives are to:</p>
      <list list-type="bullet">
        <list-item>
          <p>ensure that the method is adaptable across diverse datasets,</p>
        </list-item>
        <list-item>
          <p>achieve full automation,</p>
        </list-item>
        <list-item>
          <p>ensure scalability to accommodate large datasets.</p>
        </list-item>
      </list>
    </sec>
    <sec>
      <title>Approach</title>
      <p>Our approach centres on designing an algorithm that responds to the variability in the data distribution of probabilistic scores and stratify accordingly. The intention is for the developed method to automatically adjust its parameters, such as strata threshold and numbers based on the data’s characteristics.</p>
      <p>The research involves a comparative analysis of the performance of dynamic- and percentile-based stratification against the current standard practice of static threshold stratification.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>Tests are ongoing to compare the above methods on a variety of metrics including homogeneity of strata, total variance, and between-strata distance. Findings will be presented at the conference.</p>
    </sec>
    <sec>
      <title>Conclusions</title>
      <p>We hope to design a robust, generalisable and scalable stratification method that can be integrated into a Linkage pipeline.</p>
    </sec>
    <sec>
      <title>Implications</title>
      <p>Implementing the method will help to improve the quality of national statistics, ensuring more accurate, reliable and timely outputs are produced in a resource efficient manner.</p>
    </sec>
  </body>
</article>