<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3629</article-id>
<article-id pub-id-type="publisher-id">11:5:3629</article-id>
<article-id pub-id-type="pii">S2399490821036296</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Pilot Samples for Estimation of Linkage Error</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Thomson</surname><given-names initials="G">Gavin</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Wray</surname><given-names initials="M">Matt</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Office for National Statistics, Newport, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3629</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3629">This article is available from the IJPDS website at: https://ijpds.org/article/view/3629</self-uri>
<abstract>
<p>Efficient estimation of linkage error in large linked datasets requires careful allocation of limited clerical review resources. We investigate the use of small stratified “pilot” samples as an initial, low-cost strategy for informing sample size calculation and allocation across strata when estimating binomial error proportions. Stratified candidate links and unlinked pairs exhibit heterogeneous error rates and variances; however, without prior variance information, optimal allocation – such as Neyman allocation – cannot be applied. We show that pilot samples of approximately nh = 30 per stratum provide sufficiently informative variance estimates at minimal cost, enabling efficient allocation of subsequent sampling effort. Using binomial absolute margin-of-error behaviour as a rough guide we calculate uncertainty for error proportions up to a maximum of p = 0.5, and use simulations to show that beyond nh ≈ 30 the marginal reduction in uncertainty grows negligible. Simulations illustrate also that these small pilots successfully differentiate high- and low-variance strata, supporting targeted sampling in the final round. When pilots are combined with Neyman allocation, the resulting stratum-level confidence intervals were shown to become more homogeneous, reducing overall population-level variance relative to proportional allocation. Empirical comparisons therefore show that, for equal total clerical review effort, Neyman allocation informed by pilots consistently yields narrower population-level confidence intervals than proportional allocation (0.85% vs 0.93% MoE). This confirms that even very small pilot samples can materially improve efficiency by enabling variance-driven sample size calculation and allocation. The method offers a practical, scalable approach for organisations conducting clerical review of linkage errors in large datasets.</p>
</abstract>
</article-meta>
</front>
</article>