<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3683</article-id>
<article-id pub-id-type="publisher-id">11:5:3683</article-id>
<article-id pub-id-type="pii">S2399490821036831</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Developing an Evaluation Toolkit for Data Linkage Pipelines</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Quinn</surname><given-names initials="L">Leah</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Seguin</surname><given-names initials="M">Manon</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Hanif</surname><given-names initials="S">Sana</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>White</surname><given-names initials="Z">Zoe</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Office for National Statistics, Manchester, United Kingdom</institution></aff>
<aff id="affil-2"><label>2</label><institution>Office for National Statistics, Titchfield, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3683</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3683">This article is available from the IJPDS website at: https://ijpds.org/article/view/3683</self-uri>
<abstract>
<p>The linkage evaluation toolkit project aims to develop tools and criteria that can be used to evaluate data linkage pipelines and methods. This toolkit will not only allow assessment of a single pipeline, but will also facilitate the comparison between pipelines; enabling analysts to identify their strengths and weaknesses. There are currently limited standardised evaluation metrics for linkage pipelines, except for the precision, recall, and bias of the linked datasets. Eight key evaluation areas were identified; accuracy, bias in linkage error, flexibility in input datasets, scalability, efficiency, platform suitability, ease of use, and the ease and transparency of quality assurance. Evaluation criteria for each area were identified and quantified. Tools are being developed to enable users to assess pipelines across these criteria. These include tools for stress-testing, testing the pipelines under different parameters, assessing bias, and capturing and quantifying qualitative feedback. We are liaising with quality teams for standardised tools to assess precision and recall. Once finalised, the toolkit will be used to compare linkage pipelines when linking the same administrative datasets, to identify which pipeline performs best across each criterion. Future use will include identifying suitable linkage methods to support the UK’s 2031 Census. The evaluation toolkit will allow analysts to evaluate the strengths and weaknesses of linkage pipelines, in both their use and impact on linked outputs. The methods can be used to improve existing pipelines, and to support the development and evaluation of new methodologies in future.</p>
</abstract>
</article-meta>
</front>
</article>