<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3620</article-id>
<article-id pub-id-type="publisher-id">11:5:3620</article-id>
<article-id pub-id-type="pii">S239949082103620X</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Diversity of record pairs and mental models in clerical review: What reviewers use to decide matches</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Dharmawimala</surname><given-names initials="Y">Yashithi</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Christen</surname><given-names initials="P">Peter</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Ziyad</surname><given-names initials="S">Sumayya</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Lam</surname><given-names initials="J">Joseph</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Vidanage</surname><given-names initials="A">Anushka</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Schnell</surname><given-names initials="R">Rainer</given-names></name><xref ref-type="aff" rid="affil-4"><sup>4</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Australian National University, Canberra, Australia</institution></aff>
<aff id="affil-2"><label>2</label><institution>Australian National University, Canberra, Australia; University of Edinburgh, Edinburgh, United Kingdom</institution></aff>
<aff id="affil-3"><label>3</label><institution>University College London, London, United Kingdom</institution></aff>
<aff id="affil-4"><label>4</label><institution>University of Duisburg-Essen, Duisburg, Germany</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3620</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3620">This article is available from the IJPDS website at: https://ijpds.org/article/view/3620</self-uri>
<abstract>
<p>Clerical review is an integral step of many data linkage systems. It employs human expertise to assess record pairs for which a decision model, such as probabilistic record linkage, has not been able to make a match (two records refer to the same individual) or non-match (two records refer to different individuals) decision. While clerical review is used in many practical applications, there is surprising little research that systematically investigates how to best conduct this important step in the data linkage process.In this work, we explore one aspect of clerical review: How does the diversity of the record pairs selected for manual review affect the final linkage quality? To investigate this question, we generated diverse samples of difficult to classify record pairs from a large public population database, where each sample contained 100 pairs of true matches and 100 pairs of true non-matches. The samples differed in (1) the number of unique similarity patterns they had (calculated over a set of compared quasi-identifiers such as names and addresses), and (2) the actual set of quasi-identifiers available for manual review. All authors conducted a blinded manual assessment of all samples.Our results indicate that the diversity of the agreement patterns is less important for making informed decisions compared to the set of quasi-identifiers available for review. Furthermore, reviewers seem to make decisions based on mental models of how they think a match or non-match looks like. Our results will help design improved approaches for clerical review in practical linkage applications.</p>
</abstract>
</article-meta>
</front>
</article>