<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
  xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v10i4.3038</article-id>
      <article-id pub-id-type="publisher-id">10:3:30</article-id>
      <title-group>
        <article-title>Consistently evaluating data linkage classification results</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Christen</surname>
            <given-names initials="P">Peter</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Ziyad</surname>
            <given-names initials="S">Sumayya</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Nanayakkara</surname>
            <given-names initials="C">Charini</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>University of Edinburgh, Edinburgh, United
        Kingdom</institution></aff>
      <aff id="affil-2"><label>2</label><institution>Australian National University, Canberra,
        Australia</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>01</day>
        <month>06</month>
        <year>2025</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2025</year>
      </pub-date>
      <volume>8</volume>
      <issue>4</issue>
      <elocation-id>3038</elocation-id>
      <permissions>
        <license license-type="open-access"
          xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International
            License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/3038">This article is available from the
        IJPDS website at: https://ijpds.org/article/view/3038</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>Data linkage is commonly viewed as the problem of classifying record pairs into matches and
        non-matches. In situations where ground truth data are available, performance measures such
        as precision, recall, F-measure, sensitivity, and specificity, are commonly used to evaluate
        the quality of matches obtained with a trained data linkage classifier.</p>
    </sec>
    <sec>
      <title>Methods</title>
      <p>Comparing multiple classifiers using such measures can, however, lead to inconsistent
        evaluation because for a given measure the same numerical result can be obtained from
        different classification outcomes. This can cause a suboptimal classifier being selected and
        potentially result in linked data sets of poor quality. To overcome this problem, we propose
        the Consistent Record Linkage (CRL) measure, an application focused evaluation method that
        ensures data linkage classifiers are assessed in a fair and transparent way. The CRL-measure
        allows the definition of maximum acceptable error rates, and it provides information about
        the robustness of a classifier based on identified classification thresholds.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>Using both synthetic and real-world data sets, we illustrate how the CRL-measure can
        provide more detailed information about the performance of data linkage classification
        results compared to traditional performance measures. Based on user selected maximum
        acceptable error rates, the CRL-measure identifies the range of classification thresholds
        where error rates are below these maximums, thereby obtaining high linkage quality. This
        indicates the robustness of a classifier with regard to a varying classification threshold.
        Furthermore, the CRL-measure shows a user if a given data linkage classifier is actually
        able to achieve a certain linkage quality or not.</p>
    </sec>
    <sec>
      <title>Conclusion</title>
      <p>The CRL-measure provides users with consistent information about how multiple data linkage
        classifiers trained on the same data set perform comparatively. This will allow a better
        selection of the most suitable classifier for a given data linkage problem and lead to
        improved quality of linked data sets.</p>
    </sec>
  </body>
</article>