<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3725</article-id>
<article-id pub-id-type="publisher-id">11:5:3725</article-id>
<article-id pub-id-type="pii">S2399490821037253</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Assessing data linkage solutions in high-volume settings</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Tosco</surname><given-names initials="L">Laura</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Tuoto</surname><given-names initials="T">Tiziana</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Istat, Rome, Italy</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3725</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3725">This article is available from the IJPDS website at: https://ijpds.org/article/view/3725</self-uri>
<abstract>
<p>To design a record linkage strategy, different factors need to be taken into account, e.g. the size of the linking files, the level of errors in the matching variables, the unknown match rate. All these aspects impact on several choices: either a deterministic or probabilistic approach, (e.g. the traditional Fellegi and Sunter, a Bayesian framework, some machine learning algorithms), solutions for reducing and filtering the search space of the links, the choice of the matching variables and the metrics for comparison, the setting of resources for manual revision and checks of the results and ambiguous cases. When dealing with large volume data, as is often the case in national agencies, everything becomes even more complicated since some solutions are not scalable. In this presentation, we explore some solutions based on machine learning techniques that are very convenient for large amount of data, due to their scalability and computational efficiency. We appraise the machine learning solutions, comparing them to the traditional probabilistic solution and underlining benefits and limits. The comparison includes elements like the file size, the level of errors in the matching variables, the matching rate, along with the traditional diagnostics for the linkage results, like the false match and missing match rates, time of execution and computation resource required.</p>
</abstract>
</article-meta>
</front>
</article>