<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3746</article-id>
<article-id pub-id-type="publisher-id">11:5:3746</article-id>
<article-id pub-id-type="pii">S2399490821037460</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Rethinking data linkage through machine learning techniques</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Tuoto</surname><given-names initials="T">Tiziana</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Tosco</surname><given-names initials="L">Laura</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Istat, Rome, Italy</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3746</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3746">This article is available from the IJPDS website at: https://ijpds.org/article/view/3746</self-uri>
<abstract>
<p>In this presentation we propose a combination of machine learning techniques to perform record linkage processes in the presence of high-volume data. Indeed, record linkage is a complex process composed of several steps, each of them exploiting different techniques. Machine learning algorithms are usually exploited to reduce the pair space search, in the well-know Fellegi and Sunter probabilistic framework. However, this framework hardly manages high volume data and extends to the case of linking simultaneously multiple files. Our proposal relies on efficient algorithms based on approximate nearest neighbour (ANN) techniques to compute the distance between records and reduce the links search space, combined with linear optimization algorithms and graph techniques for the final selection of the links. The proposed solution is implemented in open-source language. It dramatically overcomes the computational performances of other fast and scalable linkage packages. Diagnostics to evaluate the linkage results are provided, and it shows to be very effective in avoiding false links while reducing the risk of missing true matches. The proposal has been tested in several scenarios, including different sizes of the linking files, different overlaps between them, and different levels of errors in the matching variables, allowing us to guide practitioners in setting the overall linkage strategy in real life cases.</p>
</abstract>
</article-meta>
</front>
</article>