<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3756</article-id>
<article-id pub-id-type="publisher-id">11:5:3756</article-id>
<article-id pub-id-type="pii">S2399490821037563</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>The influence of ethnic name characteristics on string similarities when linking large heterogeneous population databases</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Dharmawimala</surname><given-names initials="Y">Yashithi</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Christen</surname><given-names initials="P">Peter</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Ziyad</surname><given-names initials="S">Sumayya</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Vidanage</surname><given-names initials="A">Anushka</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Lam</surname><given-names initials="J">Joseph</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Schnell</surname><given-names initials="R">Rainer</given-names></name><xref ref-type="aff" rid="affil-4"><sup>4</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Australian National University, Canberra, Australia</institution></aff>
<aff id="affil-2"><label>2</label><institution>Australian National University, Canberra, Australia; University of Edinburgh, Edinburgh, United Kingdom</institution></aff>
<aff id="affil-3"><label>3</label><institution>University College London, London, United Kingdom</institution></aff>
<aff id="affil-4"><label>4</label><institution>University of Duisburg-Essen, Duisburg, Germany</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3756</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3756">This article is available from the IJPDS website at: https://ijpds.org/article/view/3756</self-uri>
<abstract>
<p>Comparing names is at the core of many data linkage applications. Names can be recorded with variations and errors (typographical, phonetic, or scanning, depending on how names have been captured). Furthermore, names can change over time. Therefore, approximate string comparison functions are commonly employed to calculate similarities between names. There are many such comparison functions, with popular ones including Jaro-Winkler, edit distance, and q-gram based techniques. Which ones are suitable for a given data linkage application is a topic that has not been explored in much detail. Selecting a suitable string comparison function is, however, crucial in the context of bias and fairness when linking population databases that can contain ethnically diverse names, as recent research has shown (for example, see Lam et al,. IPDLN 2024).In this work, we investigate the following question: Do different string comparison functions result in different levels of bias when names from different ethnic groups are being compared? We conducted extensive experiments on a large public population database where the ethnicity of each record is available. By comparing name variations of the same person using multiple string comparison functions, we show that some of these functions exhibit larger similarity ranges than others, while very different similarities are obtained for the same name pair depending upon which comparison function is employed. Our results provide important insights into how practical data linkage of large population-level databases should be conducted in order to limit linkage bias with regard to ethnic groups by employing suitable string comparison functions.</p>
</abstract>
</article-meta>
</front>
</article>