<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.2590</article-id>
      <article-id pub-id-type="publisher-id">9:5:106</article-id>
      <title-group>
        <article-title>How ethnic name variations influence data linkage results: A population level study using public voter databases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Lam</surname>
            <given-names initials="J">Joseph</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Ziyad</surname>
            <given-names initials="S">Sumayya</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Christen</surname>
            <given-names initials="P">Peter</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Schnell</surname>
            <given-names initials="R">Rainer</given-names>
          </name>
          <xref ref-type="aff" rid="affil-3">3</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>Population, Policy &amp; Practice Research and Teaching Department, UCL Great Ormond Street Institute of Child Health</institution></aff>
      <aff id="affil-2"><label>2</label><institution>The Australian National University</institution></aff>
      <aff id="affil-3"><label>3</label><institution>University Duisburg-Essen</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>2590</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2590">This article is available from the IJPDS website at: https://ijpds.org/article/view/2590</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>It has been shown that record linkage methods can induce bias in resulting datasets since the linkage quality may differ between ethno-racial groups. In this project, we used publicly available voter databases to investigate whether ethno-racial attributes can influence the similarities calculated using approximate string comparison functions on voters' names.</p>
    </sec>
    <sec>
      <title>Approach</title>
      <p>We used 11 annual snapshots of a publicly available US voter database between 2011 and 2021 with uniquely identifiable voter data. We extracted pairs of the same voter’s first and last name values from two consecutive snapshots where these values differ. We calculated string similarities and created similarity histograms for each of the available ethno-racial categories, and other sociodemographics variables such as gender and age. We characterized common patterns of name disagreements by ethno-racial groups and compared the shapes of these histograms using cumulative density distributions.</p>
    </sec>
    <sec>
      <title>Preliminary Results</title>
      <p>Across the 10 snapshot pairs, eligible voters with non-missing first and last names are included (N2011/2012 = 6,193,001, N2020/2021 = 7,852,763), describing 124,009 first name changes and 485,807 last name changes. Results will be presented as a series of density plots. Using Jaro-Winkler similarity scores of 0.7, 0.8 and 0.9 as threshold, we examined whether match rates differ by ethno-racial categories.</p>
    </sec>
    <sec>
      <title>Conclusions and Implications</title>
      <p>Using the same string similarity function on names of individuals of different ethno-racial groups may lead to different distributions of the resulting similarity values. Understanding the patterns of ethno-racial-based name changes in your particular dataset is crucial on selecting linkage parameters that would minimise ethno-racial bias.</p>
    </sec>
  </body>
</article>