<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3757</article-id>
<article-id pub-id-type="publisher-id">11:5:3757</article-id>
<article-id pub-id-type="pii">S2399490821037575</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Enhancing geocoding in Brazil through language model-based preprocessing</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Pita</surname><given-names initials="R">Robespierre</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Rocha</surname><given-names initials="G">Gabriel</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Ichihara</surname><given-names initials="M">Maria</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Harron</surname><given-names initials="K">Katie</given-names></name><xref ref-type="aff" rid="affil-4"><sup>4</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Brito</surname><given-names initials="P">Paloma</given-names></name><xref ref-type="aff" rid="affil-5"><sup>5</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Carreiro</surname><given-names initials="R">Roberto</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Almeida</surname><given-names initials="B">Bethania</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Ramos</surname><given-names initials="P">Pablo</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Barreto</surname><given-names initials="M">Marcos</given-names></name><xref ref-type="aff" rid="affil-6"><sup>6</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Barreto</surname><given-names initials="M">Mauricio</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Federal University of Bahia, Salvador, Brazil; CIDACS, Salvador, Brazil</institution></aff>
<aff id="affil-2"><label>2</label><institution>CIDACS, Salvador, Brazil; Federal University of Bahia, Salvador, Brazil</institution></aff>
<aff id="affil-3"><label>3</label><institution>CIDACS, Salvador, Brazil</institution></aff>
<aff id="affil-4"><label>4</label><institution>University College London, London, United Kingdom</institution></aff>
<aff id="affil-5"><label>5</label><institution>Federal University of Bahia, Salvador, Brazil</institution></aff>
<aff id="affil-6"><label>6</label><institution>London School of Economics and Political Science, London, United Kingdom; CIDACS, Salvador, Brazil</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3757</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3757">This article is available from the IJPDS website at: https://ijpds.org/article/view/3757</self-uri>
<abstract>
<p>Similarity-based geocoding pipelines are highly sensitive to lexical noise in address data. This work evaluates the potential of language models to enhance and standardize address inputs during preprocessing, thereby improving geocoding accuracy in CIDACS-RL, which relies on Jaro-Winkler similarity and is sensitive to string length and prefix agreement. We extended CIDACS-RL with a preprocessing step that removes street types (e.g., street, avenue, lane) and stop words and expands numeric and abbreviated address components. Two pre-trained language models were included in the experiments. Records from nine Northeastern Brazilian states in the cohort baseline (N = 54,985,455) were geocoded by linking addresses to census tracts or coordinates from the Brazilian National Register of Addresses (CNEFE; N = 30,545,117). After linkage, we manually assessed the accuracy of our extension against previous CIDACS-RL version using stratified samples (N = 2,000) from each run to determine the cutoff points. The geocoding was evaluated using accuracy and precision. Overall, the proposed extension achieved competitive performance, with higher mean precision (0.86, SD = 0.01) and accuracy (0.94, SD = 0.024) compared to the original CIDACS-RL, which achieved a mean precision of 0.84 (SD = 0.02) and a mean accuracy of 0.93 (SD = 0.024). The largest gains were observed in states with shorter average street name lengths, such as Alagoas and Pernambuco, where precision increased by approximately 3% to 10%. Our findings support that geocoding tasks may benefit from pre-trained language models for preprocessing steps, addressing the limitation of well-posed similarity measures.</p>
</abstract>
</article-meta>
</front>
</article>