<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3673</article-id>
<article-id pub-id-type="publisher-id">11:5:3673</article-id>
<article-id pub-id-type="pii">S2399490821036739</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Privacy-Preserving Record Linkage in Population Data Systems: Leveraging Synthetic HDSS Data and Deterministic Identifier Masking for Machine Learning Models</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Bhattacharjee</surname><given-names initials="T">Tathagata</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Slaymaker</surname><given-names initials="E">Emma</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Kabudula</surname><given-names initials="C">Chodziwadziwa</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Todd</surname><given-names initials="J">Jim</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>London School of Hygiene &amp; Tropical Medicine, London, United Kingdom</institution></aff>
<aff id="affil-2"><label>2</label><institution>University of the Witwatersrand, Johannesburg, South Africa</institution></aff>
<aff id="affil-3"><label>3</label><institution>London School of Hygiene &amp; Tropical Medicine, London, United Kingdom; Catholic University Of Health And Allied Sciences, Mwanza, Tanzania, United Republic of</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3673</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3673">This article is available from the IJPDS website at: https://ijpds.org/article/view/3673</self-uri>
<abstract>
<sec>
<title>Background</title>
<p>Linking population-based datasets requires identifiable attributes that raise privacy and governance concerns. In low- and middle-income countries (LMIC), Health and Demographic Surveillance Systems (HDSS) are key sources of longitudinal population data. Synthetic data offers a privacy-preserving route for developing and evaluating record linkage (RL) methodologies without exposing sensitive microdata. This study introduces a framework that leverages Conditional Tabular Generative Adversarial Network (CTGAN)-generated synthetic HDSS data and deterministic name masking to support ethical RL and the development of supervised and ensemble machine-learning linkage models.</p>
</sec>
<sec>
<title>Methods</title>
<p>A CTGAN-based pipeline was applied to clean and harmonise Kisesa HDSS data to generate a statistically comparable synthetic dataset, while sex-specific deterministic name masking ensured full replacement of identifiers. Synthetic data quality was evaluated using distributional similarity, relationship preservation, and privacy-risk assessment. Synthetic adult clinic datasets were derived from the synthetic HDSS by introducing graded, field-specific error patterns via controlled probabilistic perturbations, creating conditions to evaluate RL models. Supervised models (logistic regression, random forest, gradient boosting, support vector machines) and an ensemble meta-classifier were trained and compared on these synthetic linkage tasks.</p>
</sec>
<sec>
<title>Results</title>
<p>The synthetic HDSS dataset preserved key demographic structures and multivariate dependencies, while deterministic masking eliminated direct identifiers. Across increasingly noisy error scenarios, supervised and ensemble models maintained strong match–nonmatch discrimination, with ensemble methods exhibiting the highest robustness as data quality deteriorated.</p>
</sec>
<sec>
<title>Conclusion</title>
<p>CTGAN-derived synthetic HDSS data combined with deterministic name masking provides a secure testbed for machine-learning-based RL, enabling ethical, reproducible, and governance-aligned population data linkage research in LMICs.</p>
</sec>
</abstract>
</article-meta>
</front>
</article>