<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3548</article-id>
<article-id pub-id-type="publisher-id">11:5:3548</article-id>
<article-id pub-id-type="pii">S2399490821035485</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>SeRP Linkage Centre: A Scalable, Automated Architecture for Record Linkage</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Toby</surname><given-names initials="R">Roshan</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Edwards</surname><given-names initials="M">Michael</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Zamani</surname><given-names initials="A">Anahita</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Scully</surname><given-names initials="S">Sean</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Thompson</surname><given-names initials="S">Simon</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>SeRP, Swansea, United Kingdom; Swansea University, Swansea, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3548</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3548">This article is available from the IJPDS website at: https://ijpds.org/article/view/3548</self-uri>
<abstract>
<p>Modern population-scale record linkage and entity resolution requires technical infrastructures that support automation, reproducibility, and secure data integration. Here we present SeRP Linkage Centre, a modular and scalable end-to-end architecture to meet these requirements, enabling standardised linkage workflows and reporting within secure, federated environments. The workflow leverages a structured project directory, designed to standardise inputs, configurations, and outputs across linkage tasks. DBT (Data Build Tool) is used for data cleaning and transformation, providing version-controlled, auditable SQL models that prepare datasets for linkage. The linkage engine is implemented in Splink using a duckdb backend, orchestrated through Airflow DAGs that automate execution, quality checks, and dependency handling. All tasks are containerised and executed on Kubernetes, enabling horizontal scaling and reproducibility. Linkage jobs are triggered using JSON payloads via a REST API, allowing integration with user interfaces. To support transparency, linkage reports are produced automatically using ReportLab, including model descriptions, summary metrics and match quality indicators, and workflow audit trails from source data to linkage table. The architecture enables high-throughput, repeatable linkage execution with reduced manual overhead. Standardised project structures and containerised infrastructure improve reproducibility and portability across secure environments, whilst automated DBT transformations and reporting enhance traceability and quality assurance. The SeRP Linkage Centre tech stack demonstrates how a modular, automated, and containerised infrastructure can support population-level linkage at scale. Future development includes expanding the range of quality measures and insights, as well as evolving the platform towards a lakehouse architecture to unify storage and metadata governance.</p>
</abstract>
</article-meta>
</front>
</article>