<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3549</article-id>
<article-id pub-id-type="publisher-id">11:5:3549</article-id>
<article-id pub-id-type="pii">S2399490821035497</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Operationalising Graph-Based Linkage with Directory-Structured Project Frameworks</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Edwards</surname><given-names initials="M">Michael</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Zamani</surname><given-names initials="A">Anahita</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Toby</surname><given-names initials="R">Roshan</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Scully</surname><given-names initials="S">Sean</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Thompson</surname><given-names initials="S">Simon</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>SeRP, Swansea, United Kingdom; Swansea University, Swansea, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3549</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3549">This article is available from the IJPDS website at: https://ijpds.org/article/view/3549</self-uri>
<abstract>
<p>Motivations Consistently structuring complex linkage projects is essential for transparency, reproducibility, and effective management, and should ideally align with recurring linkage workflow tasks. We present a method that translates an abstract graph-based linkage workflow into a directory-centric representation, mapping subgraphs and composite graphs directly onto a file system. Record linkages are materialised as discrete subgraphs stored within dedicated directories containing predicted edges, metadata, quality metrics, and version history, which can be used in subsequent linkages. Composite graphs are generated by systematically aggregating edge lists followed by a clustering step, with the directory tree enabling selective inclusion of edge sets in resulting graphs. This structure establishes direct correspondence between conceptual graph units and storage locations, enabling transparent navigation from high-level composite cohorts to constituent subgraphs. Versioned directories preserve graph lineage, support provenance tracking and comparison of linkage iterations, and allow reproducible graph reconstruction. The separation between subgraph and composite graph layers supports modular updates, with subgraph revisions propagating to dependent composite structures without destabilising the broader system. By mapping graph concepts to concrete filesystem constructs, the approach reduces cognitive overhead for practitioners, operationalises edge provenance, and provides an auditable, reproducible foundation for large-scale linkage workflows. The approach abstracts across data domains, entity types, and linkage techniques, offering a scalable mechanism for managing graph complexity while maintaining control, transparency, and governance in real-world data linkage environments. Early application demonstrates improvements in linkage management, quality inspection, and reproducibility, illustrating the practical value of combining conceptual graph structures with disciplined project organisation.</p>
</abstract>
</article-meta>
</front>
</article>