<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3675</article-id>
<article-id pub-id-type="publisher-id">11:5:3675</article-id>
<article-id pub-id-type="pii">S2399490821036752</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Harmonised Health Outcomes Across Administrative Data: Lessons, Opportunities, and Challenges from UK Biobank</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Domzaridou</surname><given-names initials="E">Eleni</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Lacey</surname><given-names initials="B">Ben</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Conroy</surname><given-names initials="M">Megan</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Allen</surname><given-names initials="N">Naomi</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Li Nuffield</surname><given-names initials="Y">Yangmei</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Nuffield Department of Population Health, University of Oxford, Oxford, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3675</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3675">This article is available from the IJPDS website at: https://ijpds.org/article/view/3675</self-uri>
<abstract>
<p>To develop a resource that maps health outcomes across coding schemas in linked administrative data in UK Biobank, addressing the challenge of identifying equivalent outcomes from multiple sources. UK Biobank is a prospective cohort study of ∼500,000 adults, recruited between 2006–2010, with follow up for health outcomes through linkage with administrative data. Clinical codes include Read Version 2 (Read2) and Clinical Terms Version 3 (CTV3) from primary care, and International Classification of Diseases (ICD) 9th and 10th editions (ICD-9 and ICD-10) from hospitals, cancer registries, and death records; self-reported conditions were also reported at recruitment. We reviewed existing mapping resources and mapped clinical codes in different schemas to 4-digit ICD-10 codes. We successfully mapped 81% of Read2 (N = 12,448), 93% of CTV3 (24,188), 92% of ICD-9 (3,060), and 100% of self-reported codes (509) to ICD-10 codes. Although existing resources frequently allowed one-to-one mapping of ICD-10 codes (94% of the mapped codes for Read2, 58% of CTV3, and 79% of ICD-9), the remaining codes required extensive clinical review, which is ongoing. The conversion increased the granularity of health outcomes by 3.8 times from 2,000 3-digit to 7,500 4-digit ICD-10 codes. The mapping quality will be evaluated using phecodes, and by assessing consistency across data sources. Our approach preserves clinical detail, increases coding granularity, uncovers nuanced outcomes, and enables precise, internationally comparable research using enriched UK Biobank data. Although harmonisation supports cross-cohort research, validation of outcomes across data sources remains essential to avoid misclassification and minimise bias.</p>
</abstract>
</article-meta>
</front>
</article>