<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3495</article-id>
<article-id pub-id-type="publisher-id">11:5:3495</article-id>
<article-id pub-id-type="pii">S2399490821034959</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Feasibility research into using administrative data sources to produce multimorbidity scores for predicting general health statistics for England</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Adedire</surname><given-names initials="T">Tolu</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Munir</surname><given-names initials="W">Wajiha</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Standeven</surname><given-names initials="C">Charlotte</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Minifie</surname><given-names initials="M">Matthew</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Office of the National Statistician, Manchester, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3495</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3495">This article is available from the IJPDS website at: https://ijpds.org/article/view/3495</self-uri>
<abstract>
<p>The current primary measure of general health in England relies on census and survey data, which are limited in timeliness and granularity. Developing an administrative data-based indicator offers the potential for more frequent, detailed insights to support research and policy. This paper sets out research to develop and evaluate a multimorbidity score derived from linked administrative health datasets, combining National Health Service (NHS) hospital records, NHS General Practice data, and the Office for National Statistics’ death registrations, which are then joined to census information for residents of England. Moving beyond simple counts of conditions, or prescriptions, the multimorbidity score provides a structured measure of the burden and complexity of chronic illnesses at the individual level. We explore using this score as a key predictor of health-related outcomes, including self-reported general health status. Machine learning models were trained to predict the general health status responses from the 2021 Census. Given the subjective nature of census responses, we considered exposures on clinical indicators, healthcare utilisation metrics, demographic factors, social factors, and lifestyle factors. To ensure robustness, we experimented with multiple modelling approaches, including tree-based algorithms and regression-based methods, comparing their predictive performance across evaluation metrics. Model performance was assessed on an independent subsample using multiple evaluation metrics, and predicted probabilities were aggregated to produce breakdowns by key characteristics, which were compared against observed values. Finally, the paper considers the feasibility of generating a time series from 2015 onwards and discusses potential applications and limitations of these estimates for producing broader health-related statistics.</p>
</abstract>
</article-meta>
</front>
</article>