<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3485</article-id>
<article-id pub-id-type="publisher-id">11:5:3485</article-id>
<article-id pub-id-type="pii">S2399490821034856</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Feasibility research into using multiple administrative data sources and predictive modelling to produce disability status estimates for England</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Minifie</surname><given-names initials="M">Matthew</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Standeven</surname><given-names initials="C">Charlotte</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Ransley</surname><given-names initials="J">Jesse</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Wood</surname><given-names initials="S">Sarah</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>National Statistician’s Office, Newport, United Kingdom</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3485</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3485">This article is available from the IJPDS website at: https://ijpds.org/article/view/3485</self-uri>
<abstract>
<p>Official estimates of population by disability status for England and Wales are produced as part of the census once every ten years, but there is a pressing user need for more frequent, robust disability statistics. This presentation outlines ongoing novel research that responds to that challenge by examining the feasibility of producing population-level estimates for disability status in England using a predictive model applied to a suite of linked administrative data. We have linked nine administrative data sources from different government departments, including benefits, education, employment and health datasets. The de-identified data are linked via a population spine (the England-based residents on the Office for National Statistics’ 2021 Statistical Population Dataset with a valid disability status from Census 2021), with a sample size of 48.9 million. Machine learning algorithms were then used to train a model on a subsample of data. The binary outcome variable is Census 2021 disability status (non-disabled or disabled). Predictor variables include age, sex, geography and various benefit, education, employment and health variables from the administrative data. The performance of the model was tested and evaluated on a different subsample of the data to that on which the model was trained. Various evaluation metrics were employed to assess the performance of the model, along with predicted and observed disability prevalence rates by socio-demographic factors to further assess the coherence between the predictive model and Census 2021. Additionally, the predictive model was applied to earlier years for limited analysis of a short time series.</p>
</abstract>
</article-meta>
</front>
</article>