<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.2683</article-id>
      <article-id pub-id-type="publisher-id">9:5:195</article-id>
      <title-group>
        <article-title>Data as infrastructure: Systematic data curation addressing fundamental data content differences across the UK</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Orton</surname>
            <given-names initials="C">Chris</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Edwards</surname>
            <given-names initials="L">Lara</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Seymour</surname>
            <given-names initials="D">David</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Jones</surname>
            <given-names initials="M">Monica</given-names>
          </name>
          <xref ref-type="aff" rid="affil-3">3</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Quinlan</surname>
            <given-names initials="P">Philip</given-names>
          </name>
          <xref ref-type="aff" rid="affil-4">4</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Thompson</surname>
            <given-names initials="S">Simon</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Goble</surname>
            <given-names initials="C">Carole</given-names>
          </name>
          <xref ref-type="aff" rid="affil-5">5</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Quint</surname>
            <given-names initials="J">Jennifer</given-names>
          </name>
          <xref ref-type="aff" rid="affil-6">6</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Sheikh</surname>
            <given-names initials="A">Aziz</given-names>
          </name>
          <xref ref-type="aff" rid="affil-7">7</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>Swansea University</institution></aff>
      <aff id="affil-2"><label>2</label><institution>Health Data Research UK</institution></aff>
      <aff id="affil-3"><label>3</label><institution>University of Leeds</institution></aff>
      <aff id="affil-4"><label>4</label><institution>University of Nottingham</institution></aff>
      <aff id="affil-5"><label>5</label><institution>University of Manchester</institution></aff>
      <aff id="affil-6"><label>6</label><institution>Imperial College London</institution></aff>
      <aff id="affil-7"><label>7</label><institution>University of Edinburgh</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>2683</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2683">This article is available from the IJPDS website at: https://ijpds.org/article/view/2683</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objective and Approach</title>
      <p>Health Data Research UK, the UK national institute for health data science, is coordinating efforts alongside national academic partners to streamline data curation at disease, population, and data structure level to enhance data offerings and provide networked data infrastructure supporting whole-UK research.</p>
      <p>Due to clinical, coding, and system differences across the constituent countries of the UK, data is often not standardised for whole-UK analyses, creating burden on research teams and leading to long data preparation times in order to run even distributed analyses.</p>
      <p>The approach to solve this is multi-faceted, including deploying data curation and cohort creation algorithms into health data providers’ environments, and through the novel integration of federated analytics solutions (such as those piloted through recent national infrastructure programmes) improving data access and research deployment efficiency.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>Standardising data through clinical and structural data curation directly deployed to health data providers creates a framework for whole-UK studies to be readily achievable, and provide the data infrastructure base to integrate new technical federated analytics solutions to deploy and reproduce analytics without unnecessary large scale data migration.</p>
    </sec>
    <sec>
      <title>Conclusions and Implications</title>
      <p>Systematic curation of health data within national data providing organisations provides flexibility and choice for researchers in terms of the data they will apply for to answer vital research questions affecting the UK populace. Such advances will improve the quality and efficiency of research for all corners of the UK, and create a community of practice in terms of developing data from health systems to research environments.</p>
    </sec>
  </body>
</article>