<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.2687</article-id>
      <article-id pub-id-type="publisher-id">9:5:198</article-id>
      <title-group>
        <article-title>Enhancing the usability of health data for inequalities research: a UK quality improvement project and code list curation.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Perchyk</surname>
            <given-names initials="T">Tetyana</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Lemanska</surname>
            <given-names initials="A">Agnieszka</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Whitaker</surname>
            <given-names initials="K">Katriina L.</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Kerrison</surname>
            <given-names initials="R">Robert</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>University of Surrey</institution></aff>
      <aff id="affil-2"><label>2</label><institution>Data Sciences Department, National Physical Laboratory</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>2687</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2687">This article is available from the IJPDS website at: https://ijpds.org/article/view/2687</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objective and Approach</title>
      <p>Primary care datasets provide a rich source of information for inequalities research; however, the value derived from these depends largely on the comprehensiveness of the code lists used by researchers. To date, no standardised code list for inequalities research has been developed.</p>
      <p>The aim of this project was to develop a comprehensive code list for groups of interest to inequalities research, namely: people with learning disabilities (LDs), people with severe mental illness (SMI), people from ethnic minority groups, and people who are transgender.</p>
      <p>Existing code lists were extracted from the Clinical Practice Research Datalink (CPRD) Bibliography (the largest research dataset in the UK), and four UK code list repositories: OpenCodelists, Health Data Research UK, University of Cambridge, and London School of Hygiene and Tropical Medicine. Comprehensive code lists were then curated through collation and removal of duplicates.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>16 code lists were identified for LDs, 18 for SMI, 16 for ethnicity, and 2 for transgender. From these, 661, 733, 346 and 68 unique codes were identified, respectively. Preliminary testing of the curated code lists, in a CPRD dataset, indicated that the number of individuals belonging to these groups was increased by 90% (compared to the original code list).</p>
    </sec>
    <sec>
      <title>Conclusions and Implications</title>
      <p>Our findings suggest that the curated code lists improved data capture. These lists now need to be validated by healthcare specialists, before being made publicly available. The inclusion of code-specific definitions will enable international researchers to adapt the code lists to their medical systems.</p>
    </sec>
  </body>
</article>