<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2"  article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v7i3.2090</article-id>
      <article-id pub-id-type="publisher-id">7:03:314</article-id>
      <title-group>
        <article-title>Ancillary Data Record Linkage to characterize the completeness of data for the All of Us Research Program.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Yang</surname>
            <given-names initials="Y">Yuyang</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Rodriguez</surname>
            <given-names initials="K">Kelsey</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Basford</surname>
            <given-names initials="M">Melissa</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Nambiar</surname>
            <given-names initials="S">Sidd</given-names>
          </name>
          <xref ref-type="aff" rid="affil-3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Berman</surname>
            <given-names initials="L">Lew</given-names>
          </name>
          <xref ref-type="aff" rid="affil-3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Nambiar</surname>
            <given-names initials="DM">Sidd</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label>
        <institution>Northwestern University</institution>
      </aff>
      <aff id="affil-2"><label>2</label>
        <institution>Vanderbilt University</institution>
      </aff>
      <aff id="affil-3"><label>3</label>
        <institution>National Institute of Health</institution>
      </aff>
      <pub-date date-type="pub" publication-format="electronic"><day></day><month>09</month><year>2022</year></pub-date>
      <pub-date date-type="collection" publication-format="electronic"><year>2022</year></pub-date>
      <volume>7</volume>
      <issue>3</issue>
      <elocation-id>2090</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2090">This article is available from the IJPDS website at: https://ijpds.org/article/view/2090</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>The All of Us Research Program (AoURP) is an ambitious effort to gather health data from one million Americans to accelerate research. We linked Electronic Health Records (EHR) and insurance claims data to characterize the degree to which ancillary datasets can improve data completeness for care received by AoURP participants.</p>
    </sec>
    <sec>
      <title>Approach</title>
      <p>We sought to link EHR data for 400,000 consented AoURP participants with insurance claims data provided by IPM.AI (Swoop Analytics), a commercial analytics company who have insurance claims data for 300M (over 90% of) Americans.  We utilized a HIPAA-compliant privacy-preserving record linkage method (tokenization, provided by Datavant) to match patients between datasets. We evaluated match fidelity and the degree of overlap between AoURP EHRs and IPM.AI claims data. We characterized the association of patient and organizational level factors (demographics, healthcare provider organization, reporting site) with match performance.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>As of submission of this abstract, 41% of AoURP EHRs matched with IPM.AI claims. We compared patient healthcare encounters, diagnosis codes (DX), procedure codes (PX), and national drug codes (NDC) for matched patients by month. The union of AoU and IPM.AI data greatly increased data completeness in matched patients. Only 20% of healthcare encounters were seen by AoURP and IPM.AI concurrently while 25% were unique to AoU EHRs and 55% to IPM.AI claims on a monthly level. The number of diagnosis events compared between AoURP and IPM.AI is roughly equal (AoU +6%) while procedure events are elevated in claims data (23%) and drug counts are greatly elevated in AoURP EHR data (71%). We found that matched patients had more healthcare encounters compared to unmatched patients.</p>
    </sec>
    <sec>
      <title>Conclusion</title>
      <p>To our knowledge this is the first effort to address challenges in AoURP data completeness through complementary data linkage. Our results suggest that supplementary data linkage can improve data completeness in a large national research initiative. We identified several patient factors that require further investigation in improving match fidelity.</p>
    </sec>
  </body>
</article>