<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3481</article-id>
<article-id pub-id-type="publisher-id">11:5:3481</article-id>
<article-id pub-id-type="pii">S2399490821034819</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Application of EMR data-based LLM-based Disease Identification Framework to improve ICD data-based risk adjustment algorithms</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Lee</surname><given-names initials="S">Seungwon</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Martin</surname><given-names initials="E">Elliot</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Riazi</surname><given-names initials="K">Kiarash</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Swadas</surname><given-names initials="N">Noopur</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Eastwood</surname><given-names initials="C">Cathy</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Yao</surname><given-names initials="J">Jinli</given-names></name><xref ref-type="aff" rid="affil-4"><sup>4</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Walker</surname><given-names initials="R">Robin</given-names></name><xref ref-type="aff" rid="affil-5"><sup>5</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Southern</surname><given-names initials="D">Danielle</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Li</surname><given-names initials="B">Bing</given-names></name><xref ref-type="aff" rid="affil-6"><sup>6</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Bakal</surname><given-names initials="J">Jeffrey</given-names></name><xref ref-type="aff" rid="affil-6"><sup>6</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Provincial Research Data Services, Health Shared Services, Calgary, Canada; Centre for Health Informatics, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-2"><label>2</label><institution>Centre for Health Informatics, University of Calgary, Calgary, Canada; Department of Community Health Sciences, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-3"><label>3</label><institution>Centre for Health Informatics, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-4"><label>4</label><institution>Department of Community Health Sciences, University of Calgary, Calgary, Canada; Centre for Health Informatics, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-5"><label>5</label><institution>Department of Community Health Sciences, University of Calgary, Calgary, Canada; Applied Research &amp; Innovation,Primary Care Alberta, Calgary, Canada</institution></aff>
<aff id="affil-6"><label>6</label><institution>Provincial Research Data Services, Health Shared Services, Calgary, Canada</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3481</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3481">This article is available from the IJPDS website at: https://ijpds.org/article/view/3481</self-uri>
<abstract>
<p>Risk adjustment is essential for comparisons across populations by integrating individual health status and demographic factors to evaluate healthcare outcomes. We hypothesized that risk adjustment based on Large Language Models (LLMs) for case definitions, using electronic medical records (EMR) data, would outperform international classification of diseases (ICD) based algorithms in predicting inpatient mortality. Study Design and Methods A retrospective chart review cohort (n=10,659) consisting of randomly selected patients aged 18 years or older who were discharged from acute care settings was used. The chart review data were deterministically linked to an ICD database and an EMR database. We developed and applied an LLM-based framework (i.e., Phi-4) to EMR data for large-scale disease identification and compared it with ICD-based algorithms. To predict inpatient mortality, we calculated C-statistics using logistic regression. Among the cohort, 717 patients experienced inpatient mortality. Across all disease categories, LLM-based identification outperformed the ICD-data-based method. The ICD data-based risk adjustment algorithm achieved a C-statistic of 0.66 (95% CI 0.64 to 0.68) for in-hospital mortality, while the EMR data-based algorithm achieved a 0.76 (95% CI: 0.74 to 0.78) C-statistic. The chart review had a C-statistic of 0.71 (95% CI: 0.68 to 0.73) Effective risk adjustment for predicting health outcomes requires accurate information on patient comorbidity and demographic profiles. EMR data-based case definitions can account for disease severity (e.g., disease subtypes) and other variables (e.g., social determinants of health) that may not be readily available with historical ICD-based methods.</p>
</abstract>
</article-meta>
</front>
</article>