<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3676</article-id>
<article-id pub-id-type="publisher-id">11:5:3676</article-id>
<article-id pub-id-type="pii">S2399490821036764</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Using natural language processing on free-text notes found in family medicine EMRs to support healthcare system research</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Jaakkimainen</surname><given-names initials="L">Liisa</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Monchesky</surname><given-names initials="E">Elizabeth</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Virmani</surname><given-names initials="S">Shrey</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Lu</surname><given-names initials="V">Vanessa</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Moinshaghaghi</surname><given-names initials="A">Ali</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>ICES, Toronto, Canada; University of Toronto, Toronto, Canada</institution></aff>
<aff id="affil-2"><label>2</label><institution>ICES, Toronto, Canada</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3676</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3676">This article is available from the IJPDS website at: https://ijpds.org/article/view/3676</self-uri>
<abstract>
<sec>
<title>Objectives</title>
<p>Family physician (FP) electronic medical records (EMRs) have been linked to health administrative data to calculate primary care wait times for cancer patients. Since FP EMRs contain non-structured free-text notes, we assessed the feasibility of using natural language processing (NLP) on FP EMR notes improve the efficiency of identifying the index date for abnormal results.</p>
</sec>
<sec>
<title>Approach</title>
<p>We used a convenience sample of 385 FPs across Ontario, Canada who consented to have their EMR linked to health administrative data. In prior studies we had a trained abstractor review the entire EMR to identify the index date for patients having colon, endometrial, bladder or melanoma cancers which represented the first indication of an abnormal symptom, sign or radiological result found in their FPs EMR. Transformer models were trained using progress notes, consults and referrals using the labels identified by the abstractor.</p>
</sec>
<sec>
<title>Results</title>
<p>ClinicalBERT was best model for bladder (precision 0.73, recall 0.94, F1 0.82 and accuracy 0.96), colon (precision 0.72, recall 0.74, F1 0.73 and accuracy 0.94) and melanoma cancers (precision 0.87, recall 0.93, F1 0.90 and accuracy 0.98). Longformer was the best model for endometrial cancer (precision 0.68, recall 0.85, F1 0.75 and accuracy 0.95).</p>
</sec>
<sec>
<title>Conclusion</title>
<p>The NLP models which used FP EMR data had reasonable precision and recall in identifying cancer index dates. Implications: The use of NLP models on FP EMR free-text notes can improve the efficiency of using FP EMR linked to administrative data to support healthcare system tracking of primary care wait times.</p>
</sec>
</abstract>
</article-meta>
</front>
</article>