<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v8i1.2153</article-id>
<article-id pub-id-type="publisher-id">8:1:29</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>De-identification of free text data containing personal health information: a scoping review of reviews</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Negash</surname><given-names initials="B">Bekelu</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="corresp" rid="correspondingAurthor">*</xref></contrib>
<contrib contrib-type="author"><name><surname>Katz</surname><given-names initials="A">Alan</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-2">2</xref></contrib>
<contrib contrib-type="author"><name><surname>Neilson</surname><given-names initials="CJ">Christine J.</given-names></name><xref ref-type="aff" rid="affil-3">3</xref></contrib>
<contrib contrib-type="author"><name><surname>Moni</surname><given-names initials="M">Moniruzzaman</given-names></name><xref ref-type="aff" rid="affil-4">4</xref></contrib>
<contrib contrib-type="author"><name><surname>Nesca</surname><given-names initials="M">Marcello</given-names></name><xref ref-type="aff" rid="affil-1">1</xref></contrib>
<contrib contrib-type="author"><name><surname>Singer</surname><given-names initials="A">Alexander</given-names></name><xref ref-type="aff" rid="affil-2">2</xref></contrib>
<contrib contrib-type="author"><name><surname>Enns</surname><given-names initials="JE">Jennifer E.</given-names></name><xref ref-type="aff" rid="affil-1">1</xref></contrib>
<aff id="affil-1"><label>1</label><institution>Manitoba Centre for Health Policy, Department of Community Health Sciences, Rady Faculty of Health Sciences, University of Manitoba</institution></aff>
<aff id="affil-2"><label>2</label><institution>Department of Family Medicine, Rady Faculty of Health Sciences, University of Manitoba</institution></aff>
<aff id="affil-3"><label>3</label><institution>Neil John Maclean Health Sciences Library, University of Manitoba</institution></aff>
<aff id="affil-4"><label>4</label><institution>George &#x0026; Fay Yee Centre for Healthcare Innovation, Department of Community Health Sciences, Rady Faculty of Health Sciences, University of Manitoba</institution></aff>
</contrib-group>
<author-notes>
<corresp id="correspondingAurthor"><label>*</label>Corresponding author: Bekelu Negash <email>Bekelu.negash@umanitoba.ca</email>
</corresp>
<fn fn-type="conflict">
<label>Conflict of interest</label>
<p>The author(s) declared no potential conflicts of interest with respect to the research, and/or publication of this article.</p>
</fn>
</author-notes>
<pub-date date-type="pub" publication-format="electronic"><day>12</day><month>12</month><year>2023</year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<volume>8</volume>
<issue>1</issue>
<elocation-id>2153</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/2153">This article is available from the IJPDS website at: https://ijpds.org/article/view/2153</self-uri>
<abstract>
<title>Abstract</title>
<sec>
<title>Introduction</title>
<p>Using data in research often requires that the data first be de-identified, particularly in the case of health data, which often include Personal Identifiable Information (PII) and/or Personal Health Identifying Information (PHII). There are established procedures for de-identifying structured data, but de-identifying clinical notes, electronic health records, and other records that include free text data is more complex. Several different ways to achieve this are documented in the literature. This scoping review identifies categories of de-identification methods that can be used for free text data.</p>
</sec>
<sec>
<title>Methods</title>
<p>We adopted an established scoping review methodology to examine review articles published up to May 9, 2022, in Ovid MEDLINE; Ovid Embase; Scopus; the ACM Digital Library; IEEE Explore; and Compendex. Our research question was: What methods are used to de-identify free text data? Two independent reviewers conducted title and abstract screening and full-text article screening using the online review management tool Covidence.</p>
</sec>
<sec>
<title>Results</title>
<p>The initial literature search retrieved 3,312 articles, most of which focused primarily on structured data. Eighteen publications describing methods of de-identification of free text data met the inclusion criteria for our review. The majority of the included articles focused on removing categories of personal health information identified by the Health Insurance Portability and Accountability Act (HIPAA). The de-identification methods they described combined rule-based methods or machine learning with other strategies such as deep learning.</p>
</sec>
<sec>
<title>Conclusion</title>
<p>Our review identifies and categorises de-identification methods for free text data as rule-based methods, machine learning, deep learning and a combination of these and other approaches. Most of the articles we found in our search refer to de-identification methods that target some or all categories of PHII. Our review also highlights how de-identification systems for free text data have evolved over time and points to hybrid approaches as the most promising approach for the future.</p>
</sec>
</abstract>
<kwd-group>
<kwd>de-identification</kwd>
<kwd>Health Insurance Portability and Accountability Act</kwd>
<kwd>electronic medical records</kwd>
<kwd>machine learning</kwd>
<kwd>personal health information</kwd>
</kwd-group>
<funding-group>
<funding-statement>This research was supported by a foundation grant from the Canadian Institutes of Health Research (Foundation grant reference number 148427).</funding-statement>
</funding-group>
</article-meta>
</front>
<body>
<sec>
<title>Introduction</title>
<p>The production, collection and use of population data for research is becoming more prevalent across multiple sectors, but particularly in health and healthcare [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. For example, the use of electronic health records has seen a significant increase among researchers and clinicians [<xref ref-type="bibr" rid="ref-4">4</xref>, <xref ref-type="bibr" rid="ref-5">5</xref>]. However, population datasets often contain Personal Identifiable Information (PII) and/or Personal Health Identifying Information (PHII), which researchers have the responsibility to keep confidential. In Canada, the use of population data containing PII and PHII in research is governed by the Canadian Tri-Council Policy Statement on Ethical Conduct for Research Involving Humans, which includes three core principles: respect for persons, concern for welfare, and justice [<xref ref-type="bibr" rid="ref-6">6</xref>]. One effective way to preserve privacy and abide by this ethical framework is to de-identify the data before it is used in research. <italic>De-identification</italic> refers to the removal or masking of PII/PHII in a dataset and for research purposes, it may be preferable to anonymisation, a process that eliminates all identifying details in a data record with no way of back-tracking to link related data records together [<xref ref-type="bibr" rid="ref-7">7</xref>]. When a record number, file number or other encrypted linkage tool is retained in the original data, the data are not referred to as &#x2018;anonymised&#x2019; but are instead &#x2018;de-identified&#x2019; and can be used in data linkage applications.</p>
<p>The federally mandated Freedom of Information and Protection of Privacy Act (FIPPA) and the provincially mandated Personal Health Information Act (PHIA) provide definitions of PII and PHII and set out guidelines to inform the process of de-identifying structured data [<xref ref-type="bibr" rid="ref-7">7</xref>]. S<italic>tructured</italic> <italic>data</italic> are organised into specific value sets and are typically stored in a database [<xref ref-type="bibr" rid="ref-8">8</xref>]. Meanwhile, <italic>unstructured</italic> or <italic>free text</italic> <italic>data</italic> do not have pre-defined values; for example, reports created by physicians may contain free text data that vary widely in structure and content [<xref ref-type="bibr" rid="ref-8">8</xref>]. Currently, there is very little formal guidance available on how to de-identify free text data, and none that we could find that differentiates PII and PHII in the de-identification process. In this matter, the distinction between PII (recorded information that could identify an individual or groups of individuals) and PHII (specific health information about an individual or groups of individuals) is important, because specific approaches for de-identification are needed if health information is present in the data [<xref ref-type="bibr" rid="ref-8">8</xref>] (see <xref ref-type="table" rid="table-1">Table 1</xref> for examples [<xref ref-type="bibr" rid="ref-9">9</xref>]).</p>
<table-wrap id="table-1">
<label>Table 1: Examples of personal identifiable information and personal health identifying information</label>
<table frame="hsides" rules="groups">
<col width="100%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt;" valign="middle"><bold>Personal Identifiable Information (PII):</bold>
<list list-type="bullet">
<list-item><p>Name, contact information</p></list-item>
<list-item><p>Age, sex, sexual orientation, martial or family status</p></list-item>
<list-item><p>Ancestry, race, colour, nationality, national or ethnic origin</p></list-item>
<list-item><p>Religion, creed, religious belief, association, or activity</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle"><bold>Personal Health Identifying Information (PHII):</bold>
<list list-type="bullet">
<list-item><p>An individual&#x2019;s health or health care history, including genetic information about the individual</p></list-item>
<list-item><p>The provision of health care to the individual</p></list-item>
<list-item><p>Payment for health care provided to the individual, including personal health information number (PHIN) and any other identifying number, symbol or particular assigned to an individual</p></list-item>
<list-item><p>Any identifying information about the individual collected in the course of, and incidental to, the provision of health care or payment for health care</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The natural language processing (NLP) research community has made great strides in developing methods for automatically de-identifying data. There are currently two primary approaches in use:</p>
<list list-type="order">
<list-item><p><bold>Rule-based methods</bold> use pattern-matching with set conditions to satisfy a rule [<xref ref-type="bibr" rid="ref-7">7</xref>]. A rule-based method could, for example, be used to find names, addresses, or email addresses in data records. Advantages to rule-based methods include that they are relatively simple to create and do not require labelled data [<xref ref-type="bibr" rid="ref-10">10</xref>], and certain sub-types (e.g., generalisation, suppression, and data perturbation) can also be used to prevent individual records from being traced back and re-identified, if this is an important aspect of the research study [<xref ref-type="bibr" rid="ref-10">10</xref>]. However, developing rules can be time-consuming since it is difficult to include all possible examples in the rules, and the experts who design the rules may make assumptions about the data that could limit the effectiveness of the de-identification process [<xref ref-type="bibr" rid="ref-11">11</xref>].</p></list-item>
<list-item><p><bold>Machine learning (ML)/statistical learning methods</bold> use probabilistic or classification modelling to describe the structure of data or generate predictions based on inputs from a dataset. Machine/statistical learning algorithms are classified as either supervised or unsupervised. Supervised ML requires that a sample of data be labelled manually to support the model. The advantage of supervised ML approaches is that they automatically learn sophisticated pattern recognition [<xref ref-type="bibr" rid="ref-12">12</xref>]. However, it can be more difficult to identify sources of error in unsupervised deep learning ML model than in a rule-based approach. In addition, when it comes to rare types of information, ML methods can also have a lower performance compared to rule-based methods [<xref ref-type="bibr" rid="ref-11">11</xref>].</p></list-item>
</list>
<p>These two methods can be used for de-identification of both structured and free text data. De-identifying structured data is a relatively straight-forward process compared to de-identifying free text data, because structured data typically have a limited number of clearly identified fields; in addition, there is some literature to inform and guide the process [<xref ref-type="bibr" rid="ref-7">7</xref>]. De-identifying free text data, however, may necessitate a more sophisticated approach, since identifying information may occur anywhere in the free text and may include either PII or PHII or both. Some researchers have developed hybrid approaches in an attempt to combine the advantages of rule-based and ML methods for de-identification of PHII [<xref ref-type="bibr" rid="ref-11">11</xref>]. The hybrid approaches take advantage of the fact that certain types of PHII exhibit predictable lexical patterns and thus lend themselves well to de-identification via rule-based methods, whereas other frequently encountered PHII types, particularly those with unpredictable lexical variations, are more amenable to machine learning approaches [<xref ref-type="bibr" rid="ref-13">13</xref>]. More recently, the use of deep learning methods has been explored to de-identify electronic health records [<xref ref-type="bibr" rid="ref-14">14</xref>]. These methods have the ability to learn the most relevant features from the raw data, minimising the need for human input and making the pre-processing and feature engineering steps less time consuming [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>Data de-identification techniques are advancing quickly and have a growing number of applications in research settings. In this scoping review, we provide an overview of what is known about NLP methods used to de-identify free text data.</p>
</sec>
<sec>
<title>Methods</title>
<p>We based our scoping review approach on Arksey and O&#x2019;Malley&#x2019;s (2005) well-established scoping review framework, which comprises five stages [<xref ref-type="bibr" rid="ref-16">16</xref>]: identifying the research question; identifying relevant studies; study selection; charting the data; and collating, summarising, and reporting the results. Our research question was: What methods are used to de-identify free text data?</p>
<sec>
<title>Search strategy</title>
<p>A professional librarian and research coordinator developed the search strategy. Initially, one search strategy was tailored to the health database Ovid MEDLINE, and another was tailored to the ACM Digital Library, a computing literature database. Both strategies were independently peer-reviewed according to the Peer Review of Electronic Search Strategies (PRESS) checklist by a second librarian with the required subject specialisation [<xref ref-type="bibr" rid="ref-17">17</xref>]. The final search strategies were translated for use in Ovid Embase; Scopus; IEEE Explore; and Compendex. All searches were conducted on May 9, 2022. No date limits were used. The MEDLINE and Embase searches were limited to English language publications. Complete search histories for each database are available online (http://hdl.handle.net/1993/37168http://hdl.handle.net/1993/37168).</p>
</sec>
<sec>
<title>Article screening and selection</title>
<p>Using the study selection criteria in <xref ref-type="table" rid="table-2">Table 2</xref>, two independent reviewers examined the titles and abstracts of the search results. Articles that were ambiguous were discussed with the research coordinator and a consensus decision was reached on whether or not to include them in full-text article screening. The two reviewers then completed full-text article screening on the selected articles.</p>
<table-wrap id="table-2">
<label>Table 2: Study selection criteria</label>
<table frame="hsides" rules="groups">
<col width="50%"/>
<col width="50%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Inclusion criteria</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Exclusion criteria </bold></td>
</tr>
<tr>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Studies published in English</p></list-item>
<list-item><p>Discusses methods of de-identification</p></list-item>
<list-item><p>Focused on free text data</p></list-item>
<list-item><p>Review article</p></list-item>
</list></td>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>Focused on the accuracy and representability of the text after de-identification</p></list-item>
<list-item><p>Focused on privacy and less on the method of de-identification</p></list-item>
<list-item><p>Used de-identified text for data</p></list-item>
<list-item><p>Focused on cryptography de-identification methods</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Data extraction and analysis</title>
<p>The data categories we extracted are presented in <xref ref-type="table" rid="table-3">Table 3</xref>. We analysed and summarised the results in accordance with the PRISMA-ScR reporting checklist [<xref ref-type="bibr" rid="ref-18">18</xref>]. The data analysis was designed to provide an overview of methods used for de-identifying free text data.</p>
<table-wrap id="table-3">
<label>Table 3: Data categories extracted</label>
<table frame="hsides" rules="groups">
<col width="100%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt;" valign="middle">Article Information
<list list-type="bullet">
<list-item><p>Journal discipline (medicine, computer science, both, other)</p></list-item>
<list-item><p>Type of review</p></list-item>
<list-item><p>Publication year range of articles included in the review</p></list-item>
<list-item><p>Inclusion/exclusion criteria that the review article used</p></list-item>
<list-item><p>Number of articles cited in the review article</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="middle">Any mention of legal framework or guidelines</td></tr>
<tr>
<td align="left" valign="middle">Type of PII/PHII addressed</td></tr>
<tr>
<td align="left" valign="middle">Type of text data (medical [e.g., EHR, safety reports], social media, other)</td></tr>
<tr>
<td align="left" valign="middle">De-identification methods</td></tr>
<tr>
<td align="left" valign="middle">Evaluation metrics for the de-identification outcome</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec>
<title>Results</title>
<sec>
<title>Article screening and selection</title>
<p>As shown in <xref ref-type="fig" rid="fig-1">Figure 1</xref>, we identified 3,312 articles in the initial search, and removed 329 duplicates. The two reviewers had a 95.5% agreement rate during title and abstract screening; 4.2% (124) of the initial search results were included at this stage. After full-text article screening, 14.5% (18) of the 124 articles met the specified criteria and were included in the scoping review.</p>
<fig id="fig-1"><label>Figure 1: PRISMA diagram &#x2013; article search and selection process</label>
<graphic xlink:href="ijpds-06-2153-g001.tif"/>
</fig>
</sec>
<sec>
<title>Article characteristics</title>
<p>Of the 18 included articles, twelve were from the computer science literature. Most (83%) were literature reviews. The other information we planned to extract was scantily available &#x2013; only three of the 18 articles mentioned the databases and registers the authors used for their searches, two articles provided information regarding the year of publication of the primary articles, number of articles or the percentage of the articles included in their review, and another three articles indicated what inclusion/exclusion criteria the authors used. <xref ref-type="table" rid="table-4">Table 4</xref> presents more details on these latter articles.</p>
<table-wrap id="table-4">
<label>Table 4: Characteristics of the articles that reported inclusion and exclusion criteria</label>
<table frame="hsides" rules="groups">
<col width="10%"/>
<col width="20%"/>
<col width="20%"/>
<col width="10%"/>
<col width="20%"/>
<col width="10%"/>
<col width="10%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Author</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Databases/registries searched</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Type of review</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Year range</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Inclusion criteria</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Exclusion criteria</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Number of articles included</bold></td>
</tr>
<tr>
<td align="left" valign="top">Meystre et al., 2010 [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td align="left" valign="top">PubMed, conference proceedings, ACM Digital Library</td>
<td align="left" valign="top">Literature review</td>
<td align="left" valign="top">1995&#x2013;2010</td>
<td align="left" valign="top">Key terms: de-identification, anonymisation, text scrubbing, narrative text, and/or automated text de-identification. For the ACM Digital Library, the same terms were used, with the addition of medical, medicine, biomedical or clinical.</td>
<td align="left" valign="top">Focused on structured data, radiological or face image de-identification, manual de-identification</td>
<td align="left" valign="top">18</td></tr>
<tr>
<td align="left" valign="top">Shickel et al., 2017 [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td align="left" valign="top">Google Scholar</td>
<td align="left" valign="top">Literature review</td>
<td align="left" valign="top">Up to August 2017</td>
<td align="left" valign="top">Key terms: Electronic health records (EHRs) or electronic medical records (EMR) in conjunction with deep learning or a specific deep learning method (e.g., recurrent neural network [RNN]).</td>
<td align="left" valign="top">Not described.</td>
<td align="left" valign="top">44 articles on privacy-preserving methods, including cryptography-based methods, and approximately 4 articles on anonymisation methods of de-identification (exact number not specified in article).</td></tr>
<tr>
<td align="left" valign="top">Kushida et al., 2012 [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td align="left" valign="top">BIOSIS Previews, CINAHL, Inspec, MEDLINE, SciVerse, Scopus, Web of Science</td>
<td align="left" valign="top">Systematic review</td>
<td align="left" valign="top">Up to June 30, 2011</td>
<td align="left" valign="top">Key terms: De-identify, de-identification, anonymise, anonymisation, data scrubbing, and text scrubbing. Reviewed additional articles extracted from references of articles from search.</td>
<td align="left" valign="top">Citations for non-relevant article types (e.g., reviews, opinions, editorials, or commentaries), outside medical records domain, de-identification, or anonymisation strategy lacked sufficient detail to understand or interpret it.</td>
<td align="left" valign="top">45</td></tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Legal framework or guidelines mentioned</title>
<p>Of the 18 included articles, 12 mentioned legal acts governing data privacy. The Health Insurance Portability and Accountability Act (HIPAA) was mentioned in nine articles (75%), and four (33%) of the articles considered the European Union&#x2019;s General Data Protection Regulation (GDPR). The remaining legal acts mentioned were from Canada, China, Australia and New Zealand &#x2013; see <xref ref-type="table" rid="table-5">Table 5</xref> for more details.</p>
<table-wrap id="table-5">
<label>Table 5: Legal frameworks or guidelines referred to in the articles</label>
<table frame="hsides" rules="groups">
<col width="75%"/>
<col width="25%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Legal framework or guidelines mentioned</bold></td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Article(s)</bold></td>
</tr>
<tr>
<td align="left" valign="middle">Health Insurance Portability and Accountability Act (HIPAA)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>, <xref ref-type="bibr" rid="ref-25">25</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]</td></tr>
<tr>
<td align="left" valign="middle">General Data Protection Regulation (GDPR)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-22">22</xref>, <xref ref-type="bibr" rid="ref-23">23</xref>, <xref ref-type="bibr" rid="ref-26">26</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Health Information Technology for Economic and Clinical Health (HITECH)Act</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-21">21</xref>, <xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Personal Information Protection and Electronic Documents Act (PIPEDA)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Consumer Data Right</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="middle">China Civil Code</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Medical Practitioners Act</td>
<td align="center" valign="top"></td></tr>
<tr>
<td align="left" valign="middle">Personal Information Protection Law</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Regulations on Medical Records Management In Medical Institutions</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Children&#x2019;s Online Privacy Protection Act</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-12">12</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Genetic Information Non-discrimination Act</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-21">21</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Gramm-Leach-Bliley Act</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-12">12</xref>]</td></tr>
<tr>
<td align="left" valign="middle">Health Information Privacy Code</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-26">26</xref>]</td></tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Types of PII and PHII</title>
<p>Seventeen articles mentioned different types of PII and PHII (<xref ref-type="table" rid="table-6">Table 6</xref>).Eight of these articles identified methods that de-identified protected health information according to all 18 categories of HIPAA [<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-9">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-21">21</xref>, <xref ref-type="bibr" rid="ref-23">23</xref>, <xref ref-type="bibr" rid="ref-24">24</xref>, <xref ref-type="bibr" rid="ref-26">26</xref>, <xref ref-type="bibr" rid="ref-29">29</xref>] while others identified some of the HIPAA categories (<xref ref-type="table" rid="table-7">Table 7</xref>).</p>
<table-wrap id="table-6">
<label>Table 6: Types of PII/PHII referred to in the articles</label>
<table frame="hsides" rules="groups">
<col width="75%"/>
<col width="25%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Types of PII/PHII</bold></td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Article(s)</bold></td></tr>
<tr>
<td align="left" valign="top">Individuals&#x2019; identifiers (such as credit card records) and interaction privacy (e.g., use of voice/fingerprint)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-30">30</xref>]</td></tr>
<tr>
<td align="left" valign="top">Key attributes (e.g., ID, name, social security), quasi-identifiers (e.g., birth date, zip code, position, job, blood type), sensitive attributes (e.g., salary, medical examinations, credit card releases)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-31">31</xref>&#x2013;<xref ref-type="bibr" rid="ref-33">33</xref>]</td></tr>
<tr>
<td align="left" valign="top">7 types of PHII, including personal names, ages, geographical locations, hospitals and healthcare organisations, dates, contact information, IDs</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-19">19</xref>]</td></tr>
<tr>
<td align="left" valign="top">PHI: patient name, phone number, physician name, medical history. PII: names, addresses, contact numbers</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-22">22</xref>]</td></tr>
<tr>
<td align="left" valign="top">18 categories of PHI according to HIPAA, quasi-identifiers, 9 categories of personal information according to China Civil Code (name, birthday, ID number, biometric information, home address, phone number, email address, health condition information, and personal tracking information)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-23">23</xref>]</td></tr>
<tr>
<td align="left" valign="top">PHI according to HIPAA, doctor&#x2019;s name and years extracted from dates</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-20">20</xref>]</td></tr>
<tr>
<td align="left" valign="top">Direct identifiers (e.g., name, mailing address, email, social security number, phone number or driver&#x2019;s license number) and indirect identifiers (e.g., birth date, postal code, and sex)</td>
<td align="center" valign="top">[<xref ref-type="bibr" rid="ref-27">27</xref>, <xref ref-type="bibr" rid="ref-28">28</xref>]</td></tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-7">
<label>Table 7: HIPAA categories</label>
<table frame="hsides" rules="groups">
<col width="100%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle">
<list list-type="bullet">
<list-item><p>Names</p></list-item>
<list-item><p>All geographical subdivisions smaller than a state except the first two digits of the zip code</p></list-item>
<list-item><p>All elements of dates (except year)</p></list-item>
<list-item><p>Telephone numbers</p></list-item>
<list-item><p>Fax numbers</p></list-item>
<list-item><p>Electronic mail addresses</p></list-item>
<list-item><p>Social security numbers</p></list-item>
<list-item><p>Medical record numbers</p></list-item>
<list-item><p>Health plan numbers</p></list-item>
<list-item><p>Account numbers</p></list-item>
<list-item><p>Certificate/license numbers</p></list-item>
<list-item><p>Vehicle identifiers or serial numbers, including plate numbers</p></list-item>
<list-item><p>Device identifiers or serial numbers</p></list-item>
<list-item><p>Web URLs</p></list-item>
<list-item><p>Internet protocol addresses</p></list-item>
<list-item><p>Biometric identifiers</p></list-item>
<list-item><p>Full-face photographs and comparable images</p></list-item>
<list-item><p>Any other unique identifying number, characteristic, or code</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Types of free text data</title>
<p>Of the 18 articles, eight examined free text health data from electronic health records [<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-21">21</xref>, <xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-26">26</xref>]. Another eight articles mentioned big data [<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-28">28</xref>, <xref ref-type="bibr" rid="ref-30">30</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>] but did not elaborate further on data type. Two articles mentioned both health data and big data [<xref ref-type="bibr" rid="ref-22">22</xref>, <xref ref-type="bibr" rid="ref-27">27</xref>].</p>
</sec>
<sec>
<title>Methods of de-identification for free text data</title>
<p>The de-identification approaches for free text data we found in the literature can be categorised into four overlapping groups: rule-based methods, ML methods, deep learning (a subset of machine learning) methods, and hybrid methods. The non-automated rule-based learning approaches used are summarised in <xref ref-type="table" rid="table-8">Table 8</xref>, and all other de-identification approaches and system/software packages mentioned are presented in <xref ref-type="table" rid="table-9">Table 9</xref>.</p>
<table-wrap id="table-8">
<label>Table 8: Rule-based de-identification approaches in the included articles</label>
<table frame="hsides" rules="groups">
<col width="35%"/>
<col width="30%"/>
<col width="35%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>De-Identification process</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Rule-based learning methods</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Article(s)</bold></td>
</tr>
<tr>
<td align="left" valign="top">Anonymisation Models</td>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>K-anonymity, I-diversity, t-closeness and M-variance</p></list-item>
<list-item><p><italic>&#x0392;</italic>- likeness and suppression</p></list-item>
<list-item><p>Cluster-based Missing and Value Imputation</p></list-item>
<list-item><p>Differential privacy</p></list-item>
<list-item><p>Fuzzy-based (clustering)</p></list-item>
</list></td>
<td align="left" valign="top"><list list-type="simple">
<list-item><p>[<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-23">23</xref>, <xref ref-type="bibr" rid="ref-30">30</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-12">12</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-34">34</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-28">28</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-27">27</xref>]</p></list-item>
</list></td></tr>
<tr>
<td align="left" valign="top">Data Perturbation</td>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>Value-based (e.g., uniform perturbation, probability distribution/randomisation)</p></list-item>
<list-item><p>Dimension-based (e.g., random rotation transformation, random projection)</p></list-item>
<list-item><p>Randomisation</p></list-item></list></td>
<td align="left" valign="top"><list list-type="simple">
<list-item><p>[<xref ref-type="bibr" rid="ref-27">27</xref>, <xref ref-type="bibr" rid="ref-30">30</xref>, <xref ref-type="bibr" rid="ref-32">32</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-27">27</xref>, <xref ref-type="bibr" rid="ref-31">31</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-23">23</xref>, <xref ref-type="bibr" rid="ref-27">27</xref>, <xref ref-type="bibr" rid="ref-33">33</xref>]</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-9">
<label>Table 9: Additional categories of de-identification approaches in the included articles, including systems and software</label>
<table frame="hsides" rules="groups">
<col width="25%"/>
<col width="25%"/>
<col width="25%"/>
<col width="25%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Rule-Based automated</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Machine learning</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Deep learning</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Hybrid</bold></td></tr>
<tr>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>DE-ID [<xref ref-type="bibr" rid="ref-40">40</xref>&#x2013;<xref ref-type="bibr" rid="ref-42">42</xref>]</p></list-item>
<list-item><p>National Library of Medicine Scrubber [<xref ref-type="bibr" rid="ref-43">43</xref>, <xref ref-type="bibr" rid="ref-44">44</xref>]</p></list-item>
<list-item><p>Privacy Analytics Risk Assessment Tool [<xref ref-type="bibr" rid="ref-45">45</xref>]</p></list-item>
<list-item><p>HMS Scrubber [<xref ref-type="bibr" rid="ref-36">36</xref>, <xref ref-type="bibr" rid="ref-46">46</xref>&#x2013;<xref ref-type="bibr" rid="ref-48">48</xref>]</p></list-item>
<list-item><p>MEDTAG [<xref ref-type="bibr" rid="ref-49">49</xref>, <xref ref-type="bibr" rid="ref-50">50</xref>]</p></list-item>
<list-item><p>Regenstrief Institute System [<xref ref-type="bibr" rid="ref-51">51</xref>]</p></list-item>
<list-item><p>Concept-Match [<xref ref-type="bibr" rid="ref-52">52</xref>, <xref ref-type="bibr" rid="ref-53">53</xref>]</p></list-item>
<list-item><p>VA system [<xref ref-type="bibr" rid="ref-54">54</xref>]</p></list-item>
<list-item><p>Encryption Broker Software [<xref ref-type="bibr" rid="ref-56">56</xref>]</p></list-item>
<list-item><p>Medical information anonymisation [<xref ref-type="bibr" rid="ref-57">57</xref>]</p></list-item>
<list-item><p>N-Sanitisation [<xref ref-type="bibr" rid="ref-58">58</xref>]</p></list-item>
<list-item><p>Medical De-identification System (MEDs) [<xref ref-type="bibr" rid="ref-59">59</xref>, <xref ref-type="bibr" rid="ref-60">60</xref>]</p></list-item>
<list-item><p>MedLEE [<xref ref-type="bibr" rid="ref-42">42</xref>, <xref ref-type="bibr" rid="ref-61">61</xref>]</p></list-item>
<list-item><p>Deid-Swe [<xref ref-type="bibr" rid="ref-62">62</xref>]</p></list-item>
</list></td>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>Software based on machine learning for the CEGS N-GRID 2016 de-id shared task [<xref ref-type="bibr" rid="ref-55">55</xref>]</p></list-item>
<list-item><p>MIST (Identification Scrubber Toolkit) [<xref ref-type="bibr" rid="ref-63">63</xref>, <xref ref-type="bibr" rid="ref-64">64</xref>]</p></list-item>
<list-item><p>Stat De-id [<xref ref-type="bibr" rid="ref-65">65</xref>]</p></list-item>
<list-item><p>UCLA system [<xref ref-type="bibr" rid="ref-66">66</xref>]</p></list-item>
<list-item><p>System for the 2006 i2b2 de-identification challenge (based on the Conditional Random Field [<xref ref-type="bibr" rid="ref-67">67</xref>]</p></list-item>
<list-item><p>HIDE [<xref ref-type="bibr" rid="ref-68">68</xref>]</p></list-item>
<list-item><p>System for the 2006 i2b2 de-identification challenge (based on Support Vector Machine method) [<xref ref-type="bibr" rid="ref-25">25</xref>, <xref ref-type="bibr" rid="ref-69">69</xref>&#x2013;<xref ref-type="bibr" rid="ref-71">71</xref>]</p></list-item>
<list-item><p>System for the 2006 i2b2 de-identification challenge (based on Decision Tree method) [<xref ref-type="bibr" rid="ref-72">72</xref>, <xref ref-type="bibr" rid="ref-73">73</xref>]</p></list-item>
<list-item><p>Health Information DE-identification [<xref ref-type="bibr" rid="ref-74">74</xref>]</p></list-item>
<list-item><p>Hidden Markov Models based tagger [<xref ref-type="bibr" rid="ref-71">71</xref>, <xref ref-type="bibr" rid="ref-75">75</xref>]</p></list-item>
<list-item><p>HitzalMed [<xref ref-type="bibr" rid="ref-42">42</xref>]</p></list-item>
<list-item><p>Hidden Markov Model using Dirichlet Process [<xref ref-type="bibr" rid="ref-71">71</xref>]</p></list-item>
<list-item><p>Systems based on MALLET and conditional random field (CRF) [<xref ref-type="bibr" rid="ref-76">76</xref>]</p></list-item>
</list></td>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>NeuroNER [<xref ref-type="bibr" rid="ref-64">64</xref>]</p></list-item>
<list-item><p>System based on Bi-directional Long Short-Term Memory [<xref ref-type="bibr" rid="ref-47">47</xref>, <xref ref-type="bibr" rid="ref-64">64</xref>]</p></list-item>
<list-item><p>Frequency-filtering-based system [<xref ref-type="bibr" rid="ref-53">53</xref>, <xref ref-type="bibr" rid="ref-54">54</xref>]</p></list-item>
<list-item><p>Systems based on Bidirectional Encoder Representations from Transformers and Multilingual Bidirectional Encoder Representations from Transformers [<xref ref-type="bibr" rid="ref-59">59</xref>]</p></list-item>
<list-item><p>Systems based on two variants (Elman and Jordan) of RNN [<xref ref-type="bibr" rid="ref-74">74</xref>]</p></list-item>
<list-item><p>Text Skeleton-Recurrent Neural Network (Combination of RNN and text skeleton) [<xref ref-type="bibr" rid="ref-74">74</xref>]</p></list-item>
<list-item><p>Transfer learning with RNN [<xref ref-type="bibr" rid="ref-77">77</xref>]</p></list-item>
</list></td>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>System for the 2014 i2b2 de-identification challenge [<xref ref-type="bibr" rid="ref-37">37</xref>, <xref ref-type="bibr" rid="ref-54">54</xref>, <xref ref-type="bibr" rid="ref-73">73</xref>]</p></list-item>
<list-item><p>System based on CRF and Bi-LSTM [<xref ref-type="bibr" rid="ref-48">48</xref>, <xref ref-type="bibr" rid="ref-49">49</xref>, <xref ref-type="bibr" rid="ref-64">64</xref>]</p></list-item>
<list-item><p>System based on the combination of convolutional neural network, Bi-LSTM, and CRF [<xref ref-type="bibr" rid="ref-78">78</xref>]</p></list-item>
<list-item><p>System for the 2014 i2b2 de-identification challenge (based on combination of CRF and rule-based approaches) [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-37">37</xref>&#x2013;<xref ref-type="bibr" rid="ref-39">39</xref>]</p></list-item>
<list-item><p>System for the 2016 i2b2 de-identification challenge [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>, <xref ref-type="bibr" rid="ref-58">58</xref>]</p></list-item>
<list-item><p>Multilevel Hybrid Semi-Supervised Learning Approach (MLHSLA) [<xref ref-type="bibr" rid="ref-62">62</xref>]</p></list-item>
<list-item><p>System based on mDEID and CliDEID [<xref ref-type="bibr" rid="ref-79">79</xref>]</p></list-item>
<list-item><p>System for the 2016 i2b2 de-identification challenge (based on Bi-LSTM, CRF, and rule-based approaches) [<xref ref-type="bibr" rid="ref-80">80</xref>]</p></list-item>
<list-item><p>System based on Bi-LSTM and human-engineered features from EHRs [<xref ref-type="bibr" rid="ref-81">81</xref>]</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Nine articles referred to methods like rule-based automated learning, i.e., methods created to de-identify text data automatically using HMS Scrubber, an open-source de-identification tool that employs a three-step process to remove PHII from medical documents [<xref ref-type="bibr" rid="ref-36">36</xref>], and DE-ID rule-based automated system that uses sets of rules, pattern-matching algorithms, and dictionaries to identify PHII in medical documents [<xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-21">21</xref>]. Machine learning approaches such as MIST (MITRE Identification Scrubber Toolkit, software that uses samples of de-identified text that enable it to learn contextual features that are necessary for accuracy) were mentioned in four articles, [<xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-21">21</xref>, <xref ref-type="bibr" rid="ref-24">24</xref>] the Health Information De-identification (HIDE) system was mentioned in two articles [<xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p>System/software packages containing de-identification methods can also be further divided into specific heuristic, pattern-based and statistical learning-based systems. The systems based on deep learning use a combination of specific de-identification approaches. Some articles also mentioned hybrid systems that achieved outstanding results in various natural language processing challenges pertaining to de-identification. For example, systems developed for the 2014 i2b2 challenge is a hybrid system based on machine learning and rule-based methods [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-37">37</xref>&#x2013;<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
</sec>
</sec>
<sec>
<title>Evaluation metrics</title>
<p>The metrics mentioned to measure performance by the articles are presented in <xref ref-type="table" rid="table-10">Table 10</xref>. Six out of the 18 articles mentioned evaluation metrics for assessing the performance of NLP de-identification approaches. Some articles used terms commonly used in the computer science literature such as <italic>recall</italic> and <italic>precision</italic> while others used terms that have the same meaning from epidemiology such as <italic>sensitivity</italic> and <italic>specificity</italic>. Additionally, while the articles discuss the same metrics, some of them use different formulas in varying contexts. For instance, in Kushida et al. (2012), the term <italic>precision</italic> is employed to evaluate the performance of Stat De-id, a statistical learning-based system originally introduced in Uzuner et al. (2008) [<xref ref-type="bibr" rid="ref-20">20</xref>, <xref ref-type="bibr" rid="ref-65">65</xref>]. However, in Meystre et al. (2010), the <italic>precision</italic> formula is not provided, but instead, reference is made to how HMS Scrubber was evaluated by Beckwith et al. (2006) [<xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-36">36</xref>].</p>
<table-wrap id="table-10">
<label>Table 10: NLP metrics mentioned in the included articles</label>
<table frame="hsides" rules="groups">
<col width="60%"/>
<col width="40%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Evaluation metric</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Articles</bold></td></tr>
<tr>
<td align="left" valign="top"><list list-type="bullet">
<list-item><p>Precision</p></list-item>
<list-item><p>Accuracy</p></list-item>
<list-item><p>Area under ROC curve</p></list-item>
<list-item><p>Sensitivity</p></list-item>
<list-item><p>F-measure</p></list-item>
<list-item><p>Recall</p></list-item>
<list-item><p>Specificity</p></list-item>
</list></td>
<td align="left" valign="top"><list list-type="simple">
<list-item><p>[<xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>, <xref ref-type="bibr" rid="ref-24">24</xref>, <xref ref-type="bibr" rid="ref-25">25</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>, <xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-26">26</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>, <xref ref-type="bibr" rid="ref-24">24</xref>, <xref ref-type="bibr" rid="ref-25">25</xref>]</p></list-item>
<list-item><p>[<xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>]</p></list-item></list></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Discussion</title>
<p>Free text data contain a wealth of information that is valuable in research. To take full advantage of this information, de-identification approaches for free text data must ensure the privacy and confidentiality of individuals described in the data. The discussion of de-identification of data in health research previously focused on structured data. The growth and importance of free text data in health records and health research has resulted in the need for advances in de-identification approaches. This scoping review of reviews identifies published de-identification methods for free text data. We have categorized the methods as rule-based methods, machine learning, deep learning and a combination of these and other approaches. Most of the articles we found in our search refer to de-identification methods (primarily rule-based and machine learning methods) that target some or all categories of PHII defined by HIPAA.</p>
<p>In general, experts in the field are using rule-based methods with anonymisation models to de-identify data; in particular, they use K-anonymity, I-diversity and t-closeness. Sakpere et al. (2014) assert that K-anonymity methods are best suited for data stream anonymity, such as phone numbers [<xref ref-type="bibr" rid="ref-31">31</xref>]. However, Senosi et al. (2017) found that researchers only give anonymisation strategies an average rating for protecting privacy [<xref ref-type="bibr" rid="ref-32">32</xref>]. Additionally, Stubbs et al. (2015) observes that even if automated rule-based solutions are beneficial, some PHII is still included in the data since the success of the de-identification process depends on the dictionaries used [<xref ref-type="bibr" rid="ref-25">25</xref>]. Yogarajan et al. (2020) argues that machine learning methods for de-identification need to improve in areas such as maintaining correctness and usability of data [<xref ref-type="bibr" rid="ref-26">26</xref>]. Meystre et al. (2010) states that machine learning methods combined with rule-based approaches such as HMS Scrubber perform better than a single method at de-identification of free text data [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>Recently published articles reviewed a number of approaches, including systems based on machine learning and hybrid systems that use a combination of different de-identification methods, including deep learning methods (e.g., NeuroNER and Bidirectional Encoder Representations from Transformers (BERT) [<xref ref-type="bibr" rid="ref-22">22</xref>, <xref ref-type="bibr" rid="ref-26">26</xref>]. Shickel et al. (2018) found systems based on deep learning performed better than other methods on lexical features [<xref ref-type="bibr" rid="ref-15">15</xref>]. However, deep learning techniques require large datasets to perform effectively [<xref ref-type="bibr" rid="ref-15">15</xref>]. Deep learning methods also make validating accuracy challenging due to the nature of the method. While they do represent significant progress in de-identification, the size of the required datasets for acceptable performance is an important limitation.</p>
</sec>
<sec>
<title>Conclusion</title>
<p>This scoping review provides an overview of de-identification methods for free text data. As computation power and the availability of free text from electronic health records have increased, the importance of de-identification methods in advancing the use of text data for research has also grown. While this review sought to classify de-identification techniques, no single approach or rule-based method was found to meet the high standards required to address the needs of research privacy regulators in protecting the privacy of patients since no single approach could reliably de-identify all PHII in population data records [<xref ref-type="bibr" rid="ref-20">20</xref>]. The combination of multiple tools in a hybrid format appears to be the most promising future direction.</p>
</sec>
<sec>
<title>Ethics</title>
<p>The University of Manitoba Health Research Ethics Board does not require review of review articles.</p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>We thank Li Zhang, MLIS (University of Saskatchewan Library) for peer review of the<italic> ACM Digital Library</italic> search strategy. The librarian who peer reviewed the MEDLINE search strategy does not wish to receive formal acknowledgement, but her contribution to this review is equally valued.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="ref-1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Ngiam</surname> <given-names>KY</given-names></string-name>, <string-name><surname>Khor</surname> <given-names>IW</given-names></string-name>. <article-title>Big Data and Machine Learning Algorithms for Health-Care Delivery</article-title>. <source>Lancet Oncol</source> [Internet] <year>2019</year> <month>May</month> [cited 2023 25];<volume>20</volume>:<fpage>e262</fpage>&#x2013;<lpage>73</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/S1470-2045(19)30149-4</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Tao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>H</given-names></string-name>. <article-title>Utilization of Text Mining as a big Data Analysis Tool for Food Science and Nutrition</article-title>. <source>Compr Rev Food Sci Food Saf</source> [Internet] <year>2020</year> <month>Mar</month>. [cited 2023 25];<volume>19</volume>:<fpage>875</fpage>&#x2013;<lpage>94</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1111/1541-4337.12540</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>3</label><mixed-citation publication-type="journal"><string-name><surname>Rudrapatna</surname> <given-names>VA</given-names></string-name>, <string-name><surname>Butte</surname> <given-names>AJ</given-names></string-name>. <article-title>Opportunities and Challenges in Using Real-World Data for Health Care</article-title>. <source>J Clin Invest</source> [Internet] <year>2020</year> <month>Feb</month>. [cited 2023 25];<volume>130</volume>:<fpage>565</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1172/JCI129197</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>4</label><mixed-citation publication-type="journal"><string-name><surname>Kim</surname> <given-names>E</given-names></string-name>, <string-name><surname>Rubinstein</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Nead</surname> <given-names>KT</given-names></string-name>, <string-name><surname>Wojcieszynski</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Gabriel</surname> <given-names>PE</given-names></string-name>, <string-name><surname>Warner</surname> <given-names>JL</given-names></string-name>. <article-title>The Evolving Use of Electronic Health Records (EHR) for Research</article-title>. <source>Semin Radiat Oncol</source> <year>2019</year> <month>Oct</month>.;<volume>29</volume>:<fpage>354</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.semradonc.2019.05.010</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>5</label><mixed-citation publication-type="journal"><string-name><surname>Sarwar</surname> <given-names>T</given-names></string-name>, <string-name><surname>Seifollahi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Aksakalli</surname> <given-names>V</given-names></string-name>, <string-name><surname>Hudson</surname> <given-names>I</given-names></string-name>, <etal>et al</etal>. <article-title>The Secondary Use of Electronic Health Records for Data Mining: Data Characteristics and Challenges</article-title>. <source>ACM Comput Surv</source> [Internet] <year>2022</year> [cited 2023 25];<volume>55</volume>:<fpage>33</fpage>. Available from: <pub-id pub-id-type="doi">https://doi.org/10.1145/3490234</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>6</label><mixed-citation publication-type="website"><collab>Canadian Institutes of Health Research, Natural Sciences and Engineering Research Council of Canada, Social Sciences and Humanities Research Council</collab>. <article-title>Tri-Council Policy Statement: Ethical Conduct for Research Involving Humans [Internet]</article-title>. <source>Government of Manitoba</source> <year>2018</year> [cited 2023 2]; Available from: <uri>www.nserc-crsng.gc.ca</uri>.</mixed-citation></ref>
<ref id="ref-7"><label>7</label><mixed-citation publication-type="website"><collab>Office of the Information and Privacy Commissioner of Ontario</collab>. <article-title>De-identification Guidelines for Structured Data [Internet]</article-title>. <source>Information and Privacy Commissioner of Ontario</source> <year>2016</year> [cited 2023 28]; Available from: <uri>https://www.ipc.on.ca/resource/de-identification-guidelines-for-structured-data/</uri>.</mixed-citation></ref>
<ref id="ref-8"><label>8</label><mixed-citation publication-type="website"><string-name><surname>Abu-El-Rub</surname> <given-names>N</given-names></string-name>, <string-name><surname>Urbain</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kowalski</surname> <given-names>G</given-names></string-name>, <string-name><surname>Osinski</surname> <given-names>K</given-names></string-name>, <string-name><surname>Spaniol</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <etal>et al</etal>. <article-title>Natural Language Processing for Enterprise-scale De-identification of Protected Health Information in Clinical Notes</article-title>. [cited 2023] Available from: <uri>https://pubmed.ncbi.nlm.nih.gov/35854742/</uri>.</mixed-citation></ref>
<ref id="ref-9"><label>9</label><mixed-citation publication-type="website"><collab>University of Manitoba</collab>. <article-title>Access and Privacy [Internet]</article-title>. <source>Access and Privacy Office</source> <year>2023</year> [cited 2023 28]; Available from: <uri>https://umanitoba.ca/access-and-privacy/</uri>.</mixed-citation></ref>
<ref id="ref-10"><label>10</label><mixed-citation publication-type="website"><string-name><surname>Grolemund</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wickham</surname> <given-names>H</given-names></string-name>. <article-title>R for data science [Internet]</article-title>. <year>2019</year>. Available from: <uri>https://r4ds.had.co.nz/</uri>.</mixed-citation></ref>
<ref id="ref-11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Lee</surname> <given-names>HJ</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>K</given-names></string-name>. <article-title>A Hybrid Approach to Automatic De-identification of Psychiatric Notes</article-title>. <source>J Biomed Inform</source> [Internet] <year>2017</year> <month>Nov</month>. [cited 2023 2];<volume>75S</volume>:<fpage>S19</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2017.06.006</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12</label><mixed-citation publication-type="book"><string-name><surname>Basso</surname> <given-names>T</given-names></string-name>, <string-name><surname>Matsunaga</surname> <given-names>R</given-names></string-name>, <string-name><surname>Moraes</surname> <given-names>R</given-names></string-name>, <string-name><surname>Antunes</surname> <given-names>N</given-names></string-name>. <chapter-title>Challenges on Anonymity, Privacy, and big Data</chapter-title>. In: <source>Proceedings - 7th Latin-American Symposium on Dependable Computing</source>, <publisher-name>LADC</publisher-name> <year>2016</year>. 2016. <pub-id pub-id-type="doi">https://doi.org/10.1109/LADC.2016.34</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>13</label><mixed-citation publication-type="journal"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <etal>et al</etal>. <article-title>Automatic De-identification of Electronic Medical Records Using Token-level and Character-Level Conditional Random Fields</article-title>. <source>J Biomed Inform</source> <year>2015</year> <month>Dec</month>.;<volume>58</volume>:<fpage>S47</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2015.06.009</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14</label><mixed-citation publication-type="journal"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lyu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Bian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hogan</surname> <given-names>WR</given-names></string-name>, <etal>et al</etal>. <article-title>A study of deep learning methods for de-identification of clinical notes in cross-institute settings</article-title>. <source>BMC Med Inform Decis Mak</source> [Internet] <year>2019</year> <month>Dec</month>. [cited 2023 30];<volume>19</volume>:<fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1109/ICHI.2019.8904544</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>15</label><mixed-citation publication-type="journal"><string-name><surname>Shickel</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tighe</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Bihorac</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rashidi</surname> <given-names>P</given-names></string-name>. <article-title>Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis</article-title>. <source>Institute of Electrical and Electronics Engineers Journal of Biomedical and Health Informatics</source> <year>2018</year>;<volume>22</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1109/JBHI.2017.2767063</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>Arksey</surname> <given-names>H</given-names></string-name>, <string-name><surname>O&#x2019;Malley</surname> <given-names>L</given-names></string-name>. <article-title>Scoping Studies: Towards a Methodological Framework</article-title>. <source>International Journal of Social Research Methodology: Theory and Practice</source> <year>2005</year>;<volume>8</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1080/1364557032000119616</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>17</label><mixed-citation publication-type="journal"><string-name><surname>McGowan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sampson</surname> <given-names>M</given-names></string-name>, <string-name><surname>Salzwedel</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Cogo</surname> <given-names>E</given-names></string-name>, <string-name><surname>Foerster</surname> <given-names>V</given-names></string-name>, <string-name><surname>Lefebvre</surname> <given-names>C</given-names></string-name>. <article-title>PRESS Peer Review of Electronic Search Strategies: 2015 Guideline Statement</article-title>. <source>J Clin Epidemiol</source> <year>2016</year> <month>Jul</month>.;<volume>75</volume>:<fpage>40</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/J.JCLINEPI.2016.01.021</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>18</label><mixed-citation publication-type="journal"><string-name><surname>Tricco</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Lillie</surname> <given-names>E</given-names></string-name>, <string-name><surname>Zarin</surname> <given-names>W</given-names></string-name>, <string-name><surname>O&#x2019;Brian</surname> <given-names>KK</given-names></string-name>, <string-name><surname>Colquhoun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Levac</surname> <given-names>D</given-names></string-name>, <etal>et al</etal>. <article-title>PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation</article-title>. <source>Ann Intern Med</source> <year>2018</year>;<volume>169</volume>. <pub-id pub-id-type="doi">https://doi.org/10.7326/M18-0850</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>19</label><mixed-citation publication-type="journal"><string-name><surname>Meystre</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Friedlin</surname> <given-names>FJ</given-names></string-name>, <string-name><surname>South</surname> <given-names>BR</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Samore</surname> <given-names>MH</given-names></string-name>. <article-title>Automatic De-identification of Textual Documents in the Electronic Health Record: A Review of Recent Research</article-title>. <source>BMC Med Res Methodol</source> <year>2010</year>;<volume>10</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1186/1471-2288-10-70</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>20</label><mixed-citation publication-type="journal"><string-name><surname>Kushida</surname> <given-names>CA</given-names></string-name>, <string-name><surname>Nichols</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Jadrnicek</surname> <given-names>R</given-names></string-name>, <string-name><surname>Miller</surname> <given-names>R</given-names></string-name>, <string-name><surname>Walsh</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Griffin</surname> <given-names>K</given-names></string-name>. <article-title>Strategies for De-identification and Anonymization of Electronic Health Record Data for use in Multicenter Research Studies</article-title>. <source>Med Care</source> <year>2012</year>;<volume>50</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1097/MLR.0b013e3182585355</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Kayaalp</surname> <given-names>M</given-names></string-name>. <article-title>Patient Privacy in the era of big Data</article-title>. <source>Balkan Med J</source> <year>2018</year>;<volume>35</volume>. <pub-id pub-id-type="doi">https://doi.org/10.4274/balkanmedj.2017.0966</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>22</label><mixed-citation publication-type="journal"><string-name><surname>Mahendran</surname> <given-names>D</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>C</given-names></string-name>, <string-name><surname>McInnes</surname> <given-names>BT</given-names></string-name>. <article-title>Review: Privacy-Preservation in the Context of Natural Language Processing</article-title>. <source>Institute of Electrical and Electronics Engineers Access</source> [Internet] <year>2021</year> [cited 2023 10];<volume>9</volume>. Available from: <pub-id pub-id-type="doi">https://doi.org/10.1109/ACCESS.2021.3124163</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>23</label><mixed-citation publication-type="journal"><string-name><surname>Xiang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>W</given-names></string-name>. <article-title>Privacy Protection and Secondary Use of Health Data: Strategies and Methods</article-title>. <source>Biomed Res Int</source> <year>2021</year>;2021. <pub-id pub-id-type="doi">https://doi.org/10.1155/2021/6967166</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>24</label><mixed-citation publication-type="journal"><string-name><surname>Stubbs</surname> <given-names>A</given-names></string-name>, <string-name><surname>Filannino</surname> <given-names>M</given-names></string-name>, <string-name><surname>Uzuner</surname> <given-names>&#x00D6;</given-names></string-name>. <article-title>De-identification of Psychiatric Intake Records: Overview of 2016 CEGS N-GRID shared tasks Track 1</article-title>. <source>J Biomed Inform</source> <year>2017</year>;<volume>75</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2017.06.011</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>25</label><mixed-citation publication-type="journal"><string-name><surname>Stubbs</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kotfila</surname> <given-names>C</given-names></string-name>, <string-name><surname>Uzuner</surname> <given-names>&#x00D6;</given-names></string-name>. <article-title>Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth shared task Track 1</article-title>. <source>J Biomed Inform</source> <year>2015</year>;<volume>58</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2015.06.007</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>26</label><mixed-citation publication-type="journal"><string-name><surname>Yogarajan</surname> <given-names>V</given-names></string-name>, <string-name><surname>Pfahringer</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mayo</surname> <given-names>M</given-names></string-name>. <article-title>A Review of Automatic end-to-end De-Identification: Is High Accuracy the Only Metric?</article-title> <source>Applied Artificial Intelligence</source> <year>2020</year>;<volume>34</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1080/08839514.2020.1718343</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>27</label><mixed-citation publication-type="journal"><string-name><surname>Zainab</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Kechadi</surname> <given-names>T</given-names></string-name>. <article-title>Sensitive and private data analysis: A systematic review</article-title>. In: <source>ACM International Conference Proceeding Series</source>. <year>2019</year>. <pub-id pub-id-type="doi">https://doi.org/10.1145/3341325.3342002</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>28</label><mixed-citation publication-type="journal"><string-name><surname>Youm</surname> <given-names>HY</given-names></string-name>. <article-title>An Overview of De-identification Techniques and Their Standardization Directions</article-title>. <source>Institute of Electronics, Information and Communication EngineersTransactions on Information and Systems</source> <year>2020</year>;<fpage>E103D</fpage>. <pub-id pub-id-type="doi">https://doi.org/10.1587/transinf.2019ICI0002</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>29</label><mixed-citation publication-type="journal"><string-name><surname>Stubbs</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kotfila</surname> <given-names>C</given-names></string-name>, <string-name><surname>Uzuner</surname> <given-names>&#x00D6;</given-names></string-name>. <article-title>Automated systems for the de-identification of longitudinal clinical narratives: Overview of 2014 i2b2/UTHealth shared task Track 1</article-title>. <source>J Biomed Inform</source> <year>2015</year>;<volume>58</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2015.06.007</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>30</label><mixed-citation publication-type="journal"><string-name><surname>Binjubeir</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Ismail</surname> <given-names>MA Bin</given-names></string-name>, <string-name><surname>Sadiq</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Khurram Khan</surname> <given-names>M</given-names></string-name>. <article-title>Comprehensive Survey on big Data Privacy Protection</article-title>. <source>Institute of Electrical and Electronics Engineers</source>. Access <year>2020</year>;<volume>8</volume>:<fpage>20067</fpage>&#x2013;<lpage>79</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1109/ACCESS.2019.2962368</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>31</label><mixed-citation publication-type="journal"><string-name><surname>Sakpere</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Kayem</surname> <given-names>AVDM</given-names></string-name>. <article-title>A State-of-the-art Review of Data Stream Anonymization Schemes</article-title>. In: <source>Information Security in Diverse Computing Environments</source>. <year>2014</year>. <pub-id pub-id-type="doi">https://doi.org/10.4018/978-1-4666-6158-5.ch003</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>32</label><mixed-citation publication-type="journal"><string-name><surname>Senosi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sibiya</surname> <given-names>G</given-names></string-name>. <article-title>Classification and Evaluation of Privacy Preserving Data Mining: A review</article-title>. In: <source>Institute of Electrical and Electronics Engineers AFRICON: Science, Technology and Innovation for Africa</source>. <year>2017</year>. <pub-id pub-id-type="doi">https://doi.org/10.1109/AFRCON.2017.8095593</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>33</label><mixed-citation publication-type="book"><string-name><surname>Shanthi</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Karthikeyan</surname> <given-names>M</given-names></string-name>. <chapter-title>A Review on Privacy Preserving Data Mining</chapter-title>. <source>2012 IEEE International Conference on Computational Intelligence and Computing Research</source>, <publisher-name>ICCIC</publisher-name> <year>2012</year> 2012; <pub-id pub-id-type="doi">https://doi.org/10.1109/ICCIC.2012.6510302</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Deng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>. <article-title>Overview of Privacy Protection Data Release Anonymity Technology</article-title>. <source>International Conference on Big Data Security on Cloud (BigDataSecurity), IEEE Intl Conference on High Performance and Smart Computing, (HPSC) Conference on Intelligent Data and Security (IDS)</source> <year>2021</year> <month>May</month>;<fpage>151</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1109/BigDataSecurityHPSCIDS52275.2021.00037</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>35</label><mixed-citation publication-type="book"><string-name><surname>Shelake</surname> <given-names>VM</given-names></string-name>, <string-name><surname>Shekokar</surname> <given-names>N</given-names></string-name>. <chapter-title>A Survey of Privacy Preserving Data Integration</chapter-title>. In: <source>International Conference on Electrical, Electronics, Communication Computer Technologies and OptimizationTechniques</source>, <publisher-name>ICEECCOT</publisher-name>. <year>2017</year>. <pub-id pub-id-type="doi">https://doi.org/10.1109/ICEECCOT.2017.8284559</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>36</label><mixed-citation publication-type="journal"><string-name><surname>Beckwith</surname> <given-names>BA</given-names></string-name>, <string-name><surname>Mahaadevan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Balis</surname> <given-names>UJ</given-names></string-name>, <string-name><surname>Kuo</surname> <given-names>F</given-names></string-name>. <article-title>Development and Evaluation of an Open Source Software Tool for De-identification of Pathology Reports</article-title>. <source>BMC Med Inform Decis Mak</source> <year>2006</year>;<volume>6</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1186/1472-6947-6-12</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>37</label><mixed-citation publication-type="journal"><string-name><surname>He</surname> <given-names>B</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hua</surname> <given-names>W</given-names></string-name>. <article-title>CRFs Based De-identification of Medical Records</article-title>. <source>J Biomed Inform</source> <year>2015</year> <month>Dec</month>.;<volume>58</volume>:<fpage>S39</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/J.JBI.2015.08.012</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Garibaldi</surname> <given-names>JM</given-names></string-name>. <article-title>Automatic Detection of Protected Health Information From Clinic Narratives</article-title>. <source>J Biomed Inform</source> [Internet] <year>2015</year> <month>Dec</month>. [cited 2023 2];<volume>58</volume> <supplement>Suppl</supplement>:<fpage>S30</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2015.06.015</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>39</label><mixed-citation publication-type="journal"><string-name><surname>Dehghan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kovacevic</surname> <given-names>A</given-names></string-name>, <string-name><surname>Karystianis</surname> <given-names>G</given-names></string-name>, <string-name><surname>Keane</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Nenadic</surname> <given-names>G</given-names></string-name>. <article-title>Combining Knowledge- and Data-Driven Methods for De-identification of Clinical Narratives</article-title>. <source>J Biomed Inform</source> [Internet] <year>2015</year> <month>Dec</month>. [cited 2023 2];<volume>58</volume> <supplement>Suppl</supplement>:<fpage>S53</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2015.06.029</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>40</label><mixed-citation publication-type="journal"><string-name><surname>Neamatullah</surname> <given-names>I</given-names></string-name>, <string-name><surname>Douglass</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Lehman</surname> <given-names>LH</given-names></string-name>, <string-name><surname>Reisner</surname> <given-names>A</given-names></string-name>, <string-name><surname>Villarroel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Long</surname> <given-names>WJ</given-names></string-name>, <etal>et al</etal>. <article-title>Automated De-identification of Free-Text Medical Records</article-title>. <source>BMC Med Inform Decis Mak</source> <year>2008</year>;<volume>17</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1186/1472-6947-8-32</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>41</label><mixed-citation publication-type="journal"><string-name><surname>Gupta</surname> <given-names>D</given-names></string-name>, <string-name><surname>Saul</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gilbertson</surname> <given-names>J</given-names></string-name>. <article-title>Evaluation of a Deidentification (De-Id) Software Engine to Share Pathology Reports and Clinical Documents for Research</article-title>. <source>Am J Clin Pathol</source> <year>2004</year>;<fpage>121</fpage>. <pub-id pub-id-type="doi">https://doi.org/10.1309/e6k3-3gbp-e5c2-7fyu</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>42</label><mixed-citation publication-type="website"><string-name><surname>Lima</surname> <given-names>S</given-names></string-name>, <string-name><surname>Perez</surname> <given-names>N</given-names></string-name>, <string-name><surname>Garc&#x00ED;a-Sardi&#x00F1;a</surname> <given-names>L</given-names></string-name>, <string-name><surname>Sardi&#x00F1;a</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cuadros</surname> <given-names>M</given-names></string-name>. <article-title>HitzalMed: Anonymisation of Clinical Text in Spanish [Internet]</article-title>. In: <source>Twelfth Language Resources and Evaluation Conference</source>. <year>2020</year> [cited 2023 2]. p. <fpage>7038</fpage>&#x2013;<lpage>43</lpage>. Available from: <uri>https://aclanthology.org/2020.lrec-1.870</uri>.</mixed-citation></ref>
<ref id="ref-43"><label>43</label><mixed-citation publication-type="website"><string-name><surname>Kayaalp</surname> <given-names>M</given-names></string-name>, <string-name><surname>Browne</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Dodd</surname> <given-names>ZA</given-names></string-name>, <string-name><surname>Sagan</surname> <given-names>P</given-names></string-name>, <string-name><surname>McDonald</surname> <given-names>CJ</given-names></string-name>. <article-title>De-identification of Address, Date, and Alphanumeric Identifiers in Narrative Clinical Reports</article-title>. <source>AMIA ... Annual Symposium proceedings / AMIA Symposium. AMIA Symposium</source> [Internet] <year>2014</year> [cited 2023 10];2014. Available from: <uri>https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4419982/</uri>.</mixed-citation></ref>
<ref id="ref-44"><label>44</label><mixed-citation publication-type="journal"><string-name><surname>Kayaalp</surname> <given-names>M</given-names></string-name>, <string-name><surname>Browne</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Callaghan</surname> <given-names>FM</given-names></string-name>, <string-name><surname>Dodd</surname> <given-names>ZA</given-names></string-name>, <string-name><surname>Divita</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ozturk</surname> <given-names>S</given-names></string-name>, <etal>et al</etal>. <article-title>The Pattern of Name Tokens in Narrative Clinical Text and a Comparison of Five Systems for Redacting Them</article-title>. <source>Journal of the American Medical Informatics Association</source> <year>2014</year>;<volume>21</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1136/amiajnl-2013-001689</pub-id></mixed-citation></ref>
<ref id="ref-45"><label>45</label><mixed-citation publication-type="website"><collab>Privacy Analytics Data Anonymization Solution</collab>. <article-title>PARAT Maintenance and Support Information |Privacy Analytic&#x2019;s Privacy and Confidentiality KnowledgeBase [Internet]</article-title>. <year>2023</year> [cited 2023 2]; Available from: <uri>http://knowledgebase.privacy-analytics.com/index.php?/article/AA-00335/0/PARAT-Maintenance-and-Support-Information.html</uri>.</mixed-citation></ref>
<ref id="ref-46"><label>46</label><mixed-citation publication-type="website"><string-name><surname>Sweeney</surname> <given-names>L</given-names></string-name>. <article-title>Replacing Personally-Identifying Information in Medical Records, the Scrub System. Proceedings: a conference of the American Medical Informatics Association / ..</article-title>. <source>AMIA Annual Fall Symposium. AMIA Fall Symposium</source> [Internet] <year>1996</year> [cited 2023 10]; Available from: <uri>https://pubmed.ncbi.nlm.nih.gov/8947683/</uri>.</mixed-citation></ref>
<ref id="ref-47"><label>47</label><mixed-citation publication-type="journal"><string-name><surname>Dernoncourt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Uzuner</surname> <given-names>O</given-names></string-name>, <string-name><surname>Szolovits</surname> <given-names>P</given-names></string-name>. <article-title>De-identification of Patient Notes With Recurrent Neural Networks</article-title>. <source>Journal of the American Medical Informatics Association</source> [Internet] <year>2017</year> <month>May</month>[cited 2023 2];<volume>24</volume>:<fpage>596</fpage>&#x2013;<lpage>606</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1093/jamia/ocw156</pub-id></mixed-citation></ref>
<ref id="ref-48"><label>48</label><mixed-citation publication-type="journal"><string-name><surname>Catelli</surname> <given-names>R</given-names></string-name>, <string-name><surname>Casola</surname> <given-names>V</given-names></string-name>, <string-name><surname>De Pietro</surname> <given-names>G</given-names></string-name>, <string-name><surname>Fujita</surname> <given-names>H</given-names></string-name>, <string-name><surname>Esposito</surname> <given-names>M</given-names></string-name>. <article-title>Combining Contextualized Word Representation and Sub-document Level Analysis Through Bi-LSTM+CRF Architecture for Clinical De-identification</article-title>. <source>Knowledge-Based System</source> <year>2021</year>;<volume>213</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.knosys.2020.106649</pub-id></mixed-citation></ref>
<ref id="ref-49"><label>49</label><mixed-citation publication-type="journal"><string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>C</given-names></string-name>, <string-name><surname>He</surname> <given-names>B</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>. <article-title>De-identification of Medical Records Using Conditional Random Fields and Long Short-term Memory Networks</article-title>. <source>J Biomed Inform</source> [Internet] <year>2017</year> <month>Nov</month>. [cited 2023 2];<volume>75S</volume>:<fpage>S43</fpage>&#x2013;<lpage>53</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2017.10.003</pub-id></mixed-citation></ref>
<ref id="ref-50"><label>50</label><mixed-citation publication-type="website"><string-name><surname>Ruch</surname> <given-names>P</given-names></string-name>, <string-name><surname>Baud</surname> <given-names>RH</given-names></string-name>, <string-name><surname>Rassinoux</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Bouillon</surname> <given-names>P</given-names></string-name>, <string-name><surname>Robert</surname> <given-names>G</given-names></string-name>. <article-title>Medical Document Anonymization with a Semantic Lexicon</article-title>. <source>PMC- Proceedings AMIA Symposium</source> [Internet] <year>2000</year> [cited 2023 2];<fpage>729</fpage>&#x2013;<lpage>33</lpage>. Available from: <uri>https://pubmed.ncbi.nlm.nih.gov/11079980/</uri>.</mixed-citation></ref>
<ref id="ref-51"><label>51</label><mixed-citation publication-type="website"><string-name><surname>Thomas</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Mamlin</surname> <given-names>B</given-names></string-name>, <string-name><surname>Schadow</surname> <given-names>G</given-names></string-name>, <string-name><surname>McDonald</surname> <given-names>C</given-names></string-name>. <article-title>A Successful Technique for Removing Names in Pathology Reports Using an Augmented Search and Replace Method</article-title>. <source>Proceedings of the AMIA Symposium</source> [Internet] <year>2002</year> [cited 2023 3];<volume>777</volume>. Available from: <uri>https://pubmed.ncbi.nlm.nih.gov/12463930/</uri>.</mixed-citation></ref>
<ref id="ref-52"><label>52</label><mixed-citation publication-type="journal"><string-name><surname>Berman</surname> <given-names>JJ</given-names></string-name>. <article-title>Concept-Match Medical Data Scrubbing: How Pathology Text can be Used in Research</article-title>. <source>Arch Pathol Lab Med</source> <year>2003</year>;<volume>127</volume>. <pub-id pub-id-type="doi">https://doi.org/10.5858/2003-127-680-CMDS</pub-id></mixed-citation></ref>
<ref id="ref-53"><label>53</label><mixed-citation publication-type="book"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Rastegar-Mojarad</surname> <given-names>M</given-names></string-name>, <string-name><surname>Elayavilli</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mehrabi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal>. <chapter-title>A Frequency- Filtering Strategy of Obtaining PHI-Free Sentences From Clinical Data Repository [Internet]</chapter-title>. In: <source>ACM Conference on Bioinformatics, Computational Biology, and Health Informatics</source>. <publisher-name>Association for Computing Machinery, Inc</publisher-name>; <year>2015</year> [cited 2023 2]. p. <fpage>315</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1145/2808719.2808752</pub-id></mixed-citation></ref>
<ref id="ref-54"><label>54</label><mixed-citation publication-type="journal"><string-name><surname>Sadat</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Aziz</surname> <given-names>MM Al</given-names></string-name>, <string-name><surname>Mohammed</surname> <given-names>N</given-names></string-name>, <string-name><surname>Pakhomov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>. <article-title>A Privacy-Preserving Distributed Filtering Framework for NLP Artifacts</article-title>. <source>BMC Med Inform Decis Mak</source> [Internet] <year>2019</year> <month>Sep</month>. [cited 2023 2];<volume>19</volume>:<fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1186/s12911-019-0867-z</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>55</label><mixed-citation publication-type="journal"><string-name><surname>Bui</surname> <given-names>DDA</given-names></string-name>, <string-name><surname>Wyatt</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cimino</surname> <given-names>JJ</given-names></string-name>. <article-title>The UAB Informatics Institute and 2016 CEGS N-GRID De-identification Shared Task Challenge</article-title>. <source>J Biomed Inform</source> <year>2017</year> <month>Nov</month>.;<volume>75</volume>:<fpage>S54</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2017.05.001</pub-id></mixed-citation></ref>
<ref id="ref-56"><label>56</label><mixed-citation publication-type="journal"><string-name><surname>Pestian</surname> <given-names>JP</given-names></string-name>, <string-name><surname>Itert</surname> <given-names>L</given-names></string-name>, <string-name><surname>Andersen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Duch</surname> <given-names>W</given-names></string-name>. <article-title>Preparing Clinical Text for Use in Biomedical Research</article-title>. <source>Journal of Database Management</source> [Internet] <year>2006</year> <month>Jan</month>. [cited 2023 2]. <pub-id pub-id-type="doi">https://doi.org/10.4018/jdm.2006040101</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>57</label><mixed-citation publication-type="journal"><string-name><surname>Grouin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Rosier</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dameron</surname> <given-names>O</given-names></string-name>, <string-name><surname>Zweigenbaum</surname> <given-names>P</given-names></string-name>. <article-title>Testing Tactics to Localize De-Identification</article-title>. <source>Stud Health Technol Inform</source> [Internet] <year>2009</year> [cited 2023 3];<volume>150</volume>:<fpage>735</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.3233/978-1-60750-044-5-735</pub-id></mixed-citation></ref>
<ref id="ref-58"><label>58</label><mixed-citation publication-type="journal"><string-name><surname>Iwendi</surname> <given-names>C</given-names></string-name>, <string-name><surname>Moqurrab</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Anjum</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mohan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Srivastava</surname> <given-names>G</given-names></string-name>. <article-title>N-Sanitization: A Semantic Privacy-Preserving Framework for Unstructured Medical Datasets</article-title>. <source>Comput Commun</source> <year>2020</year>;<volume>161</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.comcom.2020.07.032</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>59</label><mixed-citation publication-type="journal"><string-name><surname>Catelli</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gargiulo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Casola</surname> <given-names>V</given-names></string-name>, <string-name><surname>De Pietro</surname> <given-names>G</given-names></string-name>, <string-name><surname>Fujita</surname> <given-names>H</given-names></string-name>, <string-name><surname>Esposito</surname> <given-names>M</given-names></string-name>. <article-title>Crosslingual named entity recognition for clinical de-identification applied to a COVID-19 Italian data set</article-title>. <source>Appl Soft Comput</source> <year>2020</year> <month>Dec</month>.;<volume>97</volume>:<fpage>106779</fpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.asoc.2020.106779</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>60</label><mixed-citation publication-type="journal"><string-name><surname>Friedlin</surname> <given-names>FJ</given-names></string-name>, <string-name><surname>McDonald</surname> <given-names>CJ</given-names></string-name>. <article-title>A Software Tool for Removing Patient Identifying Information From Clinical Documents</article-title>. <source>J Am Med Inform Assoc</source> [Internet] <year>2008</year> <month>Sep</month>. [cited 2023 2];<volume>15</volume>:<fpage>601</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1197/jamia.M2702</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>61</label><mixed-citation publication-type="journal"><string-name><surname>Morrison</surname> <given-names>FP</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Hripcsak</surname> <given-names>G</given-names></string-name>. <article-title>Repurposing the clinical record: can an existing natural language processing system de-identify clinical notes?</article-title> <source>J Am Med Inform Assoc [Internet]</source> <year>2009</year> <month>Jan</month>. [cited 2023 2];<volume>16</volume>:<fpage>37</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1197/jamia.M2862</pub-id></mixed-citation></ref>
<ref id="ref-62"><label>62</label><mixed-citation publication-type="journal"><string-name><surname>Velupillai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Dalianis</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hassel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nilsson</surname> <given-names>GH</given-names></string-name>. <article-title>Developing a Standard for De-identifying Electronic Patient Records Written in Swedish: Precision, Recall and F-measure in a Manual and Computerized Annotation Trial</article-title>. <source>Int J Med Inform [Internet]</source> <year>2009</year> <month>Dec</month>. [cited 2023 2];<volume>78</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.ijmedinf.2009.04.005</pub-id></mixed-citation></ref>
<ref id="ref-63"><label>63</label><mixed-citation publication-type="journal"><string-name><surname>Wellner</surname> <given-names>B</given-names></string-name>, <string-name><surname>Huyck</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mardis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Aberdeen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Morgan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Peshkin</surname> <given-names>L</given-names></string-name>, <etal>et al</etal>. <article-title>Rapidly Retargetable Approaches to De-identification in Medical Records</article-title>. <source>Journal of the American Medical Informatics Association</source> <year>2007</year>;<volume>14</volume>. <pub-id pub-id-type="doi">https://doi.org/10.1197/jamia.M2435</pub-id></mixed-citation></ref>
<ref id="ref-64"><label>64</label><mixed-citation publication-type="book"><string-name><surname>Dernoncourt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Szolovits</surname> <given-names>P</given-names></string-name>. <chapter-title>NeuroNER: an Easy-to-use Program for Named-Entity Recognition Based on Neural Networks [Internet]</chapter-title>. In: <source>Empirical Methods in Natural Language Processing:System Demonstrations</source>. <publisher-name>Association for Computational Linguistics (ACL)</publisher-name>; <year>2017</year> [cited 2023 2]. p. <fpage>97</fpage>&#x2013;<lpage>102</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.18653/v1/D17-2017</pub-id></mixed-citation></ref>
<ref id="ref-65"><label>65</label><mixed-citation publication-type="journal"><string-name><surname>Uzuner</surname> <given-names>&#x00D6;</given-names></string-name>, <string-name><surname>Sibanda</surname> <given-names>TC</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Szolovits</surname> <given-names>P</given-names></string-name>. <article-title>A De-identifier for Medical Discharge Summaries</article-title>. <source>Artif Intell Med</source> [Internet] <year>2008</year> <month>Jan</month>. [cited 2023 2];<volume>42</volume>:<fpage>13</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.artmed.2007.10.001</pub-id></mixed-citation></ref>
<ref id="ref-66"><label>66</label><mixed-citation publication-type="website"><string-name><surname>Taira</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Bui</surname> <given-names>AAT</given-names></string-name>, <string-name><surname>Kangarloo</surname> <given-names>H</given-names></string-name>. <article-title>Identification of Patient Name References Within Medical Documents Using Semantic Selectional Restrictions</article-title>. <source>Proceedings of the AMIA Symposium</source> [Internet] <year>2002</year> [cited 2023 3];<volume>757</volume>. Available from: <uri>https://www.ncbi.knlm.nih.gov/pmc/articles/PMC2244274/</uri>.</mixed-citation></ref>
<ref id="ref-67"><label>67</label><mixed-citation publication-type="website"><string-name><surname>Aramaki</surname> <given-names>E</given-names></string-name>, <string-name><surname>Imai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Miyo</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ohe</surname> <given-names>K</given-names></string-name>. <article-title>Automatic De-identification by Using Sentence Features and Label Consistency. i2b2 Workshop on Challenges in Natural Language Processing for Clinical Data [Internet]</article-title> <year>2006</year>; . Available from: <uri>https://www.i2b2.org/NLP/DataSets/Publications.php</uri>.</mixed-citation></ref>
<ref id="ref-68"><label>68</label><mixed-citation publication-type="journal"><string-name><surname>Gardner</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>L</given-names></string-name>. <article-title>HIDE: An Integrated System for Health Information DE-identification [Internet]</article-title>. In: <source>International Symposium on Computer-Based Medical Systems</source>. <year>2008</year> [cited 2023 2]. p. <fpage>254</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1109/CBMS.2008.129</pub-id></mixed-citation></ref>
<ref id="ref-69"><label>69</label><mixed-citation publication-type="website"><string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gaizauskas</surname> <given-names>R</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>I</given-names></string-name>, <string-name><surname>Demetriou</surname> <given-names>G</given-names></string-name>, <string-name><surname>Hepple</surname> <given-names>M</given-names></string-name>. <article-title>Identifying Personal Health Information Using Support Vector Machines</article-title>. [Internet]. In: <source>i2b2 workshop in challenges in natural language processing for clinical data</source> <year>2006</year>. 2006 [cited 2023 10]. Available from: <uri>https://www.i2b2.org/NLP/DataSets/Publications.php</uri>.</mixed-citation></ref>
<ref id="ref-70"><label>70</label><mixed-citation publication-type="book"><string-name><surname>Hara</surname> <given-names>K</given-names></string-name>. <chapter-title>Applying a SVM Based Chunker and a Text Classifier to the Deid Challenge</chapter-title>. [Internet]. In: <source>i2b2 Workshop on Challenges in Natural Language Processing for Clinical Data</source>. <publisher-loc>Washington</publisher-loc>: <year>2006</year> [cited 2023 10]. Available from: <uri>https://www.i2b2.org/NLP/DataSets/Publications.php</uri>.</mixed-citation></ref>
<ref id="ref-71"><label>71</label><mixed-citation publication-type="journal"><string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cullen</surname> <given-names>RM</given-names></string-name>, <string-name><surname>Godwin</surname> <given-names>M</given-names></string-name>. <article-title>Hidden Markov Model Using Dirichlet Process for De-identification</article-title>. <source>J Biomed Inform</source> <year>2015</year> <month>Dec</month>.;<volume>58</volume>:<fpage>S60</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/J.JBI.2015.09.004</pub-id></mixed-citation></ref>
<ref id="ref-72"><label>72</label><mixed-citation publication-type="journal"><string-name><surname>Szarvas</surname> <given-names>G</given-names></string-name>, <string-name><surname>Farkas</surname> <given-names>R</given-names></string-name>, <string-name><surname>Busa-Fekete</surname> <given-names>R</given-names></string-name>. <article-title>State-of-the-art Anonymization of Medical Records Using an Iterative Machine Learning Framework</article-title>. <source>Journal of the American Medical Informatics Association</source> <year>2007</year>;<fpage>14</fpage>. <pub-id pub-id-type="doi">https://doi.org/10.1197/j.jamia.M2441</pub-id></mixed-citation></ref>
<ref id="ref-73"><label>73</label><mixed-citation publication-type="book"><string-name><surname>Torii</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wiley</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zisook</surname> <given-names>D</given-names></string-name>. <chapter-title>De-Identification and Risk Factor Detection in Medical Records [Internet]</chapter-title>. In: <source>Seventh i2b2 Shared Task and Workshop: Challenges in Natural Language Processing for Clinical Data</source>. <publisher-loc>Washington</publisher-loc>: <year>2014</year> [cited 2023 10]. <uri>https://www.i2b2.org/NLP/HeartDisease/assets/i2b2_2014_schedule_revised.pdf</uri></mixed-citation></ref>
<ref id="ref-74"><label>74</label><mixed-citation publication-type="website"><string-name><surname>Yadav</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ekbal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Saha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bhattacharyya</surname> <given-names>P</given-names></string-name>. <article-title>Deep Learning Architecture for Patient Data De-identification in Clinical Records [Internet]</article-title>. In: <source>Clinical Natural Language Processing</source>. <year>2016</year> [cited 2023 2]. p. <fpage>32</fpage>&#x2013;<lpage>41</lpage>. Available from: <uri>https://aclanthology.org/W16-4206</uri>.</mixed-citation></ref>
<ref id="ref-75"><label>75</label><mixed-citation publication-type="website"><string-name><surname>Medlock</surname> <given-names>B</given-names></string-name>. <article-title>An Introduction to NLP-based Textual Anonymisation [Internet]</article-title>. In: <source>Fifth International Conference on Language Resources and Evaluation</source>. <year>2006</year> [cited 2023 2]. Available from: <uri>https://aclanthology.org/L06-1110/</uri>.</mixed-citation></ref>
<ref id="ref-76"><label>76</label><mixed-citation publication-type="journal"><string-name><surname>Deleger</surname> <given-names>L</given-names></string-name>, <string-name><surname>Molnar</surname> <given-names>K</given-names></string-name>, <string-name><surname>Savova</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lingren</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal>. <article-title>Large-scale evaluation of automated clinical note de-identification and its impact on information extraction</article-title>. <source>Journal of the American Medical Informatics Association [Internet]</source> <year>2013</year> <month>Jan</month>. [cited 2023 2];<volume>20</volume>:<fpage>84</fpage>&#x2013;<lpage>94</lpage>. Available from: <pub-id pub-id-type="doi">https://doi.org/10.1136/amiajnl-2012-001012</pub-id>.</mixed-citation></ref>
<ref id="ref-77"><label>77</label><mixed-citation publication-type="website"><string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Dernoncourt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Szolovits</surname> <given-names>P</given-names></string-name>. <article-title>Transfer Learning for Named-Entity Recognition with Neural Networks [Internet]</article-title>. In: <source>Clinical Natural Language Processing</source>. <year>2018</year> [cited 2023 2]. Available from: <uri>https://aclanthology.org/L18-1708</uri>.</mixed-citation></ref>
<ref id="ref-78"><label>78</label><mixed-citation publication-type="journal"><string-name><surname>Moqurrab</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Ayub</surname> <given-names>U</given-names></string-name>, <string-name><surname>Anjum</surname> <given-names>A</given-names></string-name>, <string-name><surname>Asghar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Srivastava</surname> <given-names>G</given-names></string-name>. <article-title>An Accurate Deep Learning Model for Clinical Entity Recognition from Clinical Notes</article-title>. <source>Institute of Electrical and Electronics Engineers Journal of Biomedical and Health Informatics</source> <year>2021</year>;<fpage>25</fpage>. <pub-id pub-id-type="doi">https://doi.org/10.1109/JBHI.2021.3099755</pub-id></mixed-citation></ref>
<ref id="ref-79"><label>79</label><mixed-citation publication-type="journal"><string-name><surname>Dehghan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kovacevic</surname> <given-names>A</given-names></string-name>, <string-name><surname>Karystianis</surname> <given-names>G</given-names></string-name>, <string-name><surname>Keane</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Nenadic</surname> <given-names>G</given-names></string-name>. <article-title>Learning to identify Protected Health Information by integrating knowledge- and data-driven algorithms: A case study on psychiatric evaluation notes</article-title>. <source>J Biomed Inform</source> [Internet] <year>2017</year> <month>Nov</month>. [cited 2023 2];<volume>75S</volume>:<fpage>S28</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1016/j.jbi.2017.06.005</pub-id></mixed-citation></ref>
<ref id="ref-80"><label>80</label><mixed-citation publication-type="journal"><string-name><surname>Young Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dernoncourt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Uzuner</surname> <given-names>O</given-names></string-name>, <string-name><surname>Szolovits</surname> <given-names>P</given-names></string-name>, <string-name><surname>Albany</surname> <given-names>S</given-names></string-name>. <article-title>Feature-Augmented Neural Networks for Patient Note De-identification [Internet]</article-title>. In: <source>Clinical Natural Language Processing</source>. <year>2016</year> [cited 2023 2]. p. <fpage>17</fpage>&#x2013;<lpage>22</lpage>. Available from: <uri>https://doi.org/wbreak10.1016/j.jbi.2017.06.005</uri>.</mixed-citation></ref>
<ref id="ref-81"><label>81</label><mixed-citation publication-type="journal"><string-name><surname>Li</surname> <given-names>XB</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>J</given-names></string-name>. <article-title>Anonymizing and Sharing Medical Text Records</article-title>. <source>PMC- Author Manuscripts</source> [Internet] <year>2017</year> <month>Apr</month>. [cited 2023 2];<volume>28</volume>:<fpage>332</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">https://doi.org/10.1287/isre.2016.0676</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>