<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v6i1.1757</article-id>
<article-id pub-id-type="publisher-id">6:1:19</article-id>
<article-id pub-id-type="pii">S2399490821017572</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A scoping review of preprocessing methods for unstructured text data to assess data quality</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Nesca</surname><given-names initials="M">Marcello</given-names></name><xref ref-type="aff" rid="affil-1">1</xref></contrib>
<contrib contrib-type="author"><name><surname>Katz</surname><given-names initials="A">Alan</given-names></name><xref ref-type="aff" rid="affil-2">2</xref></contrib>
<contrib contrib-type="author"><name><surname>Leung</surname><given-names initials="CK">Carson K.</given-names></name><xref ref-type="aff" rid="affil-3">3</xref></contrib>
<contrib contrib-type="author"><name><surname>Lix</surname><given-names initials="LM">Lisa M.</given-names></name><xref ref-type="aff" rid="affil-4">4</xref><xref ref-type="corresp" rid="correspondingAurthor">*</xref></contrib>
<aff id="affil-1"><label>1</label><institution>Department of Community Health Sciences &#x0026; Manitoba Centre for Health Policy, University of Manitoba, Winnipeg, MB, Canada</institution></aff>
<aff id="affil-2"><label>2</label><institution>Department of Community Health Sciences, Department of Family Medicine, &#x0026; Manitoba Centre for Health Policy, University of Manitoba, Winnipeg, MB, Canada</institution></aff>
<aff id="affil-3"><label>3</label><institution>Department of Computer Science, University of Manitoba, Winnipeg, MB, Canada</institution></aff>
<aff id="affil-4"><label>4</label><institution>Department of Community Health Sciences, Manitoba Centre for Health Policy, &#x0026; George &#x0026; Fay Yee Centre for Healthcare Innovation, University of Manitoba, Winnipeg, MB, Canada</institution></aff>
</contrib-group>
<author-notes>
<corresp id="correspondingAurthor"><label>*</label>Corresponding author: Lisa M Lix <email>Lisa.Lix@umanitoba.ca</email>
</corresp>
<fn fn-type="conflict">
<label>Statement on conflicts of interest</label>
<p>The authors declare that they have no conflicts to report.</p>
</fn>
</author-notes>
<pub-date date-type="pub" publication-format="electronic"><day>04</day><month>10</month><year>2022</year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year>2022</year></pub-date>
<volume>7</volume>
<issue>1</issue>
<elocation-id>1757</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/1757">This article is available from the IJPDS website at: https://ijpds.org/article/view/1757</self-uri>
<abstract>
<title>Abstract</title>
<sec>
<title>Introduction</title>
<p>Unstructured text data (UTD) are increasingly found in many databases that were never intended to be used for research, including electronic medical record (EMR) databases. Data quality can impact the usefulness of UTD for research. UTD are typically prepared for analysis (i.e., preprocessed) and analyzed using natural language processing (NLP) techniques. Different NLP methods are used to preprocess UTD and may affect data quality.</p>
</sec>
<sec>
<title>Objective</title>
<p>Our objective was to systematically document current research and practices about NLP preprocessing methods to describe or improve the quality of UTD, including UTD found in EMR databases.</p>
</sec>
<sec>
<title>Methods</title>
<p>A scoping review was undertaken of peer-reviewed studies published between December 2002 and January 2021. Scopus, Web of Science, ProQuest, and EBSCOhost were searched for literature relevant to the study objective. Information extracted from the studies included article characteristics (i.e., year of publication, journal discipline), data characteristics, types of preprocessing methods, and data quality topics. Study data were presented using a narrative synthesis.</p>
</sec>
<sec>
<title>Results</title>
<p>A total of 41 articles were included in the scoping review; over 50% were published between 2016 and 2021. Almost 20% of the articles were published in health science journals. Common preprocessing methods included removal of extraneous text elements such as stop words, punctuation, and numbers, word tokenization, and parts of speech tagging. Data quality topics for articles about EMR data included misspelled words, security (i.e., de-identification), word variability, sources of noise, quality of annotations, and ambiguity of abbreviations.</p>
</sec>
<sec>
<title>Conclusions</title>
<p>Multiple NLP techniques have been proposed to preprocess UTD, with some differences in techniques applied to EMR data. There are similarities in the data quality dimensions used to characterize structured data and UTD. While a few general-purpose measures of data quality that do not require external data; most of these focus on the measurement of noise.</p>
</sec>
</abstract>
<kwd-group>
<kwd>review</kwd>
<kwd>data quality</kwd>
<kwd>natural language processing</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>Introduction</title>
<p>Routinely collected electronic health data are generated during the process of managing and monitoring the healthcare system [<xref ref-type="bibr" rid="ref-1">1</xref>, <xref ref-type="bibr" rid="ref-2">2</xref>]. Unstructured text data (UTD) are common in electronic medical records (EMRs), which is one type of routinely collected electronic health data. Further examples of UTD found in other types of routinely collected electronic health data, are laboratory testing results and clinical registry files.</p>
<p>The quality of data has been defined as their fitness for use [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>], that is, that data meets the needs of the user for a specific task or purpose, such as identifying individuals with a specific health condition. Given that routinely collected electronic health data are increasingly being used for research, it is important to consider their fitness for research, including epidemiologic studies or health services utilization studies. There are several consequences of poor data quality. For example, Kiefer noted that poor data quality can &#x201C;slow down innovation processes&#x201D; [<xref ref-type="bibr" rid="ref-6">6</xref>]. Poor data quality may also increase the time required to prepare a dataset for use, which can impact the timeliness of research outputs. Data quality is a multidimensional construct; it encompasses such dimensions as relevance, consistency, accuracy, comparability, timeliness, accessibility and usability [<xref ref-type="bibr" rid="ref-4">4</xref>, <xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-7">7</xref>, <xref ref-type="bibr" rid="ref-8">8</xref>]. Most data quality frameworks and assessment methods have been developed for structured data. However, Kiefer [<xref ref-type="bibr" rid="ref-6">6</xref>] argued that most, if not all, data quality dimensions developed for structured data are also relevant for UTD, although she emphasized the importance of relevance, interpretability, and accuracy when assessing UTD fitness for use. Kiefer also noted that there has been little research about data quality dimensions and indicators of these dimensions for UTD [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>The usability of UTD for research generally requires the application of natural language processing (NLP) techniques, including topic modeling, sentiment analysis, aspect mining (e.g., identifying different parts of speech), text summarization, and named entity recognition (e.g., identifying people, places, and other entities in unstructured data) [<xref ref-type="bibr" rid="ref-9">9</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>]. To prepare UTD for one or more of these NLP techniques, preprocessing of the data is an essential step. Preprocessing of UTD includes such actions as removing stop words (i.e., common words in a language), removing punctuation, tagging (i.e., identifying or labelling) parts of speech, and transforming abbreviations into words or phrases so that they can be easily interpreted. Accordingly, some researchers have suggested that indicators that measure the outputs of data preprocessing steps, such as the number or percent of abbreviations and the number or percent of spelling errors, could be used to characterize UTD quality. Some types of NLP, such as named entity recognition, involve training classification models to learn to identify data entities, such as parts of speech, diseases, or geographic locations [<xref ref-type="bibr" rid="ref-10">10</xref>]. This requires the use of annotated databases as a &#x201C;gold standard&#x201D;, which have been tagged (e.g., parts of speech have been documented or labelled) using manual or automated methods [<xref ref-type="bibr" rid="ref-10">10</xref>]. The quality of these annotated databases and the accuracy of classification models based on NLP applications involving annotated databases have also been proposed as indicators of UTD quality. Additionally, several unsupervised or supervised methods, which are used for both internal or external validation have been used for preprocessing and may be used to develop indicators for data quality. For example, in a study by Zennaki et al. [<xref ref-type="bibr" rid="ref-15">15</xref>], a recurrent neural network was used for unsupervised and semi-supervised parts of speech tagging in languages for which there are no labeled training data.</p>
<p>The data quality paradigm places a high value on the representation of truth from the perspective of the patient [<xref ref-type="bibr" rid="ref-16">16</xref>]. In a citizens&#x2019; jury study by Ford et al. [<xref ref-type="bibr" rid="ref-16">16</xref>], a representative sample of citizens listened to subject matter experts about the sharing of UTD within EMRs for research. The jury then deliberated to reach a conclusion from questions they were asked. With respect to data quality, the jurors noted that that text data may contain information about patients, judgments and offhand comments that may be misinterpreted by the researcher [<xref ref-type="bibr" rid="ref-16">16</xref>]. The concern with the representation of truth is a form of external validation (i.e., assessing data veracity). A study by Pantazos et al. [<xref ref-type="bibr" rid="ref-17">17</xref>], discussed the preservation of medical correctness, readability and consistency after EMR records were de-identified, to ensure data quality from a representativeness perspective. At the same time, NLP technologies still have difficulties with context and understanding language [<xref ref-type="bibr" rid="ref-18">18</xref>, <xref ref-type="bibr" rid="ref-19">19</xref>]. Preprocessing activities to prepare text data for research do not address the contextual concerns that the jurors raised. Thus, it should be noted that this study primarily focuses on assessing the goodness of fit of text data for analytics.</p>
<p>In summary, there are potentially many indicators that could be used to describe UTD quality, and these are primarily based on the use of NLP techniques to preprocess UTD. However, there is little relevant literature and few, if any, guidelines on the data quality indicators that might be recommended for inclusion in data quality frameworks for UTD, or that might be used to guide data preprocessing in studies that apply NLP methods to UTD. In addition, there have been few studies that have investigated the impact of UTD quality assessment on the performance of text analyses using NLP. The objective was to systematically document current research and practices about NLP preprocessing methods for UTD to describe or improve its quality.</p>
</sec>
<sec>
<title>Methods</title>
<p>To achieve the research objective, we undertook a scoping review of published literature about NLP preprocessing methods and data quality. The purpose of a scoping review is to map and describe the literature on a new topic or research area and identify key concepts, gaps in the research area, and types and sources of evidence to inform future research [<xref ref-type="bibr" rid="ref-20">20</xref>]. We adopted the Arksey and O&#x2019;Malley framework [<xref ref-type="bibr" rid="ref-20">20</xref>] for scoping reviews, which has the following steps: 1) define the research question, 2) identify relevant studies, 3) select studies, 4) chart the data, and 5) collate, summarize, and report results.</p>
<sec>
<title>Search strategy</title>
<p>The search strategy included the concepts of (1) data quality, (2) NLP, and (3) data preprocessing (see <xref ref-type="fig" rid="fig-1">Figure 1</xref>). The selection of search terms was informed by a systematic review on extracting text from EMRs to improve case detection [<xref ref-type="bibr" rid="ref-21">21</xref>], a scoping review about quality of routinely-collected electronic health data [<xref ref-type="bibr" rid="ref-22">22</xref>], and keywords related to preprocessing identified from an initial search of the literature [<xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>]. We consulted a librarian who assisted in developing and refining the list of search terms. Our initial literature review revealed few articles that included NLP, data quality, and preprocessing in the health science discipline, thus we expanded our search to include relevant literature in all disciplines.</p>
<fig id="fig-1"><label>Figure 1: Scoping review search strategy results</label>
<graphic xlink:href="ijpds-06-1757-g001.tif"/>
<attrib>Note: Numbers not in parentheses are for the search completed in April 2021; number in parentheses are for the search completed in May 2021.</attrib>
</fig>
<p>The review included empirical research articles and review articles and was conducted over two time periods. An initial search was executed with an unrestricted minimum date criterion, with an end date of April 15, 2020. We updated the search to the end of May 15, 2021. We searched Scopus, Web of Science, EBSCOhost and ProQuest. In addition, the reference sections of the selected articles were hand searched to identify additional, relevant articles.</p>
</sec>
<sec>
<title>Inclusion and exclusion criteria</title>
<p>An article was selected for inclusion if it met one or both of the following criteria: (1) it described research about preprocessing methods for UTD or UTD quality measures or methods, or it was a review article that discussed preprocessing methods to restructure or reorganize UTD for analysis; (2) it was about methods or processes to create a gold standard (or reference) dataset to validate UTD.</p>
<p>An article was excluded if it met one or more of the following criteria: (1) it was about methods for sentiment analysis, ontologies, semantic models, geo-spatial analysis, or qualitative research (e.g., methods to analyze interview or focus group data); (2) it was about methods to construct lexicon databases, dictionaries, or language databases; (3) it focused on the creation of software programs or proprietary solutions for text analysis; (4) it was not available in English; (5) it was an article from the ProQuest database that was neither an empirical article nor a scholarly article.</p>
</sec>
<sec>
<title>Article screening</title>
<p>Title and abstract screening were conducted for all articles identified through the implementation of the search strategy, after duplicates were removed. Training was undertaken first for the title, and abstract screening was completed by two authors on a 10% sample of all articles identified from the application of the initial search strategy, to ensure consistency when applying the inclusion and exclusion criteria. Percent agreement and its 95% confidence interval (CI) were calculated. After training was completed, all remaining articles were screened by one author. Rayyan [<xref ref-type="bibr" rid="ref-26">26</xref>], a web application for systematic and scoping reviews, was used to manage and organize articles through the process of title and abstract screening. Differences of opinion on the title and abstract screening were resolved by consensus. Full text screening of all articles selected after title and abstract screening was then conducted, to identify the articles to retain in the scoping review.</p>
</sec>
<sec>
<title>Data extraction and analysis</title>
<p>The following types of information were extracted from each article: (1) characteristics of the article, (2) characteristics of the text data, and (3) characteristics of preprocessing methods to restructure or reorganize the UTD. The systematic review conducted by Hinds et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] was used to inform the types of information extracted from the articles, such as the characteristics of articles and validation methods. The reviews conducted by Spasic et al. [<xref ref-type="bibr" rid="ref-6">6</xref>, <xref ref-type="bibr" rid="ref-28">28</xref>] and Kiefer [<xref ref-type="bibr" rid="ref-27">27</xref>] provided guidance on the characteristics of text data and the preprocessing methods that were included in the data extraction form [<xref ref-type="bibr" rid="ref-6">6</xref>, <xref ref-type="bibr" rid="ref-27">27</xref>, <xref ref-type="bibr" rid="ref-28">28</xref>]. We extracted information about the specific dimensions (i.e., types) of data quality that were mentioned in each of the articles. These dimensions were identified from existing data quality frameworks, such as those developed by the Manitoba Centre for Health Policy [<xref ref-type="bibr" rid="ref-7">7</xref>]. Lastly, an inspection of the topics for data quality for UTD that used EMRs were explored. Information extracted included: (1) methods, (2) strengths and limitations, and (3) use cases.</p>
<p>Data extraction training was completed by two authors on a 10% sample of the articles selected for full data extraction. A data dictionary was created to ensure consistency of the data extraction methods. Percent agreement and its 95% CI were calculated. Variables with open-ended responses were excluded from this calculation. Any differences in agreement were resolved by consensus.</p>
</sec>
</sec>
<sec>
<title>Results</title>
<p>The initial search (<xref ref-type="fig" rid="fig-1">Figure 1</xref>, numbers not in parentheses), which encompassed the period up to April 15, 2020, identified a total of 1134 articles. The updated search (<xref ref-type="fig" rid="fig-1">Figure 1</xref>, numbers in parentheses) identified an additional 154 articles. Thus, in total 1288 articles were retrieved from the search before duplicates were removed, and 1226 remained after duplicates were removed (<xref ref-type="fig" rid="fig-2">Figure 2</xref>). Initial training for title and abstract screening yielded 83.3% (95% CI: 75.2%, 89.2%) agreement. After full text screening, a total of 41 articles remained for data extraction (<xref ref-type="fig" rid="fig-2">Figure 2</xref>).</p>
<fig id="fig-2"><label>Figure 2: Scoping review PRISMA flowchart</label>
<graphic xlink:href="ijpds-06-1757-g002.tif"/>
<attrib>Note: Numbers not in parentheses are for the search completed in April 2021; number in parentheses are for the search completed in May 2021.</attrib>
</fig>
<sec>
<title>Data extraction results</title>
<p>Ten percent of articles from the initial search were selected for the calculation of agreement for full text extraction; the overall agreement was 94.6% (95% CI: 88.9%, 97.4%). <xref ref-type="table" rid="table-1">Table 1</xref> summarizes the characteristics of the articles selected for full data extraction.</p>
<table-wrap id="table-1">
<label>Table 1: Characteristics of the scoping review articles (<italic>n</italic> = 41)</label>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="middle" align="left"></th>
<th valign="middle" align="center"><bold>n</bold></th>
<th valign="middle" align="center">%</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Article type</bold></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Empirical research</td>
<td valign="middle" align="center">37</td>
<td valign="middle" align="center">90.2</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Review article</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">7.3</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Case study</td>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">2.4</td>
</tr>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Publication year</bold></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">2002&#x2013;2009</td>
<td valign="middle" align="center">7</td>
<td valign="middle" align="center">17.1</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">2010&#x2013;2015</td>
<td valign="middle" align="center">13</td>
<td valign="middle" align="center">31.7</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">2016&#x2013;2021</td>
<td valign="middle" align="center">21</td>
<td valign="middle" align="center">51.2</td>
</tr>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Disciplinary area</bold></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Computer Science and Engineering</td>
<td valign="middle" align="center">25</td>
<td valign="middle" align="center">61.0</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Health Sciences</td>
<td valign="middle" align="center">8</td>
<td valign="middle" align="center">19.5</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Social Sciences, Humanities, Business</td>
<td valign="middle" align="center">8</td>
<td valign="middle" align="center">19.5</td>
</tr>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Type of data</bold><sup>*</sup></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Electronic Medical Records</td>
<td valign="middle" align="center">8</td>
<td valign="middle" align="center">19.5</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Lexical (e.g., language treebanks)</td>
<td valign="middle" align="center">8</td>
<td valign="middle" align="center">19.5</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Organizational Documents</td>
<td valign="middle" align="center">6</td>
<td valign="middle" align="center">14.6</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Product Reviews</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">9.8</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">News Articles</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">9.8</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Corpora (e.g., biomedical text of gene entities)</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">9.8</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Abstracts and Articles from Scientific Journals</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">9.8</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Social Media</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">7.3</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Administrative</td>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">2.4</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Other</td>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">12.2</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><sup>*</sup>Categories are not mutually exclusive.</p>
</table-wrap-foot>
</table-wrap>
<p>In total, 90.2% of the articles reported the results of empirical research and another 7.3% (<italic>n</italic> = 3) were review articles. Only one article was classified as a case study. More than half (51.2%) of the articles were published between 2016 and 2021. No articles were published prior to 2002. In terms of disciplinary area, 61% of articles were deemed to be from computer science and engineering disciplines, while 39% were from the health sciences, social sciences, humanities, and business.</p>
<p>A variety of types of text data were represented in the selected articles including EMRs (i.e., clinical notes, progress notes, patient safety records [<xref ref-type="bibr" rid="ref-17">17</xref>, <xref ref-type="bibr" rid="ref-30">30</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>]), lexical documents (i.e., language treebanks which are bodies of text that have been parsed semantically and syntactically, WordNet database [<xref ref-type="bibr" rid="ref-37">37</xref>&#x2013;<xref ref-type="bibr" rid="ref-43">43</xref>]), organizational documents (i.e., maintenance logs/data, accident reports, requirements documentation [<xref ref-type="bibr" rid="ref-44">44</xref>&#x2013;<xref ref-type="bibr" rid="ref-47">47</xref>]), abstracts and scientific articles (i.e., PubMed and various engineering journals [<xref ref-type="bibr" rid="ref-29">29</xref>, <xref ref-type="bibr" rid="ref-48">48</xref>&#x2013;<xref ref-type="bibr" rid="ref-50">50</xref>]), various bodies of text (corpora) (i.e., non-language corpora, non-medical/medical/biomedical corpora, language corpus [<xref ref-type="bibr" rid="ref-50">50</xref>&#x2013;<xref ref-type="bibr" rid="ref-53">53</xref>]), social media data (i.e., Twitter, meme tracker from various social media websites [<xref ref-type="bibr" rid="ref-54">54</xref>&#x2013;<xref ref-type="bibr" rid="ref-56">56</xref>]), product reviews (i.e., general product, Chinese tourism, Amazon product [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-57">57</xref>, <xref ref-type="bibr" rid="ref-58">58</xref>]), and news articles (i.e., magazines, newswires, consumer reports [<xref ref-type="bibr" rid="ref-54">54</xref>, <xref ref-type="bibr" rid="ref-59">59</xref>, <xref ref-type="bibr" rid="ref-60">60</xref>]).</p>
<p>Almost all empirical articles (85.4%) described preprocessing methods to improve NLP algorithm performance. However, one article [<xref ref-type="bibr" rid="ref-55">55</xref>] offered an empirical approach to compare a new preprocessing methodology to an existing (i.e., baseline) preprocessing method on short text similarity measures; results showed that the newly-proposed method outperformed the existing baseline method [<xref ref-type="bibr" rid="ref-55">55</xref>]. Other empirical articles discussed methods for manual or automated annotation. For example, Westpfahl et al. [<xref ref-type="bibr" rid="ref-53">53</xref>] discussed the creation of a gold standard corpus for teaching and research about spoken German. In this article, the corpus was manually annotated with part of speech tagging (i.e., identifying and annotating words that are nouns, adverbs, verbs, and other parts of speech) and lemmatization, a form of word stemming, where the morphological base of a word is returned [<xref ref-type="bibr" rid="ref-53">53</xref>].</p>
<p>In terms of data size and terms used to describe data size, the UTD described in the articles were characterized in many ways (<xref ref-type="table" rid="table-2">Table 2</xref>). Documents and words were mentioned most frequently (e.g., citing how many documents were used or how many words were found in a type of UTD). More than half (53.7%) of the articles mentioned the volume of documents, and 46.3% of the articles mentioned the count of words in each document. Elements of size often used to describe the size of structured data such as &#x201C;how many rows (records)&#x201D; or &#x201C;how many columns (features)&#x201D; were among the least common ways to describe UTD size.</p>
<table-wrap id="table-2">
<label>Table 2: Characteristics of text data and quality assessment in the scoping review articles (<italic>n</italic> = 41)</label>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="middle" align="left"></th>
<th valign="middle" align="center"><bold>n</bold></th>
<th valign="middle" align="center">%</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Measures used to describe data size</bold><sup>*</sup></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Documents</td>
<td valign="middle" align="center">22</td>
<td valign="middle" align="center">53.7</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Words</td>
<td valign="middle" align="center">19</td>
<td valign="middle" align="center">46.3</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Phrases (sentences)</td>
<td valign="middle" align="center">8</td>
<td valign="middle" align="center">19.5</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Rows (records)</td>
<td valign="middle" align="center">6</td>
<td valign="middle" align="center">14.6</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Features</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">9.8</td>
</tr>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Data annotation<sup>*</sup></bold></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Manual</td>
<td valign="middle" align="center">20</td>
<td valign="middle" align="center">48.8</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Automated</td>
<td valign="middle" align="center">17</td>
<td valign="middle" align="center">41.5</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">No annotation</td>
<td valign="middle" align="center">15</td>
<td valign="middle" align="center">36.6</td>
</tr>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Preprocessing methods<sup>*</sup></bold></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Reorganizing or restructuring methods</td>
<td valign="middle" align="center">35</td>
<td valign="middle" align="center">85.4</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Internal validation</td>
<td valign="middle" align="center">23</td>
<td valign="middle" align="center">56.1</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">External validation</td>
<td valign="middle" align="center">11</td>
<td valign="middle" align="center">26.8</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Other</td>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">9.8</td>
</tr>
<tr>
<td valign="middle" align="left" colspan="3"><bold>Data quality dimensions mentioned<sup>*</sup></bold></td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Accuracy</td>
<td valign="middle" align="center">28</td>
<td valign="middle" align="center">68.3</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Relevance</td>
<td valign="middle" align="center">14</td>
<td valign="middle" align="center">34.1</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Comparability</td>
<td valign="middle" align="center">13</td>
<td valign="middle" align="center">31.7</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Usability</td>
<td valign="middle" align="center">7</td>
<td valign="middle" align="center">17.1</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Completeness</td>
<td valign="middle" align="center">7</td>
<td valign="middle" align="center">17.1</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Validity</td>
<td valign="middle" align="center">7</td>
<td valign="middle" align="center">17.1</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Readability</td>
<td valign="middle" align="center">6</td>
<td valign="middle" align="center">14.6</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Accessibility</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">7.3</td>
</tr>
<tr>
<td valign="middle" align="left" style="padding-left:1em">Timeliness</td>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">7.3</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><sup>*</sup>Categories are not mutually exclusive.</p>
</table-wrap-foot>
</table-wrap>
<p>Overall, over one third (36.6%) of the articles did not employ text annotation, while 48.8% of the articles discussed manual annotation (e.g., an experienced coder assigning codes using the International Classification of Diseases - Clinical Modification (ICD-9-CM) [<xref ref-type="bibr" rid="ref-61">61</xref>]), and 41.5% of articles discussed automated annotation (e.g., words were automatically annotated for part of speech analyses, grammar tagging, and assistance with manual annotation processes [<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-29">29</xref>, <xref ref-type="bibr" rid="ref-37">37</xref>, <xref ref-type="bibr" rid="ref-39">39</xref>, <xref ref-type="bibr" rid="ref-52">52</xref>, <xref ref-type="bibr" rid="ref-53">53</xref>, <xref ref-type="bibr" rid="ref-55">55</xref>, <xref ref-type="bibr" rid="ref-62">62</xref>]).</p>
<p>Almost all (85.4%) of the articles discussed using restructuring and reorganizing methods to prepare UTD for analysis (e.g., removing stop words, punctuation, removing URLs). <xref ref-type="fig" rid="fig-3">Figure 3</xref> describes the types of restructuring and reorganizing methods that were used for all articles and for the subset of articles from health science disciplines.</p>
<fig id="fig-3"><label>Figure 3: Comparison of restructuring and reorganizing methodsfor all articles and for health science articles</label>
<graphic xlink:href="ijpds-06-1757-g003.tif"/>
</fig>
<p>We identified the articles that explicitly mentioned a data quality dimension (i.e., we counted the words pertaining to data quality). There was no prior determination of the quality dimension criteria when capturing words in relation to data quality. The three data quality dimensions that were most frequently mentioned were: accuracy (68.3%), relevance (34.1%), and comparability (31.7%). Furthermore, &#x201C;data quality&#x201D; or &#x201C;quality&#x201D; as terms were described or referenced in several ways among the 41 articles. Several articles discussed quality either from the perspective of data quality (or information quality), or using terminology from data or information quality dimensions (e.g., accuracy, correctness, interpretability) [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-17">17</xref>, <xref ref-type="bibr" rid="ref-47">47</xref>, <xref ref-type="bibr" rid="ref-56">56</xref>, <xref ref-type="bibr" rid="ref-58">58</xref>]. Other articles discussed enhancing data quality by focusing on utilizing or improving preprocessing methods [<xref ref-type="bibr" rid="ref-31">31</xref>, <xref ref-type="bibr" rid="ref-34">34</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>, <xref ref-type="bibr" rid="ref-37">37</xref>, <xref ref-type="bibr" rid="ref-40">40</xref>, <xref ref-type="bibr" rid="ref-42">42</xref>, <xref ref-type="bibr" rid="ref-46">46</xref>, <xref ref-type="bibr" rid="ref-50">50</xref>, <xref ref-type="bibr" rid="ref-52">52</xref>, <xref ref-type="bibr" rid="ref-54">54</xref>&#x2013;<xref ref-type="bibr" rid="ref-56">56</xref>, <xref ref-type="bibr" rid="ref-63">63</xref>, <xref ref-type="bibr" rid="ref-64">64</xref>]. Several of these articles only mentioned &#x201C;data quality&#x201D; in passing; the main focus of these articles were the preprocessing methods. Other articles discussed quality from an &#x201C;annotation&#x201D; perspective whereby tagging documents or words was intended to enhance the quality of algorithms either through the creation of treebanks or manual annotation using expert opinion [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-30">30</xref>, <xref ref-type="bibr" rid="ref-32">32</xref>, <xref ref-type="bibr" rid="ref-33">33</xref>, <xref ref-type="bibr" rid="ref-38">38</xref>, <xref ref-type="bibr" rid="ref-39">39</xref>, <xref ref-type="bibr" rid="ref-47">47</xref>&#x2013;<xref ref-type="bibr" rid="ref-49">49</xref>, <xref ref-type="bibr" rid="ref-51">51</xref>, <xref ref-type="bibr" rid="ref-53">53</xref>, <xref ref-type="bibr" rid="ref-59">59</xref>, <xref ref-type="bibr" rid="ref-65">65</xref>, <xref ref-type="bibr" rid="ref-66">66</xref>]. Lastly, some articles discussed quality through utilizing preprocessing methods that were primarily focused on achieving a more &#x201C;accurate or relevant&#x201D; outcome for algorithmic model performance [<xref ref-type="bibr" rid="ref-29">29</xref>, <xref ref-type="bibr" rid="ref-41">41</xref>, <xref ref-type="bibr" rid="ref-43">43</xref>&#x2013;<xref ref-type="bibr" rid="ref-45">45</xref>, <xref ref-type="bibr" rid="ref-57">57</xref>, <xref ref-type="bibr" rid="ref-58">58</xref>, <xref ref-type="bibr" rid="ref-60">60</xref>, <xref ref-type="bibr" rid="ref-62">62</xref>, <xref ref-type="bibr" rid="ref-67">67</xref>, <xref ref-type="bibr" rid="ref-68">68</xref>]. Overall, while articles mentioned terminology such as &#x201C;accuracy&#x201D; or &#x201C;relevance&#x201D;, data quality as a concept (or its individual dimensions), was referenced or described in multiple ways. Articles that focused exclusively on data quality dimensions or their measurement were rare.</p>
<p>Overall, many of the articles described an improvement of algorithm outcome performance through a variety of preprocessing techniques to enhance UTD quality. However, there were no articles that reported on indicators of UTD quality before preprocessing. While data quality dimensions were not mapped to specific preprocessing methods, &#x201C;accuracy&#x201D; was used to evaluate outcomes in a confusion matrix, and it was also used as a descriptor for utilising preprocessing methods to enable unsupervised and supervised algorithms outcomes to be more accurate. That is, all preprocessing methods were closely tied to the data quality dimension of accuracy.</p>
</sec>
<sec>
<title>UTD quality topics for EMR data</title>
<p>Seven data quality topics were discussed specifically in reference to EMR data. The topic areas were: (1) misspelled words, (2) security/de-identification, (3) reducing word variability, (4) sources of noise, (5) quality of annotations, (6) ambiguous abbreviations, and (7) reducing manual annotations. Further details such as the methods used, strengths and limitations of the methods, and the data used (i.e., use case) are provided in <xref ref-type="supplementary-material" rid="sup-a">Supplementary Appendix 1</xref>.</p>
<p>Several of these quality topics focused on the reduction of possible error or variability (e.g., correct misspelled words, reduction of word variability, addressing sources of noise, or addressing ambiguity in abbreviations). Two articles reported on the assessment of misspelled words. Some of the methods utilized were a combination of rule-based approaches or machine learning approaches such as dictionaries, regular expressions, a string-to-string edit distance known as the Levenshtein-Damerau distance, and word sense disambiguation for words with similar parts of speech (e.g., distinguishing between two similar words that are classified as verbs). One article by Assale et al. [<xref ref-type="bibr" rid="ref-31">31</xref>], was about reducing word variability in typographical corrections, where typographical errors in one word can create variability in how the word is spelled. Rule-based approaches were used in this article such as preprocessing methods to restructure and reorganize UTD (e.g., removal of stop words), counting word frequencies, and utilizing the Levenshtein-Damerau distance metric. One article addressed sources of noise where the removal of noise in text involves preprocessing methods that restructure and reorganize UTD. Some of these methods include (but are not limited to): tokenization, converting uppercase letters to lowercase letters, and removal of stop words. Lastly, ambiguity in abbreviations was addressed in one article where the authors used deep learning methods (i.e., convolutional neural networks); no feature engineering or preprocessing was required. The convolutional neural network was trained on word embeddings, which are representations of words in a list. These word embeddings were extracted from journal articles found in PubMed.</p>
<p>Other topics focused on the evaluation or reduction of manual efforts such as assessing the quality of annotations or solutions involving reducing manual annotation of text data. Two articles focused on the quality of annotations. One focused on manual annotation of clinical notes for fall-related injuries. In the second article, tags for parts of speech and named entities were applied by annotators with backgrounds in computational linguistics, while physicians in training solved any disagreements between annotators.</p>
</sec>
</sec>
<sec>
<title>Discussion</title>
<p>The scoping review has documented practices for preprocessing UTD to describe or improve its quality. Few articles in the scoping review discussed the quality of UTD before analysis or preprocessing was initiated. The main topics raised in the selected studies were about the challenges of defining data quality, the choice of data quality assessments for UTD and how this is influenced by the context of the text data, and differences in the data quality challenges associated with EMR UTD when compared to other types of UTD.</p>
<p>The scoping review reveals the following key points: 1) The most common preprocessing methods used in health science articles were different from the most common preprocessing methods used in all disciplines combined. To elaborate, the most common preprocessing methods for health science articles included removal of stop words, removal of punctuation, removal of numbers, tokenization, parts of speech tagging, and converting characters to lowercase. 2) Few dimensions of data quality were considered in assessed UTD. Accuracy, relevance, and comparability were the most commonly-reported dimensions. 3) Quality indicator topics addressed potential challenges in preprocessing, such as the quality of annotations, presence of spelling errors, and presence of ambiguous abbreviations.</p>
<p>One difficulty with describing the quality of UTD is the lack of standardized terminology. Strong et al. [<xref ref-type="bibr" rid="ref-69">69</xref>] differentiated &#x201C;information quality&#x201D; from &#x201C;data quality&#x201D; in terms of its specific goals; information quality is about assessing the needs of information users while data quality refers to the fitness of data for its intended use. However, Chen and Tseng [<xref ref-type="bibr" rid="ref-58">58</xref>] did not make that distinction; their indicators of information quality are similar to those found in existing data quality frameworks.</p>
<p>Amongst the data quality measures identified from the scoping review, some were tailored specifically for the data being assessed. This emphasizes that quality of UTD depends on the context or type of data that is being assessed. Language usage must be contextual to the environment (i.e., vernacular used in product reviews differ from vernacular used in EMR notes). For example, several of the data quality measures for UTD that Chen and Tseng&#x2019;s article reference do not necessarily apply to data in EMRs such as the data quality dimension for &#x201C;objectivity&#x201D;, and assessments of whether a product review is an opinion rather than factual [<xref ref-type="bibr" rid="ref-58">58</xref>].</p>
<p>Some of the challenges with quality of UTD in EMRs are different than the challenges associated with UTD found in social media data or organizational reports. One of the biggest challenges of the former is ensuring privacy, anonymity, and confidentiality of patient data. Pantazos et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] stated that as UTD in EMRs increase in usage, it must achieve two goals (1) that it is de-identified and anonymized and (2) EMRs that are de-identified must contain accurate information about the un-identified patient and be coherent to the reader. The scoping review revealed that high-quality UTD within EMRs must be readable, correct, and consistent with a patient&#x2019;s record.</p>
<p>This research has some limitations. The scoping review was restricted to English language articles and grey literature was not searched. Accordingly, the scoping review may not represent all published articles about UTD quality. Furthermore, articles that discussed preprocessing methods to improve algorithmic modelling outcomes were not included. It was not feasible to address all different modelling techniques for text data (i.e., collect all articles that conducted a sentiment analysis or other methods). The articles selected were those that focused on quality of text data, or that focused on preprocessing methods and also mentioned data quality. Choosing key words for the scoping review search strategy was challenging. This was due in large part to the lack of standardized terminology and the diverse terminology within the NLP and data quality literatures and may have resulted in some articles being missed.</p>
<p>Despite these limitations, the major strength of this research is that it used a systematic approach to examine preprocessing methods to describe data quality for UTD. Data quality for UTD is an important area for research in multiple fields, including in the health sciences. Written language in EMRs and other patient-related documents is different from other types of text data found in textbooks or social media due to its nature in short text and point form.</p>
<p>Several opportunities for future research exist. First, a scoping review could be conducted to identify operational definitions for data quality dimensions specific to text data from routinely collected electronic health data. Akin to Weiskopf et al.&#x2019;s scoping review to discover methods and dimensions [<xref ref-type="bibr" rid="ref-70">70</xref>], the scoping review could be used to develop definitions for dimensions of data quality for UTD. This research has shown that there are unique characteristics of text data that are not present in numerical data (e.g., grammar rules or punctuation). Thus, operational data quality definitions that specifically address text data properties is an important step towards structuring a quality framework for text data. Second, a documentation project that maps operational definitions for data quality dimensions to the data quality indicators for routinely collected electronic health data should be explored [<xref ref-type="bibr" rid="ref-70">70</xref>]. Third, a case study could be undertaken for several text data quality indicator topics identified from this research (e.g., disambiguation of abbreviations). While there have been studies to disambiguate abbreviations [<xref ref-type="bibr" rid="ref-34">34</xref>] and correct spelling/typographical errors [<xref ref-type="bibr" rid="ref-31">31</xref>, <xref ref-type="bibr" rid="ref-35">35</xref>], it would also be beneficial to identify the impact of word variability on algorithm outputs in text data. Lastly, additional scoping or systematic reviews could be conducted. In particular, quality indicators for short text documents such as social media posts, reviews, and EMRs, could be explored [<xref ref-type="bibr" rid="ref-13">13</xref>, <xref ref-type="bibr" rid="ref-54">54</xref>&#x2013;<xref ref-type="bibr" rid="ref-56">56</xref>, <xref ref-type="bibr" rid="ref-58">58</xref>, <xref ref-type="bibr" rid="ref-71">71</xref>]. Since EMRs are characterized by short texts, it would be interesting to examine other text quality indicators appropriate for these types of data.</p>
</sec>
<sec>
<title>Conclusion</title>
<p>Data quality is a multidimensional construct that depends on the context in which the data will be used. However, there are many similarities between the dimensions of data quality for structured and unstructured text and methods to assess data quality. Assessing data quality in UTD often requires access to specialized gold standard datasets or dictionaries. However, there are a few general-purpose measures of data quality that do not require external data; most of these focus on the measurement of noise in the data.</p>
</sec>
<sec sec-type="supplementary-material">
<title>Supplementary Files</title>
<supplementary-material id="sup-a">
<label>Supplementary Table 1</label> 
<media mimetype="application" mime-subtype="pdf" xlink:href="ijpds-06-1757-s001.pdf"/>
</supplementary-material>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>MN was supported by the Visual and Automated Disease Analytics program during his research program. LML is supported by a Tier I Canada Research Chair in Methods for Electronic Health Data Quality. This research was supported by the Canada Research Chairs program.</p>
</ack>
<sec>
<title>Ethics statement</title>
<p>Ethical approval was not required because this research did not use patient data.</p>
</sec>
<fn-group>
<fn fn-type="other">
<label>Supplementary appendices</label>
<p>Supplementary Appendix 1 present data quality topics for UTD found in articles that used EMR data. The following information was collected: 1) the quality topic, 2) author of publication, 3) the methods used, 4) the strengths and limitations of the method, and 5) the use case that describes what data the authors used.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="ref-1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Benchimol</surname> <given-names>EI</given-names></string-name>, <string-name><surname>Smeeth</surname> <given-names>L</given-names></string-name>, <string-name><surname>Guttmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Moher</surname> <given-names>D</given-names></string-name>, <string-name><surname>Petersen</surname> <given-names>I</given-names></string-name>, <string-name><surname>S&#x00F8;rensen</surname> <given-names>HT</given-names></string-name>, <string-name><surname>von Elm</surname> <given-names>E</given-names></string-name>, <string-name><surname>Langan</surname> <given-names>SM</given-names></string-name>. <article-title>The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) statement</article-title>. <source>PLoS Med [Internet]</source>. <year>2015</year>;<volume>12</volume>(<issue>10</issue>). Available from: <uri>https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4595218/</uri>.</mixed-citation></ref>
<ref id="ref-2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Nicholls</surname> <given-names>SG</given-names></string-name>, <string-name><surname>Langan</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Benchimol</surname> <given-names>EI</given-names></string-name>. <article-title>Routinely collected data: the importance of high-quality diagnostic coding to research</article-title>. <source>Can Med Assoc J [Internet]</source>. <year>2017</year>;<volume>189</volume>(<issue>33</issue>):<fpage>E1054</fpage>&#x2013;<lpage>5</lpage>. Available from: <uri>https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5566604/</uri>.</mixed-citation></ref>
<ref id="ref-3"><label>3</label><mixed-citation publication-type="journal"><string-name><surname>Smith</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lix</surname> <given-names>LM</given-names></string-name>, <string-name><surname>Azimaee</surname> <given-names>M</given-names></string-name>, <string-name><surname>Enns</surname> <given-names>JE</given-names></string-name>, <string-name><surname>Orr</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Roos</surname> <given-names>LL</given-names></string-name>. <article-title>Assessing the quality of administrative data for research: a framework from the Manitoba Centre for Health Policy</article-title>. <source>J Am Med Inform Assoc [Internet]</source>. <year>2018</year> [cited <year>2018</year> <month>Nov</month> <day>18</day>];<volume>25</volume>(<issue>3</issue>):<fpage>224</fpage>&#x2013;<lpage>9</lpage>. Available from: <uri>http://academic.oup.com/jamia/article/25/3/224/4102320</uri>.</mixed-citation></ref>
<ref id="ref-4"><label>4</label><mixed-citation publication-type="book"><string-name><surname>Agboola</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Camara</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Brown</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kendall</surname> <given-names>O</given-names></string-name>. <chapter-title>The PHAC Data Quality Framework: A useful tool to serve surveillance programs</chapter-title>. In <publisher-loc>Canada</publisher-loc>: <publisher-name>Public Health Agency of Canada</publisher-name>; <year>2011</year>. Available from: <uri>https://www.researchgate.net/publication/305469008_The_PHAC_Data_Quality_Framework_A_Useful_Tool_to_Serve_Surveillance_Programs</uri>.</mixed-citation></ref>
<ref id="ref-5"><label>5</label><mixed-citation publication-type="website"><collab>Canadian Institute for Health Information</collab>. <article-title>CIHI&#x2019;s information quality framework [Internet]</article-title>. <source>Canadian Institute for Health Information</source>; <year>2017</year>. Available from: <uri>https://www.cihi.ca/sites/default/files/document/iqf-summary-july-26-2017-en-web_0.pdf</uri>.</mixed-citation></ref>
<ref id="ref-6"><label>6</label><mixed-citation publication-type="website"><string-name><surname>Kiefer</surname> <given-names>C</given-names></string-name>. <chapter-title>Assessing the quality of unstructured data: An initial overview</chapter-title>. In: <source>LWDA [Internet]</source>. <year>2016</year>. p. <fpage>62</fpage>&#x2013;<lpage>73</lpage>. Available from: <uri>https://www.semanticscholar.org/paper/Assessing-the-Quality-of-Unstructured-Data%3A-An-Kiefer/a7d31b09498a16201f7044eb77ff14d27c1c559b</uri>.</mixed-citation></ref>
<ref id="ref-7"><label>7</label><mixed-citation publication-type="website"><string-name><surname>Azimaee</surname> <given-names>M</given-names></string-name>, <string-name><surname>Smith</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lix</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ostapyk</surname> <given-names>T</given-names></string-name>, <string-name><surname>Burchill</surname> <given-names>C</given-names></string-name>, <string-name><surname>Orr</surname> <given-names>J</given-names></string-name>. <article-title>MCHP data quality framework [Internet]</article-title>. <source>Winnipeg Manitoba Canada University of Manitoba MCHP</source>; <year>2018</year>. Available from: <uri>http://umanitoba.ca/faculties/health_sciences/medicine/units/chs/departmental_units/mchp/protocol/media/Data_Quality_Framework.pdf</uri>.</mixed-citation></ref>
<ref id="ref-8"><label>8</label><mixed-citation publication-type="website"><collab>Government of Canada SC</collab>. <article-title>Statistics Canada&#x2019;s quality assurance framework, 2017 [Internet]</article-title>. <source>Statistics Canada</source>; <year>2017</year>. Available from: <uri>https://www150.statcan.gc.ca/n1/pub/12-586-x/12-586-x2017001-eng.pdf</uri>.</mixed-citation></ref>
<ref id="ref-9"><label>9</label><mixed-citation publication-type="website"><string-name><surname>Sonntag</surname> <given-names>D</given-names></string-name>. <chapter-title>Assessing the quality of natural language text data</chapter-title>. In <source>Gesellschaft f&#x00FC;r Informatik e.V.</source>; <year>2004</year>. p. <fpage>259</fpage>&#x2013;<lpage>63</lpage>. Available from: <uri>https://dl.gi.de/handle/20.500.12116/28866</uri>.</mixed-citation></ref>
<ref id="ref-10"><label>10</label><mixed-citation publication-type="website"><string-name><surname>Ritter</surname> <given-names>A</given-names></string-name>, <string-name><surname>Clark</surname> <given-names>S</given-names></string-name>, <string-name><surname>Etzioni</surname> <given-names>O</given-names></string-name>. <chapter-title>Named entity recognition in tweets: an experimental study</chapter-title>. In: <source>Proceedings of the 2011 conference on empirical methods in natural language processing [Internet]</source>. <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2011</year>. p. <fpage>1524</fpage>&#x2013;<lpage>34</lpage>. Available from: <uri>https://aclanthology.org/D11-1141</uri>.</mixed-citation></ref>
<ref id="ref-11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Kee</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jieyi Tang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chuang</surname> <given-names>KL</given-names></string-name>. <article-title>Scoping review of mindfulness research: a topic modelling approach</article-title>. <source>Mindfulness [Internet]</source>. <year>2019</year>;<volume>10</volume>:<fpage>1474</fpage>&#x2013;<lpage>88</lpage>. Available from: <uri>https://link-springer-com.uml.idm.oclc.org/article/10.1007%2Fs12671-019-01136-4</uri>.</mixed-citation></ref>
<ref id="ref-12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Briesch</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hobbs</surname> <given-names>R</given-names></string-name>, <string-name><surname>Jaja</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kjersten</surname> <given-names>B</given-names></string-name>, <string-name><surname>Voss</surname> <given-names>C</given-names></string-name>. <article-title>Training and evaluating a statistical part of speech tagger for natural language applications using Kepler Workflows</article-title>. <source>Procedia Comput Sci [Internet]</source>. <year>2012</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>9</volume>:<fpage>1588</fpage>&#x2013;<lpage>94</lpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S1877050912002955</uri>.</mixed-citation></ref>
<ref id="ref-13"><label>13</label><mixed-citation publication-type="journal"><string-name><surname>Yang</surname> <given-names>HL</given-names></string-name>, <string-name><surname>Chao</surname> <given-names>AFY</given-names></string-name>. <article-title>Sentiment annotations for reviews: an information quality perspective</article-title>. <source>Online Inf Rev [Internet]</source>. <year>2018</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>42</volume>(<issue>5</issue>):<fpage>579</fpage>&#x2013;<lpage>94</lpage>. Available from: <pub-id pub-id-type="doi">10.1108/OIR-04-2017-0114.</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14</label><mixed-citation publication-type="book"><string-name><surname>Subha</surname> <given-names>R</given-names></string-name>, <string-name><surname>Palaniswami</surname> <given-names>S</given-names></string-name>. <chapter-title>Quality factor assessment and text summarization of unambiguous natural language requirements</chapter-title>. In: <string-name><surname>Unnikrishnan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Surve</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bhoir</surname> <given-names>D</given-names></string-name>, editors. <source>Advances in Computing, Communication, and Control [Internet]</source>. <publisher-name>Berlin, Heidelberg</publisher-name>: <publisher-loc>Springer</publisher-loc>; <year>2013</year>. p. <fpage>131</fpage>&#x2013;<lpage>46</lpage>. Available from: <uri>https://doi-org.uml.idm.oclc.org/10.1007/978-3-642-36321-4_12</uri>.</mixed-citation></ref>
<ref id="ref-15"><label>15</label><mixed-citation publication-type="book"><string-name><surname>Zennaki</surname> <given-names>O</given-names></string-name>, <string-name><surname>Semmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Besacier</surname> <given-names>L</given-names></string-name>. <chapter-title>Unsupervised and lightly supervised part-of-speech tagging using recurrent neural networks</chapter-title>. In: <source>29th Pacific Asia Conference on Language, Information and Computation (PACLIC) [Internet]</source>. <publisher-loc>Shangai, China</publisher-loc>; <year>2015</year> [cited <year>2021</year> <month>Jul</month> <day>18</day>]. Available from: <uri>https://hal.archives-ouvertes.fr/hal-01350113</uri>.</mixed-citation></ref>
<ref id="ref-16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>Ford</surname> <given-names>E</given-names></string-name>, <string-name><surname>Oswald</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hassan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bozentko</surname> <given-names>K</given-names></string-name>, <string-name><surname>Nenadic</surname> <given-names>G</given-names></string-name>, <string-name><surname>Cassell</surname> <given-names>J</given-names></string-name>. <article-title>Should free-text data in electronic medical records be shared for research? A citizens&#x2019; jury study in the UK</article-title>. <source>Journal of Medical Ethics [Internet]</source>. <year>2020</year> <month>Jun</month> <day>1</day> [cited <year>2022</year> <month>Jul</month> <day>5</day>];<volume>46</volume>(<issue>6</issue>):<fpage>367</fpage>&#x2013;<lpage>77</lpage>. Available from: <uri>https://jme.bmj.com/content/46/6/367</uri>.</mixed-citation></ref>
<ref id="ref-17"><label>17</label><mixed-citation publication-type="journal"><string-name><surname>Pantazos</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lauesen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lippert</surname> <given-names>S</given-names></string-name>. <article-title>Preserving medical correctness, readability and consistency in de-identified health records</article-title>. <source>Health Informatics J [Internet]</source>. <year>2017</year> <month>Dec</month> <day>1</day> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>23</volume>(<issue>4</issue>):<fpage>291</fpage>&#x2013;<lpage>303</lpage>. Available from: <pub-id pub-id-type="doi">10.1177/1460458216647760.</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>18</label><mixed-citation publication-type="journal"><string-name><surname>Baclic</surname> <given-names>O</given-names></string-name>, <string-name><surname>Tunis</surname> <given-names>M</given-names></string-name>, <string-name><surname>Young</surname> <given-names>K</given-names></string-name>, <string-name><surname>Doan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Swerdfeger</surname> <given-names>H</given-names></string-name>, <string-name><surname>Schonfeld</surname> <given-names>J</given-names></string-name>. <article-title>Challenges and opportunities for public health made possible by advances in natural language processing</article-title>. <source>Can Commun Dis Rep [Internet]</source>. <year>2020</year> <month>Jun</month> <day>4</day> [cited <year>2022</year> <month>Jul</month> <day>5</day>];<volume>46</volume>(<issue>6</issue>):<fpage>161</fpage>&#x2013;<lpage>8</lpage>. Available from: <uri>https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7343054/</uri>.</mixed-citation></ref>
<ref id="ref-19"><label>19</label><mixed-citation publication-type="website"><string-name><surname>van der Aa</surname> <given-names>H</given-names></string-name>, <string-name><surname>Carmona Vargas</surname> <given-names>J</given-names></string-name>, <string-name><surname>Leopold</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mendling</surname> <given-names>J</given-names></string-name>, <string-name><surname>Padr&#x00F3;</surname> <given-names>L</given-names></string-name>. <chapter-title>Challenges and opportunities of applying natural language processing in business process management</chapter-title>. In <source>Association for Computational Linguistics</source>; <year>2018</year> [cited <year>2022</year> <month>Jul</month> <day>5</day>]. p. <fpage>2791</fpage>&#x2013;<lpage>801</lpage>. Available from: <uri>https://upcommons.upc.edu/handle/2117/121682</uri>.</mixed-citation></ref>
<ref id="ref-20"><label>20</label><mixed-citation publication-type="journal"><string-name><surname>Arksey</surname> <given-names>H</given-names></string-name>, <string-name><surname>O&#x2019;Malley</surname> <given-names>L</given-names></string-name>. <article-title>Scoping studies: towards a methodological framework</article-title>. <source>Int J Soc Res Methodol [Internet]</source>. <year>2005</year>;<volume>8</volume>(<issue>1</issue>):<fpage>19</fpage>&#x2013;<lpage>32</lpage>. Available from: <pub-id pub-id-type="doi">10.1080/1364557032000119616.</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Ford</surname> <given-names>E</given-names></string-name>, <string-name><surname>Carroll</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Smith</surname> <given-names>HE</given-names></string-name>, <string-name><surname>Scott</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cassell</surname> <given-names>JA</given-names></string-name>. <article-title>Extracting information from the text of electronic medical records to improve case detection: a systematic review</article-title>. <source>J Am Med Inform Assoc [Internet]</source>. <year>2016</year> [cited <year>2019</year> <month>May</month> <day>21</day>];<volume>23</volume>(<issue>5</issue>):<fpage>1007</fpage>&#x2013;<lpage>15</lpage>. Available from: <uri>http://academic.oup.com/jamia/article/23/5/1007/2379833</uri>.</mixed-citation></ref>
<ref id="ref-22"><label>22</label><mixed-citation publication-type="journal"><string-name><surname>Hinds</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lix</surname> <given-names>LM</given-names></string-name>, <string-name><surname>Smith</surname> <given-names>M</given-names></string-name>, <string-name><surname>Quan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sanmartin</surname> <given-names>C</given-names></string-name>. <article-title>Quality of administrative health databases in Canada: A scoping review</article-title>. <source>Can J Public Health [Internet]</source>. <year>2016</year>;<volume>107</volume>(<issue>1</issue>):<fpage>e56</fpage>-<lpage>61</lpage>. Available from: <uri>https://link-springer-com.uml.idm.oclc.org/article/10.17269/cjph.107.5244</uri>.</mixed-citation></ref>
<ref id="ref-23"><label>23</label><mixed-citation publication-type="book"><string-name><surname>Da&#x0159;ena</surname> <given-names>F</given-names></string-name>, <string-name><surname>&#x017D;i&#x017E;ka</surname> <given-names>J</given-names></string-name>. <chapter-title>Interdependence of text mining quality and the input data preprocessing</chapter-title>. In: <string-name><surname>Silhavy</surname> <given-names>R</given-names></string-name>, <string-name><surname>Senkerik</surname> <given-names>R</given-names></string-name>, <string-name><surname>Oplatkova</surname> <given-names>ZK</given-names></string-name>, <string-name><surname>Prokopova</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Silhavy</surname> <given-names>P</given-names></string-name>, editors. <source>Artificial Intelligence Perspectives and Applications [Internet]</source>. <publisher-name>Cham</publisher-name>: <publisher-loc>Springer International Publishing</publisher-loc>; <year>2015</year>. p. <fpage>141</fpage>&#x2013;<lpage>50</lpage>. (Advances in Intelligent Systems and Computing). Available from: <uri>citeashttps://link-springer-com.uml.idm.oclc.org/chapter/10.1007/978-3-319-18476-0_15#citeas</uri>.</mixed-citation></ref>
<ref id="ref-24"><label>24</label><mixed-citation publication-type="journal"><string-name><surname>Vijayarani</surname> <given-names>M</given-names></string-name>. <article-title>Preprocessing techniques for text mining - an overview</article-title>. <source>Int J Comput Netw Commun [Internet]</source>. <year>2015</year>;<volume>5</volume>(<issue>1</issue>):<fpage>7</fpage>&#x2013;<lpage>16</lpage>. Available from: <uri>https://www.researchgate.net/profile/Vijayarani-Mohan/publication/339529230_Preprocessing_Techniques_for_Text_Mining_-_An_Overview/links/5e57a0f7299bf1bdb83e7505/Preprocessing-Techniques-for-Text-Mining-An-Overview.pdf</uri>.</mixed-citation></ref>
<ref id="ref-25"><label>25</label><mixed-citation publication-type="book"><string-name><surname>Malak</surname> <given-names>P</given-names></string-name>. <chapter-title>Text preprocessing: a tool of information visualization and digital humanities [Internet]</chapter-title>. <publisher-name>Information Visualization Techniques in the Social Sciences and Humanities</publisher-name>. <publisher-loc>IGI Global</publisher-loc>; <year>2018</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>]. p. <fpage>86</fpage>&#x2013;<lpage>104</lpage>. Available from: <uri>http://www.igi-global.com/chapter/text-preprocessing/201305</uri>.</mixed-citation></ref>
<ref id="ref-26"><label>26</label><mixed-citation publication-type="journal"><string-name><surname>Ouzzani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hammady</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fedorowicz</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Elmagarmid</surname> <given-names>A</given-names></string-name>. <article-title>Rayyan&#x2014;a web and mobile app for systematic reviews</article-title>. <source>Systematic Reviews [Internet]</source>. <year>2016</year>;<volume>5</volume>(<issue>1</issue>):<fpage>210</fpage>. Available from: <pub-id pub-id-type="doi">10.1186/s13643-016-0384-4.</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>27</label><mixed-citation publication-type="journal"><string-name><surname>Spasic</surname> <given-names>I</given-names></string-name>, <string-name><surname>Nenadic</surname> <given-names>G</given-names></string-name>. <article-title>Clinical text data in machine learning: systematic review</article-title>. <source>JMIR Med Inform [Internet]</source>. <year>2020</year>;<volume>8</volume>(<issue>3</issue>):<fpage>e17984</fpage>. Available from: <uri>https://medinform.jmir.org/2020/3/e17984/</uri>.</mixed-citation></ref>
<ref id="ref-28"><label>28</label><mixed-citation publication-type="book"><string-name><surname>Kiefer</surname> <given-names>C</given-names></string-name>. <chapter-title>Quality indicators for text data</chapter-title>. In: <source>BTW 2019&#x2013;Workshopband [Internet]</source>. <publisher-name>Gesellschaft f&#x00FC;r Informatik</publisher-name>, <publisher-loc>Bonn</publisher-loc>; <year>2019</year>. p. <fpage>145</fpage>&#x2013;<lpage>54</lpage>. Available from: <uri>https://dl.gi.de/handle/20.500.12116/21801</uri>.</mixed-citation></ref>
<ref id="ref-29"><label>29</label><mixed-citation publication-type="journal"><string-name><surname>Gero</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>J</given-names></string-name>. <article-title>PMCVec: Distributed phrase representation for biomedical text processing</article-title>. <source>J Biomed Inform [Internet]</source>. <year>2019</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>100</volume>:<fpage>100047</fpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S2590177X19300460</uri>.</mixed-citation></ref>
<ref id="ref-30"><label>30</label><mixed-citation publication-type="journal"><string-name><surname>Berndt</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>McCart</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Finch</surname> <given-names>DK</given-names></string-name>, <string-name><surname>Luther</surname> <given-names>SL</given-names></string-name>. <article-title>A case study of data quality in text mining clinical progress notes</article-title>. <source>ACM Trans Manag Inf Syst [Internet]</source>. <year>2015</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>6</volume>(<issue>1</issue>):<fpage>1:1</fpage>-<lpage>1:21</lpage>. Available from: <pub-id pub-id-type="doi">10.1145/2669368.</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>31</label><mixed-citation publication-type="journal"><string-name><surname>Assale</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dui</surname> <given-names>LG</given-names></string-name>, <string-name><surname>Cina</surname> <given-names>A</given-names></string-name>, <string-name><surname>Seveso</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cabitza</surname> <given-names>F</given-names></string-name>. <article-title>The revival of the notes field: leveraging the unstructured content in electronic health records</article-title>. <source>Front Med [Internet]</source>. <year>2019</year>;<volume>6</volume>:<fpage>66</fpage>. Available from: <uri>https://www.frontiersin.org/articles/10.3389/fmed.2019.00066/full</uri>.</mixed-citation></ref>
<ref id="ref-32"><label>32</label><mixed-citation publication-type="journal"><string-name><surname>He</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>B</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Qu</surname> <given-names>C</given-names></string-name>. <article-title>Building a comprehensive syntactic and semantic corpus of Chinese clinical texts</article-title>. <source>J Biomed Inform [Internet]</source>. <year>2017</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>69</volume>:<fpage>203</fpage>&#x2013;<lpage>17</lpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S153204641730076X</uri>.</mixed-citation></ref>
<ref id="ref-33"><label>33</label><mixed-citation publication-type="journal"><string-name><surname>Liang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Christian</surname> <given-names>N.</given-names></string-name>, <string-name><surname>Kuziemsky</surname> <given-names>C.E.</given-names></string-name>, <string-name><surname>Kushniruk</surname> <given-names>A.W.</given-names></string-name>, <string-name><surname>Borycki</surname> <given-names>E.M</given-names></string-name>. <article-title>Enhancing patient safety event reporting by K-nearest neighbor classifier</article-title>. <source>Stud Health Technol Informatics [Internet]</source>. <year>2015</year>;<volume>218</volume>:<fpage>93</fpage>&#x2013;<lpage>9</lpage>. Available from: <uri>https://www.scopus.com/inward/record.uri?eid=2-s2.0-84951923066&#x0026;doi=10.3233%2f978-1-61499-574-6-93&#x0026;partnerID=40&#x0026;md5=9dd7ea3d83936147533292dec0f1447b</uri>.</mixed-citation></ref>
<ref id="ref-34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Joopudi</surname> <given-names>V</given-names></string-name>, <string-name><surname>Dandala</surname> <given-names>B</given-names></string-name>, <string-name><surname>Devarakonda</surname> <given-names>M</given-names></string-name>. <article-title>A convolutional route to abbreviation disambiguation in clinical text</article-title>. <source>J Biomed Inform</source>. <year>2018</year>;<volume>86</volume>:<fpage>71</fpage>&#x2013;<lpage>8</lpage>. Available from: <uri>https://www-sciencedirect-com.uml.idm.oclc.org/science/article/pii/S1532046418301552</uri>.</mixed-citation></ref>
<ref id="ref-35"><label>35</label><mixed-citation publication-type="journal"><string-name><surname>Lai</surname> <given-names>KH</given-names></string-name>, <string-name><surname>Topaz</surname> <given-names>M</given-names></string-name>, <string-name><surname>Goss</surname> <given-names>FR</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>L</given-names></string-name>. <article-title>Automated misspelling detection and correction in clinical free-text records</article-title>. <source>J Biomed Inform [Internet]</source>. <year>2015</year> [cited <year>2019</year> <month>Aug</month> <day>3</day>];<volume>55</volume>:<fpage>188</fpage>&#x2013;<lpage>95</lpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S1532046415000751</uri>.</mixed-citation></ref>
<ref id="ref-36"><label>36</label><mixed-citation publication-type="journal"><string-name><surname>Ruch</surname> <given-names>P</given-names></string-name>, <string-name><surname>Baud</surname> <given-names>R</given-names></string-name>, <string-name><surname>Geissb&#x00FC;hler</surname> <given-names>A</given-names></string-name>. <article-title>Using lexical disambiguation and named-entity recognition to improve spelling correction in the electronic patient record</article-title>. <source>Artificial Intelligence in Medicine [Internet]</source>. <year>2003</year> <month>Sep</month> <day>1</day> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>29</volume>(<issue>1</issue>):<fpage>169</fpage>&#x2013;<lpage>84</lpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S0933365703000526</uri>.</mixed-citation></ref>
<ref id="ref-37"><label>37</label><mixed-citation publication-type="journal"><string-name><surname>Hoste</surname> <given-names>V</given-names></string-name>, <string-name><surname>Hendrickx</surname> <given-names>I</given-names></string-name>, <string-name><surname>Daelemans</surname> <given-names>W</given-names></string-name>, <string-name><surname>Bosch</surname> <given-names>AVD</given-names></string-name>. <article-title>Parameter optimization for machine-learning of word sense disambiguation</article-title>. <source>Natural Language Engineering [Internet]</source>. <year>2002</year> <month>Dec</month> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>8</volume>(<issue>4</issue>):<fpage>311</fpage>&#x2013;<lpage>25</lpage>. Available from: <uri>http://www.cambridge.org/core/journals/natural-language-engineering/article/parameter-optimization-for-machinelearning-of-word-sense-disambiguation/ED5DAA7F32FEFBC5D1C81B22DEA5F43B#</uri>.</mixed-citation></ref>
<ref id="ref-38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Nguyen</surname> <given-names>QT</given-names></string-name>, <string-name><surname>Miyao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Le</surname> <given-names>HTT</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>NTH</given-names></string-name>. <article-title>Ensuring annotation consistency and accuracy for Vietnamese treebank</article-title>. <source>Lang Resour Eval [Internet]</source>. <year>2018</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>52</volume>(<issue>1</issue>):<fpage>269</fpage>&#x2013;<lpage>315</lpage>. Available from: <uri>https://doi-org.uml.idm.oclc.org/10.1007/s10579-017-9398-3</uri>.</mixed-citation></ref>
<ref id="ref-39"><label>39</label><mixed-citation publication-type="journal"><string-name><surname>Nguyen</surname> <given-names>PT</given-names></string-name>, <string-name><surname>Le</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>TB</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>VH</given-names></string-name>. <article-title>Vietnamese treebank construction and entropy-based error detection</article-title>. <source>Lang Resour Eval [Internet]</source>. <year>2015</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>49</volume>(<issue>3</issue>):<fpage>487</fpage>&#x2013;<lpage>519</lpage>. Available from: <uri>https://doi-org.uml.idm.oclc.org/10.1007/s10579-015-9308-5</uri>.</mixed-citation></ref>
<ref id="ref-40"><label>40</label><mixed-citation publication-type="website"><string-name><surname>Qin</surname> <given-names>kan</given-names></string-name>, <string-name><surname>Yujiu</surname> <given-names>Yang</given-names></string-name>, <string-name><surname>Wenhuang</surname> <given-names>Liu</given-names></string-name>, <string-name><surname>Xiaodong</surname> <given-names>Liu</given-names></string-name>. <chapter-title>An integrated approach for detecting approximate duplicate records</chapter-title>. In: <source>2009 Asia-Pacific Conference on Computational Intelligence and Industrial Applications (PACIIA)</source>. <year>2009</year>. p. <fpage>381</fpage>&#x2013;<lpage>4</lpage>. Available from: <uri>https://ieeexplore-ieee-org.uml.idm.oclc.org/document/5406409</uri>.</mixed-citation></ref>
<ref id="ref-41"><label>41</label><mixed-citation publication-type="book"><string-name><surname>Uejima</surname> <given-names>H</given-names></string-name>, <string-name><surname>Miura</surname> <given-names>T</given-names></string-name>, <string-name><surname>Shioya</surname> <given-names>I</given-names></string-name>. <chapter-title>Improving text categorization by resolving semantic ambiguity</chapter-title>. In: <source>2003 IEEE Pacific Rim Conference on Communications Computers and Signal Processing (PACRIM 2003) (Cat No03CH37490) [Internet]</source>. <publisher-name>IEEE</publisher-name>; <year>2003</year>. p. <fpage>796</fpage>&#x2013;<lpage>9</lpage>. Available from: <uri>https://ieeexplore-ieee-org.uml.idm.oclc.org/abstract/document/1235901</uri>.</mixed-citation></ref>
<ref id="ref-42"><label>42</label><mixed-citation publication-type="journal"><string-name><surname>Zurini</surname> <given-names>M</given-names></string-name>. <article-title>Word sense disambiguation using aggregated similarity based on WordNet graph representation</article-title>. <source>Inform econ [Internet]</source>. <year>2013</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>17</volume>(<issue>3</issue>):<fpage>169</fpage>&#x2013;<lpage>80</lpage>. Available from: <uri>https://ideas.repec.org/a/aes/infoec/v17y2013i3p169-180.html</uri>.</mixed-citation></ref>
<ref id="ref-43"><label>43</label><mixed-citation publication-type="book"><string-name><surname>&#x0160;najder</surname> <given-names>J</given-names></string-name>. <chapter-title>DerivBase.hr: A high-coverage derivational morphology resource for croatian</chapter-title>. In: <source>Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC&#x2019;14) [Internet]</source>. <publisher-loc>Reykjavik, Iceland</publisher-loc>: <publisher-name>European Language Resources Association (ELRA)</publisher-name>; <year>2014</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>]. p. <fpage>3371</fpage>&#x2013;<lpage>7</lpage>. Available from: <uri>http://www.lrec-conf.org/proceedings/lrec2014/pdf/1090_Paper.pdf</uri>.</mixed-citation></ref>
<ref id="ref-44"><label>44</label><mixed-citation publication-type="book"><string-name><surname>Edwards</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zatorsky</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nayak</surname> <given-names>R</given-names></string-name>. <chapter-title>Clustering and classification of maintenance logs using text data mining</chapter-title>. In: <source>Conf Res Pract Inf Technol Ser [Internet]</source>. <publisher-name>CRC for Integrated Engineering Asset Management, Faculty of Information Technology, Queensland University of Technology</publisher-name>, <publisher-loc>PO Box 2434, Brisbane 4001, QLD, Australia: AusDM</publisher-loc>; <year>2008</year>. p. <fpage>193</fpage>&#x2013;<lpage>9</lpage>. Available from: <uri>https://www.scopus.com/inward/record.uri?eid=2-s2.0-84870484659&#x0026;partnerID=40&#x0026;md5=f34d58eba076a357439884fae90dc336</uri>.</mixed-citation></ref>
<ref id="ref-45"><label>45</label><mixed-citation publication-type="journal"><string-name><surname>Marzec</surname> <given-names>M</given-names></string-name>, <string-name><surname>Uhl</surname> <given-names>T</given-names></string-name>, <string-name><surname>Michalak</surname> <given-names>D</given-names></string-name>. <article-title>Verification of text mining techniques accuracy when dealing with urban buses maintenance data</article-title>. <source>Diagnostyka [Internet]</source>. <year>2017</year> <month>Dec</month> <day>20</day> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>15</volume>(<issue>3</issue>):<fpage>51</fpage>&#x2013;<lpage>7</lpage>. Available from: <uri>http://www.diagnostyka.net.pl/Verification-of-text-mining-techniques-accuracy-when-dealing-with-urban-buses-maintenance,81413,0,2.html</uri>.</mixed-citation></ref>
<ref id="ref-46"><label>46</label><mixed-citation publication-type="website"><string-name><surname>Stewart</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cardell-Oliver</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>R</given-names></string-name>. <chapter-title>Short-Text Lexical Normalisation on Industrial Log Data</chapter-title>. In: <source>2018 IEEE International Conference on Big Knowledge (ICBK)</source>. <year>2018</year>. p. <fpage>113</fpage>&#x2013;<lpage>22</lpage>. Available from: <uri>https://ieeexplore-ieee-org.uml.idm.oclc.org/document/8588782</uri>.</mixed-citation></ref>
<ref id="ref-47"><label>47</label><mixed-citation publication-type="book"><string-name><surname>Subha</surname> <given-names>R</given-names></string-name>, <string-name><surname>Palaniswami</surname> <given-names>S</given-names></string-name>. <chapter-title>Quality Factor Assessment and Text Summarization of Unambiguous Natural Language Requirements</chapter-title>. In: <string-name><surname>Unnikrishnan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Surve</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bhoir</surname> <given-names>D</given-names></string-name>, editors. <source>Advances in Computing, Communication, and Control</source>. <publisher-name>Berlin, Heidelberg</publisher-name>: <publisher-loc>Springer</publisher-loc>; <year>2013</year>. p. <fpage>131</fpage>&#x2013;<lpage>46</lpage>. Available from: <uri>https://doi-org.uml.idm.oclc.org/10.1007/978-3-642-36321-4_12</uri>.</mixed-citation></ref>
<ref id="ref-48"><label>48</label><mixed-citation publication-type="journal"><string-name><surname>Do&#x011F;an</surname> <given-names>RI</given-names></string-name>, <string-name><surname>Leaman</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>. <article-title>NCBI disease corpus: a resource for disease name recognition and concept normalization</article-title>. <source>J Biomed Inform [Internet]</source>. <year>2014</year>;<volume>47</volume>:<fpage>1</fpage>&#x2013;<lpage>10</lpage>. Available from: <uri>https://www-sciencedirect-com.uml.idm.oclc.org/science/article/pii/S1532046413001974?via%3Dihub</uri>.</mixed-citation></ref>
<ref id="ref-49"><label>49</label><mixed-citation publication-type="journal"><string-name><surname>Jung</surname> <given-names>Y</given-names></string-name>. <article-title>A semantic annotation framework for scientific publications</article-title>. <source>Qual Quant [Internet]</source>. <year>2017</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>51</volume>(<issue>3</issue>):<fpage>1009</fpage>&#x2013;<lpage>25</lpage>. Available from: <pub-id pub-id-type="doi">10.1007/s11135-016-0369-3.</pub-id></mixed-citation></ref>
<ref id="ref-50"><label>50</label><mixed-citation publication-type="journal"><string-name><surname>Kang</surname> <given-names>N</given-names></string-name>, <string-name><surname>van Mulligen</surname> <given-names>EM</given-names></string-name>, <string-name><surname>Kors</surname> <given-names>JA</given-names></string-name>. <article-title>Comparing and combining chunkers of biomedical text</article-title>. <source>Journal of Biomedical Informatics [Internet]</source>. <year>2011</year> <month>Apr</month> <day>1</day> [cited <year>2020</year> <month>Jun</month> <source>6</source>];<volume>44</volume>(<issue>2</issue>):<fpage>354</fpage>&#x2013;<lpage>60</lpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S1532046410001577</uri>.</mixed-citation></ref>
<ref id="ref-51"><label>51</label><mixed-citation publication-type="journal"><string-name><surname>Tissot</surname> <given-names>H</given-names></string-name>, <string-name><surname>Del Fabro</surname> <given-names>MD</given-names></string-name>, <string-name><surname>Derczynski</surname> <given-names>L</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>. <article-title>Normalisation of imprecise temporal expressions extracted from text</article-title>. <source>Knowl Inf Syst [Internet]</source>. <year>2019</year> <month>Dec</month> <day>1</day> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>61</volume>(<issue>3</issue>):<fpage>1361</fpage>&#x2013;<lpage>94</lpage>. Available from: <pub-id pub-id-type="doi">10.1007/s10115-019-01338-1.</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>52</label><mixed-citation publication-type="journal"><string-name><surname>V&#x00ED;tovec</surname> <given-names>P</given-names></string-name>, <string-name><surname>Kl&#x00E9;ma</surname> <given-names>J</given-names></string-name>. <article-title>Gene interaction extraction from biomedical texts by sentence skeletonization</article-title>. <source>CEUR Workshop Proc [Internet]</source>. <year>2011</year>;<volume>802</volume>:NA. Available from: <uri>https://www.scopus.com/inward/record.uri?eid=2-s2.0-84891772107&#x0026;partnerID=40&#x0026;md5=4c64e3778af84417d329e274327ffff7</uri>.</mixed-citation></ref>
<ref id="ref-53"><label>53</label><mixed-citation publication-type="book"><string-name><surname>Westpfahl</surname> <given-names>S</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>T</given-names></string-name>. <chapter-title>FOLK-Gold - A Gold Standard for Part-of-Speech-Tagging of Spoken German</chapter-title>. In: <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC&#x2019;16) [Internet]</source>. <publisher-loc>Portoro&#x017E;, Slovenia</publisher-loc>: <publisher-name>European Language Resources Association (ELRA)</publisher-name>; <year>2016</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>]. p. <fpage>1493</fpage>&#x2013;<lpage>9</lpage>. Available from: <uri>https://www.aclweb.org/anthology/L16-1237</uri>.</mixed-citation></ref>
<ref id="ref-54"><label>54</label><mixed-citation publication-type="journal"><string-name><surname>Allen</surname> <given-names>TT</given-names></string-name>, <string-name><surname>Sui</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Akbari</surname> <given-names>K</given-names></string-name>. <article-title>Exploratory text data analysis for quality hypothesis generation</article-title>. <source>Qual Eng [Internet]</source>. <year>2018</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>30</volume>(<issue>4</issue>):<fpage>701</fpage>&#x2013;<lpage>12</lpage>. Available from: <pub-id pub-id-type="doi">10.1080/08982112.2018.1481216.</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>55</label><mixed-citation publication-type="book"><string-name><surname>Alnajran</surname> <given-names>N</given-names></string-name>, <string-name><surname>Crockett</surname> <given-names>K</given-names></string-name>, <string-name><surname>McLean</surname> <given-names>D</given-names></string-name>, <string-name><surname>Latham</surname> <given-names>A</given-names></string-name>. <chapter-title>A heuristic based pre-processing methodology for short text similarity measures in microblogs</chapter-title>. In: <source>Proc - Int Conf High Perform Comput Commun, Int Conf Smart City Int Conf Data Sci Syst, HPCC/SmartCity/DSS [Internet]</source>. <publisher-loc>Institute of Electrical and Electronics Engineers Inc.</publisher-loc>; <year>2019</year>. p. <fpage>1627</fpage>&#x2013;<lpage>33</lpage>. Available from: <uri>https://ieeexplore-ieee-org.uml.idm.oclc.org/document/8623003</uri>.</mixed-citation></ref>
<ref id="ref-56"><label>56</label><mixed-citation publication-type="journal"><string-name><surname>Christen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gayler</surname> <given-names>RW</given-names></string-name>, <string-name><surname>Tran</surname> <given-names>KN</given-names></string-name>, <string-name><surname>Fisher</surname> <given-names>J</given-names></string-name>, <string-name><surname>Vatsalan</surname> <given-names>D</given-names></string-name>. <article-title>Automatic discovery of abnormal values in large textual databases</article-title>. <source>J Data Inf Qual [Internet]</source>. <year>2016</year>;<volume>7</volume>(<issue>1&#x2013;2</issue>):<fpage>1</fpage>&#x2013;<lpage>31</lpage>. Available from: <pub-id pub-id-type="doi">10.1145/2889311.</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>57</label><mixed-citation publication-type="website"><string-name><surname>Gharatkar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ingle</surname> <given-names>A</given-names></string-name>, <string-name><surname>Naik</surname> <given-names>T</given-names></string-name>, <string-name><surname>Save</surname> <given-names>A</given-names></string-name>. <chapter-title>Review preprocessing using data cleaning and stemming technique</chapter-title>. In: <source>2017 International Conference on Innovations in Information, Embedded and Communication Systems (ICIIECS)</source>. <year>2017</year>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>. Available from: <uri>https://ieeexplore-ieee-org.uml.idm.oclc.org/document/8276011</uri>.</mixed-citation></ref>
<ref id="ref-58"><label>58</label><mixed-citation publication-type="journal"><string-name><surname>Chen</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Tseng</surname> <given-names>YD</given-names></string-name>. <article-title>Quality evaluation of product reviews using an information quality framework</article-title>. <source>Decis Support Syst [Internet]</source>. <year>2011</year> [cited <year>2020</year> <month>Jun</month> <day>5</day>];<volume>50</volume>(<issue>4</issue>):<fpage>755</fpage>&#x2013;<lpage>68</lpage>. Available from: <uri>http://www.sciencedirect.com/science/article/pii/S0167923610001478</uri>.</mixed-citation></ref>
<ref id="ref-59"><label>59</label><mixed-citation publication-type="journal"><string-name><surname>Xue</surname> <given-names>N</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>F</given-names></string-name>, Chiou F dong, <string-name><surname>Palmer</surname> <given-names>M</given-names></string-name>. <article-title>The Penn Chinese TreeBank: Phrase structure annotation of a large corpus</article-title>. <source>Nat Lang Eng [Internet]</source>. <year>2005</year> <month>Jun</month> <day>1</day> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>11</volume>(<issue>2</issue>):<fpage>207</fpage>&#x2013;<lpage>38</lpage>. Available from: <pub-id pub-id-type="doi">10.1017/S135132490400364X.</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>60</label><mixed-citation publication-type="journal"><string-name><surname>Scharkow</surname> <given-names>M</given-names></string-name>. <article-title>Thematic content analysis using supervised machine learning: An empirical evaluation using German online news</article-title>. <source>Qual Quant [Internet]</source>. <year>2013</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>47</volume>(<issue>2</issue>):<fpage>761</fpage>&#x2013;<lpage>73</lpage>. Available from: <pub-id pub-id-type="doi">10.1007/s11135-011-9545-7.</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>61</label><mixed-citation publication-type="website"><string-name><surname>Laur&#x00ED;a</surname> <given-names>EJM</given-names></string-name>, <string-name><surname>March</surname> <given-names>AD</given-names></string-name>. <chapter-title>Effect of dirty data on free text discharge diagnoses used for automated ICD-9-CM coding</chapter-title>. In: <source>AMCIS [Internet]</source>. <year>2006</year>. Available from: <uri>https://aisel.aisnet.org/amcis2006/188</uri>.</mixed-citation></ref>
<ref id="ref-62"><label>62</label><mixed-citation publication-type="website"><string-name><surname>Abad</surname> <given-names>ZSH</given-names></string-name>, <string-name><surname>Karras</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ghazi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Glinz</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ruhe</surname> <given-names>G</given-names></string-name>, <string-name><surname>Schneider</surname> <given-names>K</given-names></string-name>. <chapter-title>What works better? A study of classifying requirements</chapter-title>. In: <source>2017 IEEE 25th International Requirements Engineering Conference [Internet]</source>. <year>2017</year>. p. <fpage>496</fpage>&#x2013;<lpage>501</lpage>. Available from: <uri>https://arxiv.org/abs/1707.02358</uri>.</mixed-citation></ref>
<ref id="ref-63"><label>63</label><mixed-citation publication-type="journal"><string-name><surname>HaCohen-Kerner</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Miller</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yigal</surname> <given-names>Y</given-names></string-name>. <article-title>The influence of preprocessing on text classification using a bag-of-words representation</article-title>. <source>PLOS ONE [Internet]</source>. <year>2020</year> <month>May</month> <day>1</day> [cited <year>2021</year> <month>Feb</month> <day>21</day>];<volume>15</volume>(<issue>5</issue>):<fpage>e0232525</fpage>. Available from: <uri>https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0232525</uri>.</mixed-citation></ref>
<ref id="ref-64"><label>64</label><mixed-citation publication-type="journal"><string-name><surname>Song</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>KC</given-names></string-name>. <article-title>Observational Studies: Cohort and Case-Control Studies</article-title>. <source>Plast Reconstr Surg [Internet]</source>. <year>2010</year> <month>Dec</month> [cited <year>2019</year> <month>Nov</month> <day>5</day>];<volume>126</volume>(<issue>6</issue>):<fpage>2234</fpage>&#x2013;<lpage>42</lpage>. Available from: <uri>https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2998589/</uri>.</mixed-citation></ref>
<ref id="ref-65"><label>65</label><mixed-citation publication-type="journal"><string-name><surname>Moreno</surname> <given-names>I</given-names></string-name>, <string-name><surname>Boldrini</surname> <given-names>E</given-names></string-name>, <string-name><surname>Moreda</surname> <given-names>P</given-names></string-name>, <string-name><surname>Rom&#x00E1;-Ferri</surname> <given-names>MT</given-names></string-name>. <article-title>DrugSemantics: A corpus for named entity recognition in Spanish summaries of product characteristics</article-title>. <source>J Biomed Inform [Internet]</source>. <year>2017</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>72</volume>:<fpage>8</fpage>&#x2013;<lpage>22</lpage>. Available from: <pub-id pub-id-type="doi">10.1016/j.jbi.2017.06.013.</pub-id></mixed-citation></ref>
<ref id="ref-66"><label>66</label><mixed-citation publication-type="journal"><string-name><surname>Song</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>B</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>. <article-title>Intelligent assessment of 95598 speech transcription text quality based on topic model</article-title>. <source>IOP Conf Ser: Mater Sci Eng [Internet]</source>. <year>2019</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>563</volume>(<issue>4</issue>):<fpage>2001</fpage>. Available from: <pub-id pub-id-type="doi">10.1088%2F1757-899x%2F563%2F4%2F042001.</pub-id></mixed-citation></ref>
<ref id="ref-67"><label>67</label><mixed-citation publication-type="journal">Rianto, Mutiara Achmad Benny, Wibowo Eri Prasetyo, <string-name><surname>Santosa</surname> <given-names>PI</given-names></string-name>. <article-title>Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation</article-title>. <source>J Big Data [Internet]</source>. <year>2021</year>;<volume>8</volume>(<issue>1</issue>):<fpage>26</fpage>. Available from: <uri>https://journalofbigdata.springeropen.com/articles/10.1186/s40537-021-00413-1</uri>.</mixed-citation></ref>
<ref id="ref-68"><label>68</label><mixed-citation publication-type="journal"><string-name><surname>Laur&#x00ED;a</surname> <given-names>EJM</given-names></string-name>, <string-name><surname>March</surname> <given-names>AD</given-names></string-name>. <article-title>Combining bayesian text classification and shrinkage to automate healthcare coding: A data quality analysis</article-title>. <source>ACM J Data Inf Qual [Internet]</source>. <year>2011</year> [cited <year>2020</year> <month>Jun</month> <day>6</day>];<volume>2</volume>(<issue>3</issue>):<fpage>13:1</fpage>-<lpage>13:22</lpage>. Available from: <pub-id pub-id-type="doi">10.1145/2063504.2063506.</pub-id></mixed-citation></ref>
<ref id="ref-69"><label>69</label><mixed-citation publication-type="book"><string-name><surname>Strong</surname> <given-names>DM</given-names></string-name>. <chapter-title>Information quality: managing information as a product</chapter-title>. In: <string-name><surname>LIU</surname> <given-names>L</given-names></string-name>, <string-name><surname>&#x00D6;ZSU</surname> <given-names>MT</given-names></string-name>, editors. <source>Encyclopedia of Database Systems [Internet]</source>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Springer US</publisher-name>; <year>2009</year> [cited <year>2021</year> <month>Jun</month> <day>12</day>]. p. <fpage>1502</fpage>&#x2013;<lpage>8</lpage>. Available from: <pub-id pub-id-type="doi">10.1007/978-0-387-39940-9_497.</pub-id></mixed-citation></ref>
<ref id="ref-70"><label>70</label><mixed-citation publication-type="journal"><string-name><surname>Weiskopf</surname> <given-names>NG</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>C</given-names></string-name>. <article-title>Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical research</article-title>. <source>J Am Med Inform Assoc [Internet]</source>. <year>2013</year>;<volume>20</volume>(<issue>1</issue>):<fpage>144</fpage>&#x2013;<lpage>51</lpage>. Available from: <uri>https://doi-org.uml.idm.oclc.org/10.1136/amiajnl-2011-000681</uri>.</mixed-citation></ref>
<ref id="ref-71"><label>71</label><mixed-citation publication-type="book"><string-name><surname>Pantazos</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lauesen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lippert</surname> <given-names>S</given-names></string-name>. <chapter-title>De-identifying an EHR database-Anonymity, correctness and readability of the medical record [Internet]</chapter-title>. <publisher-name>Stud. Health Technol. Informatics</publisher-name>. <publisher-loc>IOS Press</publisher-loc>; <year>2011</year>. <fpage>862</fpage>&#x2013;<lpage>866</lpage> p. (Health Technol. Informatics; vol. 169). Available from: <uri>https://ebooks.iospress.nl/publication/14293</uri>.</mixed-citation></ref>
</ref-list>
<glossary>
<title>Abbreviations</title>
<array>
<tbody>
<tr>
<td>UTD</td>
<td>Unstructured Text Data</td>
</tr>
<tr>
<td>EMR</td>
<td>Electronic Medical Records</td>
</tr>
<tr>
<td>CI</td>
<td>Confidence Interval</td>
</tr>
<tr>
<td>NLP</td>
<td>Natural Language Processing</td>
</tr>
<tr>
<td>PRISMA</td>
<td>Preferred Reporting Items for Systematic Reviews and Meta-Analysis</td></tr>
</tbody>
</array>
</glossary>
</back>
</article>