<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v8i1.2115</article-id>
<article-id pub-id-type="publisher-id">8:1:2115</article-id>
<article-id pub-id-type="pii">S2399490821021157</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Thirty-three myths and misconceptions about population data: from data capture and processing to linkage</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Christen</surname><given-names initials="P">Peter</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-2">2</xref><xref ref-type="corresp" rid="correspondingAurthor">*</xref></contrib>
<contrib contrib-type="author"><name><surname>Schnell</surname><given-names initials="R">Rainer</given-names></name><xref ref-type="aff" rid="affil-3">3</xref></contrib>
<aff id="affil-1"><label>1</label><institution>School of Computing, The Australian National University, Canberra, ACT 2600, Australia</institution></aff>
<aff id="affil-2"><label>2</label><institution>Scottish Centre for Administrative Data Research (SCADR), University of Edinburgh. UK</institution></aff>
<aff id="affil-3"><label>3</label><institution>Methodology Research Group, University Duisburg-Essen, Germany</institution></aff>
</contrib-group>
<author-notes>
<corresp id="correspondingAurthor"><label>*</label>Corresponding author: Peter Christen <email>peter.christen@anu.edu.au</email>
</corresp>
<fn fn-type="conflict">
<label>Statement on conflicts of interest</label>
<p>The authors have no conflicts of interest.</p>
</fn>
</author-notes>
<pub-date date-type="pub" publication-format="electronic"><day>31</day><month>01</month><year>2023</year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<volume>8</volume>
<issue>1</issue>
<elocation-id>2115</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/2115">This article is available from the IJPDS website at: https://ijpds.org/article/view/2115</self-uri>
<abstract>
<title>Abstract</title>
<p>Databases covering all individuals of a population are increasingly used for research and decision-making. The massive size of such databases is often mistaken as a guarantee for valid inferences. However, population data have characteristics that make them challenging to use. Various assumptions on population coverage and data quality are commonly made, including how such data were captured and what types of processing have been applied to them. Furthermore, the full potential of population data can often only be unlocked when such data are linked to other databases. Record linkage often implies subtle technical problems, which are easily missed. We discuss a diverse range of myths and misconceptions relevant for anybody capturing, processing, linking, or analysing population data. Remarkably, many of these myths and misconceptions are due to the social nature of data collections and are therefore missed by purely technical accounts of data processing. Many are also not well documented in scientific publications. We conclude with a set of recommendations for using population data.</p>
</abstract>
<kwd-group>
<kwd>data quality</kwd>
<kwd>record linkage</kwd>
<kwd>data linkage</kwd>
<kwd>personal data</kwd>
<kwd>administrative data</kwd>
<kwd>data editing</kwd>
<kwd>data errors</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>Introduction</title>
<p>Many domains of science increasingly use large administrative or operational databases that cover whole populations to replace &#x2013; or at least enrich &#x2013; traditional data collection methods such as surveys or experiments [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. This kind of data are now seen as a crucial strategic resource for research [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. Governments and businesses have also recognised the value that large population databases can provide to improve decision-making [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>Due to the perceived advantages of population data, the number of projects adopting existing databases for research and planning is increasing. The use of buzzwords like big data, AI, and machine learning, in the context of population data seems to suggest for non-technical users and decision makers that any kind of question can be answered when analysing population databases [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>However, often neither data quality issues (how population data were captured and processed) nor the techniques used to link population data, are clear to decision makers and researchers used to smaller data sets. Although there is much work on general data quality, very little specific to population data is published or taught in data science courses.</p>
<p>Therefore, the kind of problems we consider in this paper are usually underestimated by non-specialists, leading to inflated expectations. Such over-expectations might cause costly mismanagement in areas such as public health or in government decision-making. Furthermore, failing population data projects, such as census operations or health surveillance, might even result in the loss of trust in governments and science by the public [<xref ref-type="bibr" rid="ref-8">8</xref>, <xref ref-type="bibr" rid="ref-11">11</xref>]. In the context of research, myths and misconceptions<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> about population data can lead to wrong outcomes of research studies that can result in conclusions with severe negative impact [<xref ref-type="bibr" rid="ref-12">12</xref>, <xref ref-type="bibr" rid="ref-13">13</xref>].</p>
<p>We focus on personal data about individuals covering (nearly) whole populations. Following a recent definition of <italic>Population Data Science</italic> [<xref ref-type="bibr" rid="ref-14">14</xref>], we define population data as <italic>data about people at the level of a population</italic>. The focus on populations is important, as it refers to the scale and complexity of the data being considered. These make (manual) processing and assessment of data quality that are normally conducted on the much smaller data sets used in traditional medical studies or social science surveys challenging.</p>
<p>Personal data include personally identifiable information (PII) [<xref ref-type="bibr" rid="ref-15">15</xref>], such as the names, addresses, or dates of birth of people. Most administrative and operational data collected by governments and businesses can also be categorised as personal data, including electronic health records, (online) shopping baskets, and people&#x2019;s educational, social security, financial, and taxation data [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>A crucial aspect of population data is that they are not primarily collected for research, but rather for operational or administrative purposes [<xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-10">10</xref>, <xref ref-type="bibr" rid="ref-14">14</xref>, <xref ref-type="bibr" rid="ref-17">17</xref>]. As a result, researchers have much less control over such data and how they are processed, and only limited ways to learn about the data&#x2019;s provenance, making it unclear if such data are fit for the purpose of a specific research study [<xref ref-type="bibr" rid="ref-18">18</xref>]. Quoting Brown et al. [<xref ref-type="bibr" rid="ref-19">19</xref>], &#x201C;in science, three things matter: the data, the methods used to collect the data (which give them their probative value), and the logic connecting the data and methods to conclusions&#x201D;. When both the data and their collection are outside the control of a researcher, conducting proper science can become challenging.</p>
<p>Most of the discussion about the use of population data for research has been about privacy and confidentiality [<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>, <xref ref-type="bibr" rid="ref-21">21</xref>]. Much less consideration has been given to how data quality and assumptions about such data can influence the outcomes of a research study. For administrative data, decades of experience have shown that many unexpected patterns (such as unusual combinations of attribute values) in large databases are due to data errors instead of anything of actual interest [<xref ref-type="bibr" rid="ref-10">10</xref>]. The same can be said about population covering databases of personal data.</p>
<p>Not much has been written about the characteristics of personal data and how they differ from other types of data, such as scientific or financial data. During the writing of this article, we conducted extensive literature research to find scientific publications, government reports, or technical white papers that describe experiences or challenges when dealing with population databases. The number of identified publications was scarce, indicating that many challenges encountered and lessons learnt are not being shared, even though these would be of high value to anybody who works with population data.</p>
<p>Many of the misconceptions we discuss below are therefore drawn from our several decades of experience working with large real-world population databases and in collaborative projects with both private and public sector organisations across diverse disciplines (including health, finance, and national statistics).</p>
<p>We do not advocate to abandon the use of population data for research or decision making. Rather, with this work we aim to improve the handling of population databases. By presenting common misconceptions about population data, we hope to show how researchers could identify and avoid the resulting traps.</p>
</sec>
<sec>
<title>Characteristics and uses of population data</title>
<p>Population data are observational data that most commonly occur in administrative or operational databases as held by government or business organisations [<xref ref-type="bibr" rid="ref-10">10</xref>]. As we illustrate in <xref ref-type="fig" rid="fig-1">Figure 1</xref>, in their most general form each <italic>entity</italic> (person) in a population is represented by one or more <italic>records</italic> (like rows in a spreadsheet) in such a database, each consisting of a set of <italic>attributes</italic> (or <italic>variables</italic>).</p>
<fig id="fig-1"><label>Figure 1: Two small example population databases, where only one contains unique entity identifiers (the PID attribute), but both contain quasi-identifiers (QIDs) and microdata</label>
<graphic xlink:href="ijpds-06-2115-g001.tif"/>
<attrib>The first database consists of eight attributes and the second of six, where both have three records. The QIDs have been used to link records across these two databases. Note the variations and missing values in QIDs.</attrib>
</fig>
<p>The attributes that represent entities in population data can be categorised into <italic>identifiers</italic> and <italic>microdata</italic>. The first category can either be an <italic>entity identifier</italic> (such as patient identifiers) that is (supposed to be) unique for each person in a population [<xref ref-type="bibr" rid="ref-22">22</xref>]. Alternatively they can be a group of attributes that, when enough are combined, become unique for each individual. Such attributes are known as <italic>quasi-identifiers</italic> (QIDs), and they include names, addresses, date and place of birth, and so on [<xref ref-type="bibr" rid="ref-15">15</xref>]. Many of the misconceptions about population data are about entity identifiers and QIDs, and how they are captured, processed, and used to link population databases to transform them into a form suitable for a research study.</p>
<p>The second component of personal data are known as <italic>microdata</italic> [<xref ref-type="bibr" rid="ref-15">15</xref>] (or payload data), and they include the data of interest for research, such as the medical, educational, financial, or location details of individuals. Much of these microdata are highly sensitive if they are connected to QID values because combined they can reveal private personal details about an individual. Research into data anonymisation and disclosure control [<xref ref-type="bibr" rid="ref-23">23</xref>, <xref ref-type="bibr" rid="ref-24">24</xref>] is addressing how sensitive microdata can be made available to researchers in anonymised form.</p>
<p>It is commonly recognised that isolated population databases are of limited use when trying to investigate and solve today&#x2019;s complex challenges [<xref ref-type="bibr" rid="ref-14">14</xref>, <xref ref-type="bibr" rid="ref-18">18</xref>], such as how a pandemic spreads through a population. Therefore, many projects that are based on population data employ <italic>data</italic> or <italic>record linkage</italic> [<xref ref-type="bibr" rid="ref-22">22</xref>, <xref ref-type="bibr" rid="ref-25">25</xref>] to link all records that correspond to the same person across two or more databases. Linked records at the level of individuals, rather than aggregated data, are generally required to allow the development of accurate predictive models.</p>
<p>We illustrate the general pathway of population data in <xref ref-type="fig" rid="fig-2">Figure 2</xref>, and discuss each stage (how data are captured, processed, and linked) and the corresponding misconceptions below. While many of these misconceptions seem obvious, they are often not taken into consideration when population data are used in research studies or for decision making.</p>
<fig id="fig-2"><label>Figure 2: The general pathway of population data from their source to a researcher&#x2019;s computer</label>
<graphic xlink:href="ijpds-06-2115-g002.tif"/>
<attrib>We assume <bold><italic>D</italic></bold><sub>1</sub> to <bold><italic>D</italic></bold><sub>n</sub> (with n ≥ 2) are the source databases that likely cover different subsets of a real population. These databases are generally processed independently by their respective data owners into the databases <bold>D</bold><sup>&#x2032;</sup><sub>1</sub> to <bold>D</bold><sup>&#x2032;</sup><sub>n</sub>, and then linked and processed further (often by different organisations) into the linked database <bold><italic>D</italic></bold><sub>L</sub> to make it fit for the purpose of a research study</attrib>
</fig>
<p>Many data issues are due to humans being involved in the processes that generate population data, including the mistakes and choices people make, changing requirements, novel computing and data entry systems, limited resources and time, as well as decision making influenced by political or for economical reasons. Research managers, and policy and decision makers, often also assume any kind of question can be answered with highly accurate and unbiased results when analysing population databases [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>While the literature on general data quality is broad [<xref ref-type="bibr" rid="ref-26">26</xref>, <xref ref-type="bibr" rid="ref-27">27</xref>], given the wide-spread use of personal data at the level of populations it is surprising that only little published work seems to discuss data quality aspects specific to personal data [<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-28">28</xref>, <xref ref-type="bibr" rid="ref-29">29</xref>]. One reason is due to the perceived sensitivity of this kind of data: Population data are generally covered by privacy regulations, such as the European General Data Protection Regulation (GDPR) or the US Health Insurance Portability and Accountability Act (HIPAA) [<xref ref-type="bibr" rid="ref-15">15</xref>], and the processes and methods employed are often covered by confidentiality agreements. Furthermore, detailed data aspects are generally not included in scientific publications where the focus is on presenting the results obtained in a research study rather than the steps taken to obtain these results.</p>
</sec>
<sec>
<title>Misconceptions about population data</title>
<p>Following from <xref ref-type="fig" rid="fig-2">Figure 2</xref>, we categorise the identified misconceptions due to how population data are captured, how such data are processed, and how they are linked. We do not discuss any misconceptions related to the analysis of population data &#x2013; how to prevent pitfalls in statistical data analysis and machine learning has been discussed extensively elsewhere [<xref ref-type="bibr" rid="ref-30">30</xref>, <xref ref-type="bibr" rid="ref-31">31</xref>].</p>
<sec>
<title>Misconceptions due to data capturing</title>
<p>Misconceptions under this category can occur due to how information about people is captured. We refer to data capture as any processes and methods that convert information from a source into electronic format. This involves decisions about selecting or sampling individuals from an actual real population to be included into a population database, how to measure the characteristics of these people, and the methods employed to actually collect and record this information.</p>
<p>Many data capturing methods and processes involve humans who can make mistakes or behave in unexpected ways, or equipment that can malfunction or be misconfigured. Data capturing methods include manual data entry, optical character recognition, automatic speech recognition, and sensor readings (like biometrics from fingerprint readers and smart watches, or location traces from smart phones). While each of these methods can introduce specific data quality issues (such as keyboard typing or scanning errors), there are some common misconceptions about capturing population data.</p>
<sec>
<title>(1) A population database contains all individuals in a population</title>
<p>Even databases that are supposed to cover whole populations, such as taxation or census databases, very likely have subpopulations that are under-represented or absent (for example children or residents without a fixed address). Individuals who do not have a national health identifier number, like tourists or international students, and therefore are not eligible for government health services might be missing from population health databases. There are also always individuals who refuse to participate in government services for personal reasons, which can be influenced by their ethnicity, religion, or political beliefs.</p>
<p>The digital divide [<xref ref-type="bibr" rid="ref-32">32</xref>], the division of people who have access to digital services and media versus those who do not, likely also results in biased population databases. Younger and more affluent individuals are more likely included in databases collected using digital services compared to older people, migrants, or those with a lower socio-economic status.</p>
<p>Organisations might also refuse to provide their data for commercial reasons or due to confidentiality concerns. Not all patients will be included in a health database if, for example, certain private hospitals refuse to participate in a research study. This will likely introduce bias [<xref ref-type="bibr" rid="ref-33">33</xref>], since patients with higher income will more likely go to exclusive private hospitals.</p>
<p>The assumption that all individuals in a population are represented in a database leads to the illusion that it will be possible to identify any subpopulation of interest and explore this group of people, even if it is very small [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
</sec>
<sec>
<title>(2) The population covered in a database is well defined</title>
<p>The reasons for records about individuals to be included (or not) into a database are crucial to understanding the population covered in that database. Some population databases can be based on mandatory inclusion of individuals (think of government taxation or residency databases) while others are based on voluntary or self-selected inclusion (think of health databases such as registries where patients can opt-out when asked to provide their medical details).</p>
<p>The definitions and rules used to extract records about individuals into a population database might not be known to those who are processing and linking the database, and even less likely to the researchers who will be analysing it [<xref ref-type="bibr" rid="ref-18">18</xref>]. Furthermore, these rules might differ between organisations or change over time, or they might contain minor inconsistencies or mistakes such as ill-defined age or date ranges. For example, COVID-19 cases can be included into a database based on the date of symptom onset, collection of samples, or diagnosis [<xref ref-type="bibr" rid="ref-34">34</xref>], where each of these will result in different numbers of records being added to a database each day.</p>
</sec>
<sec>
<title>(3) Population databases contain complete information for all records in the database</title>
<p>Many population databases contain different pieces of information for different sets of records. This sparseness is because large databases are commonly generated by compiling different individual databases each only covering a part of a population, or by collecting records over time where changes in regulations and data capturing methods and processes can lead to different attributes (both QIDs and microdata) being collected. The resulting sparse patterns may prevent many statistical analyses of the data if not all relevant information has been captured.</p>
</sec>
<sec>
<title>(4) All records in a population database are within the scope of interest</title>
<p>Individuals might have left the population of interest for a research study because the criteria for inclusion in a database are no longer given. For example, a person may have died (although this might be of interest in itself) or have left the geographical area of a study. Because many organisations are not notified of an individual&#x2019;s death or relocation until sometime after the event (in some cases never), at any given time a population database will likely contain records that are outside the scope of interest. Including records about these people can affect research studies as well as operational systems.</p>
</sec>
<sec>
<title>(5) Each individual in a population is represented by a single record in a database</title>
<p>A common assumption is that population databases are generated and maintained as person-level databases, with one record per person in a population. However, it is not uncommon for a population database to contain duplicate records referring to the same person due to errors or variations in QID values. While data entry validation may prevent exact duplicates, fuzzy or approximate duplicates [<xref ref-type="bibr" rid="ref-22">22</xref>] might be missed by (automatic) checks. In real-world settings, the same person can therefore be registered multiple times at different institutions and their duplicate records are not being identified as referring to the same individual.</p>
<p>Some duplicates are very difficult to find, for example women who change their last name and address when they get married and therefore only their first name, gender, and place and date of birth stay the same. The flip side is that several people with highly similar personal details (similar QID values), such as twins who only have different first names, might not be recognised as two individuals but instead as duplicates.</p>
<p>Duplicate records are possible even if entity identifiers (such as social security numbers or patient identifiers) are available that should prevent multiple records by the same individual from being added to a population database. Due to human behaviour and errors such identifiers are not always provided or entered correctly. It might therefore be beneficial to apply deduplication (also known as internal linkage) [<xref ref-type="bibr" rid="ref-22">22</xref>] to a population database to identify any possible duplicate records that refer to the same person, and then handle such duplicates appropriately.</p>
<p>An example are people in a social security database who should have one record but were registered multiple times because they changed their names or addresses and might have forgotten their previous registration (and their unique identifier number), or they might be interested in multiple registrations to obtain social benefits more than once. Duplicate records have even been identified in domains where high data quality is crucial, such as in voter registration databases [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
</sec>
<sec>
<title>(6) Records in a population database always refer to real people</title>
<p>Real-world databases can contain records of people who never existed [<xref ref-type="bibr" rid="ref-36">36</xref>]. These can be records added to train data entry personnel or test software systems. Often these records are not removed from a database. While potentially easy to identify by humans (&#x2018;Tony Test&#x2019; living in &#x2018;Testville&#x2019;), these records are difficult to detect by data cleaning algorithms because they were designed to have the characteristics of real people and often contain high amounts of variations and errors.</p>
<p>In population databases collected from social media platforms, complete records might furthermore correspond to fake users, and increasingly even AI (artificial intelligence) bots that generate human-like content [<xref ref-type="bibr" rid="ref-17">17</xref>, <xref ref-type="bibr" rid="ref-18">18</xref>]. In one reported case of research fraud, records have been created in order to boost the size of a database being analysed in order to make research findings more convincing [<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
</sec>
<sec>
<title>(7) Errors in personal data are not intentional</title>
<p>There are social, cultural, as well as personal reasons why individuals would decide to provide incorrect personal details. These include fear of surveillance by governments, trying to prevent unsolicited advertisements from businesses, or simply the desire to keep sensitive personal data private. Fear of data breaches and how personal data are being (mis)used by organisations are clearly influencing the reluctance of individuals to provide their details unless deemed necessary [<xref ref-type="bibr" rid="ref-8">8</xref>, <xref ref-type="bibr" rid="ref-38">38</xref>]. In domains such as policing and criminal justice, faked QID values such as name aliases occur commonly as criminals try to hide their actual identities [<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
<p>Incorrectly provided data might only be modified slightly from a correct value (such as a date of birth a few days off), be changed completely (a different occupation given), not be provided at all (no value entered if an input field is not mandatory), or be made up (such as a telephone number in the form of &#x2018;1234 5678&#x2019;).</p>
<p>The decision to provide incorrect or withhold personal details is dependent upon the context in which this information is being collected. Generally, it is less likely for an individual to provide incorrect values (for non-crucial QIDs) on an official government form or when opening a bank account compared to when ordering a book in an online store. However, the opposite might be true in cases where an individual does not trust the institution that is collecting their data [<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
</sec>
<sec>
<title>(8) Certain personal details do not change over time</title>
<p>While some personal details, such as names and addresses, are known to change over time for many individuals, it is often assumed that others are fixed at birth. These include ethnic and gender identification, as well as place and country of birth. In many population databases, ethnic identification is self-reported, where the available categories depend upon how a society values different subpopulations. For example, the US Bureau of the Census has repeatedly changed the details of their policy regarding ethnic groups<xref ref-type="fn" rid="fn2"><sup>2</sup></xref>. Socially influenced attributes might be changed by individuals over time, a recent example being the Black Lives Matter movement which made many individuals become more proud to be of colour. It can therefore be problematic to use these values in the context of, for example, longitudinal data analysis or record linkage [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>It is even possible for values fixed at birth to change in the real world. An example is the Eastern German city of Chemnitz, which from 1953 until 1990 was named Karl-Marx-Stadt. Individuals born during that period have a country of birth (German Democratic Republic) and a place of birth that both do not exist anymore.</p>
</sec>
<sec>
<title>(9) Personal name variations are incorrect</title>
<p>People&#x2019;s names are a key component of QIDs in many population databases. Unlike with most general words, for many personal names there are multiple spelling variations (such as &#x2018;Gail&#x2019;, &#x2018;Gayle&#x2019;, and &#x2018;Gale&#x2019;), and all of them are correct [<xref ref-type="bibr" rid="ref-22">22</xref>]. When data are entered, for example over the telephone, differently sounding name variations might be recorded for the same individual due to mispronunciation or misunderstanding.</p>
<p>Furthermore, there are many cultural aspects of names, including different name orders and structures, ambiguous transliterations from non-roman into the roman alphabet, or name changes over time for religious reasons, to name a few. Name variations are a known problem when names are compared between records when linking databases [<xref ref-type="bibr" rid="ref-41">41</xref>]. Working with names can therefore be a challenging undertaking that requires expertise in the cultural and ethnic aspects of names [<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
</sec>
<sec>
<title>(10) Coding systems do not change over time</title>
<p>Categorical QID values and microdata are often coded using systems such as the International Standard Classification of Occupations (ISCO)<xref ref-type="fn" rid="fn3"><sup>3</sup></xref> or the International Classification of Diseases (ICD)<xref ref-type="fn" rid="fn4"><sup>4</sup></xref>, the latter currently in its eleventh revision. It is commonly assumed that such codes are fixed over time and unique in that a certain item, such as an occupation or disease, is only assigned one code, and that this assignment does not change. However, many coding systems are revised over time with new codes being added, outdated and unused codes being removed, and whole groups of codes being recoded (including codes being swapped). A database might therefore contain codes which are no longer valid. Furthermore, at a given time, different revisions of coding systems might occur in a population database, for example if the transition to a new version of a system is not conducted at the same time by all organisations that contribute to that database.</p>
<p>An example are the codes of the Australian Pharmaceutical Benefit Scheme (PBS) [<xref ref-type="bibr" rid="ref-43">43</xref>], where the antidepressant Venlafaxine had the code N06AE06 until 1995, when it was given the code N06AA22, which was then changed to N06AX16 in 1999. Using such codes to group or categorise records can therefore lead to wrong results of a research study if records have been collected over time.</p>
</sec>
<sec>
<title>(11) Data definitions are unambiguous</title>
<p>Many population databases contain information that is based on definitions such as how to categorise records or create categorical QID or microdata values. As with coding systems, data definitions can change over time, and they can also be interpreted differently. A recent example are the definitions for death or hospitalisations due to COVID-19 infections, where different US states used various definitions that resulted in databases that could not be used for comparative analysis [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>Unless metadata (see misconception 25 below) are available that clearly describe such definitions and their changes, it can be difficult to identify the effects of any changed definitions because any such change might have subtle effects on the characteristics of only some individuals in the population of interest for a research study.</p>
</sec>
<sec>
<title>(12) Temporal data aspects do not matter</title>
<p>Given the dynamic nature of personal details, the time and date when population data are captured and stored in a database can be crucial because differences in data lag can lead to inconsistent data that are not suitable for research studies [<xref ref-type="bibr" rid="ref-1">1</xref>, <xref ref-type="bibr" rid="ref-12">12</xref>]. If it takes different amounts of time for different organisations to capture data about the same events then clearly these data are not comparable, resulting in misreporting for example of the numbers of daily COVID-19 deaths [<xref ref-type="bibr" rid="ref-34">34</xref>], vaccination rates [<xref ref-type="bibr" rid="ref-1">1</xref>, <xref ref-type="bibr" rid="ref-44">44</xref>], or education levels of migrants [<xref ref-type="bibr" rid="ref-45">45</xref>].</p>
<p>Daily, weekly, monthly, or seasonal aspects can influence data measurements, as can events such as public holidays and religious festivities which likely only affect certain subpopulations. For example, daily reporting of new COVID-19 infections might be limited on weekends and the beginning of a week due to less testing and delayed laboratory diagnosis on weekends. Similar delays will happen during and after public holidays.</p>
<p>Data corrections are not uncommon, especially in applications where there is an urgent need to provide initial data as quickly as possible, for example to better understand a global pandemic [<xref ref-type="bibr" rid="ref-1">1</xref>, <xref ref-type="bibr" rid="ref-34">34</xref>]. Later updates and corrections of data might not be considered, leading to wrong conclusions of research studies.</p>
</sec>
<sec>
<title>(13) The meaning of data is always known</title>
<p>It is not uncommon for population databases to contain attributes that are not (well) documented. These can include codes without known meaning, irrelevant sequence numbers, or temporary values that have been added at some point in time for some specific purpose. If no documentation is available, database managers are generally reluctant to remove such attributes. As a result, spurious patterns might be detected if such attributes are included into a data analysis.</p>
</sec>
<sec>
<title>(14) Missing data have no meaning</title>
<p>Missing data are common in many databases [<xref ref-type="bibr" rid="ref-46">46</xref>]. They can lead to problems with data processing, linking, and analysis [<xref ref-type="bibr" rid="ref-15">15</xref>]. Missing data can occur at the level of missing records (no information is available about certain individuals in a population), missing attributes (some QIDs or microdata contain no data values for all records in a database), or missing QID or microdata values for individual records (specific missing attribute values) [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>There are different categories of missing data [<xref ref-type="bibr" rid="ref-46">46</xref>, <xref ref-type="bibr" rid="ref-47">47</xref>]. In some cases a missing value does not contain any valuable information, in others it can be the only correct value (children under a certain age should not have an occupation), or it can have multiple interpretations. A missing value for a question about religion in a census, for example, can mean an individual does not have a religion or they choose not to disclose it. Missing data can also occur in settings where resources are limited and therefore data entries had to be prioritised, such as in busy emergency departments [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>Care must therefore be taken when considering missing data. Removing attributes or even records with missing values, or imputing missing values [<xref ref-type="bibr" rid="ref-41">41</xref>, <xref ref-type="bibr" rid="ref-47">47</xref>], can result in errors and structural bias being introduced into a population database that can lead to incorrect outcomes of a research study.</p>
</sec>
<sec>
<title>(15) All records in a population database were captured using the same process</title>
<p>Since population databases are often collected over long periods of time and wide geographical areas, records are commonly generated or entered by a large number of staff, giving rise to different interpretations of data entry rules. For example, if an input field requires a mandatory value, humans will enter all kinds of unstandardised indicators for missingness, ranging from single symbols (like &#x2018;&#x2013;&#x2019; or &#x2018;.&#x2019;), acronyms (&#x2018;NA&#x2019; or &#x2018;MD&#x2019;), to texts explaining the missing data (like &#x2018;unknown&#x2019;). If a population database is compiled from independent organisations these different interpretations of data entry rules will cause the need for standardisation before analysis. Manual data entry such as typing can furthermore lead to different error characteristics between data entry personnel [<xref ref-type="bibr" rid="ref-15">15</xref>], resulting in subsets of records in a population database with different data quality.</p>
<p>Data might be captured at different temporal and spatial resolution, such as postcodes or city names only versus detailed street addresses. As a result, the characteristics of both QID and microdata values can differ between subsets of records in a population database, making their comparison and analysis challenging [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
</sec>
<sec>
<title>(16) Attribute values are correct and valid</title>
<p>Any data values captured, either by some form of sensor or manually entered into a database, can be subject to errors coming from equipment malfunction, human data entry (typing mistake), or cognitive mistakes (such as confusion about the data required or difficulties recalling correct information), or even malicious intent [<xref ref-type="bibr" rid="ref-18">18</xref>, <xref ref-type="bibr" rid="ref-27">27</xref>].</p>
<p>In the medical domain, manual typing errors, wrong interpretations of forms (think of handwritten prescriptions by doctors), entering values into the wrong input fields, or making mistakes interpreting instructions (when prescribing drugs) are commonly occurring mistakes, where rates ranging from 2 to 514 mistakes per 1,000 prescriptions have been reported [<xref ref-type="bibr" rid="ref-48">48</xref>].</p>
<p>While data validation tools can detect values outside the range or domain of what is valid (such as 31st of February) [<xref ref-type="bibr" rid="ref-27">27</xref>], without external validation it is generally not feasible to ascertain the correctness of any given value. If a patient is really 42 years old can only be validated if authoritative information (likely from an external database) about the patient&#x2019;s true age is available.</p>
<p>Furthermore, while individual QID values in a given record can each be valid, they might contradict each other. For example, a record with first name &#x2018;John&#x2019; and gender &#x2019;f&#x2019; likely contains one QID value that is incorrect. Many, but not all, such contradictions can be identified and corrected using appropriate edit constraints [<xref ref-type="bibr" rid="ref-41">41</xref>].</p>
</sec>
<sec>
<title>(17) Data values are in their correct attributes</title>
<p>Data entry personnel do not always enter values into the correct attribute. Many Asian and some Western names, for example, can be used interchangeably as first and last names, leading to misinterpretation. For example, &#x2018;Paul&#x2019;, &#x2018;Thomas&#x2019;, &#x2018;Chris&#x2019;, and &#x2018;Dennis&#x2019; are all used as first and last names. The ordering of how first and last names are written can also depend on the culture and origin of an individual.</p>
</sec>
<sec>
<title>(18) Data validation rules produce correct data</title>
<p>To ensure data of high quality, many data management systems contain rules that need to be fulfilled when data are being captured. For example, registering a new patient in a hospital requires both a valid address and a valid date of birth. In some cases, such as in emergency admissions, not all of this information will be known. Due to such data validation rules, default values are often used. A common example is the 1st of January being used for individuals with unknown day and month of birth. While these are valid, if not handled properly such defaults can result in skewed data distributions that can adversely affect research studies. Data entry personnel might also have ad-hoc rules they apply in order to bypass data entry requirements and to ensure any entered records fulfil all data validation steps.</p>
</sec>
<sec>
<title>(19) All relevant data have been captured</title>
<p>Because the primary purpose of most population databases is not their use for research studies, not all relevant information that is of importance for a given study might be available for all records in a database, or it might only be available in subsets of records. This can, for example, be due to changes in data entry requirements over time, or because data might have been withheld by the owner due to confidentiality concerns or for commercial reasons, or data might only be provided in aggregated or anonymised form. If a statistical model is generated from such data, a probable causal variable might be missed, since it is not captured at all.</p>
<p>Data that are not available are known as <italic>dark data</italic> [<xref ref-type="bibr" rid="ref-46">46</xref>], data we do not know about but that could be of interest to a research study. As a result, certain required or desired information might be missing for a given research study, making a given population database less useful or requiring the use of alternative data for that study.</p>
</sec>
<sec>
<title>(20) Population data provide the same answers as survey data</title>
<p>Population data, as captured from administrative or operational databases, are about what people are and what they do [<xref ref-type="bibr" rid="ref-10">10</xref>]. This is unlike survey data where commonly questions about attitudes, beliefs, expectations, or intentions are asked with the aim to understand the behaviour of people. Factual information about people can provide different answers compared to questions about what they claim to do, while inferring people&#x2019;s beliefs from their behaviour might not be possible.</p>
</sec>
<sec>
<title>(21) Population data are always of value</title>
<p>Both private and public sector organisations increasingly make databases publicly available to facilitate their analysis by researchers. However, many of these databases either lack metadata or context for them to be of use, or they are aggregated or anonymised due to privacy and confidentiality concerns [<xref ref-type="bibr" rid="ref-24">24</xref>]. A main reason for this is because past experiences have shown that sensitive personal information about individuals can sometimes be re-identified even from supposedly anonymised databases [<xref ref-type="bibr" rid="ref-38">38</xref>, <xref ref-type="bibr" rid="ref-49">49</xref>].</p>
<p>Population data without context are unlikely to be of use for research studies. A database of QID values (such as names and addresses) without any (or only limited) microdata is, by itself, of little value for research. Having the educational level of individuals in a database only becomes useful if this database can be linked with other data at the level of individuals. Furthermore, due to the dynamic nature of people&#x2019;s lives, population data become out of date quickly and therefore need to be updated regularly. Without adequate metadata (see misconception 25 below), context, useful detailed microdata, and regular updates, many publicly available databases are of little value for research.</p>
</sec>
</sec>
<sec>
<title>Misconceptions due to data processing</title>
<p>It is rare for population databases to be used for research without any processing being conducted. The organisation(s) that collect population data, and those that further aggregate, link (or otherwise integrate) such data, as well as the researchers who will analyse the data, all will likely apply some form of data processing [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>Processing can include data cleaning and standardisation, parsing of free text values, transformation of values, numerical normalisation, recoding into categories, imputation of missing values, and data aggregation. The use of different database management systems and data analysis software can furthermore result in data being reformatted internally before being stored and later extracted for further processing and analysis. Each component of a data pipeline can result in both explicit (user applied) as well as implicit (internally to software) data processing being conducted, leading to various misconceptions.</p>
<sec>
<title>(22) Data processing can be fully automated</title>
<p>Much of the processing of population data has to be conducted in an iterative fashion, where data exploration and profiling lead to a better understanding of a database which in turns helps to apply appropriate data processing techniques [<xref ref-type="bibr" rid="ref-27">27</xref>]. This process requires manual exploration, programming of data specific functionalities, domain expertise with regard to the provenance and content of a database, as well understanding of the final use of a database. Data processing is often the most time-consuming and resource intensive step of the overall data analytics pipeline, commonly requiring substantial domain as well as data expertise. In national statistical agencies, it has been reported that as much as 40% of resources are used on data processing [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Time and resource constraints might mean not all desired data processing can be accomplished. Manual editing and evaluation is also unlikely to be possible on large and complex population databases, and therefore compromises have to be made between data quality and timeliness for a database to be made available for research [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
</sec>
<sec>
<title>(23) Data processing is always correct</title>
<p>There are often multiple methods available to process data, for example to normalise numerical values, impute missing data, or standardise free-format text [<xref ref-type="bibr" rid="ref-26">26</xref>, <xref ref-type="bibr" rid="ref-41">41</xref>]. Converting &#x2018;dirty&#x2019; data into &#x2018;clean&#x2019; data can therefore result in incorrectly cleaned data [<xref ref-type="bibr" rid="ref-22">22</xref>]. Sometimes there is no single correct value for a given ambiguous input value. For example, within a street address, the abbreviation &#x2018;St&#x2019; can stand for either &#x2018;Street&#x2019; or &#x2018;Saint&#x2019; (as used in a town name like &#x2018;Saint Marys&#x2019;).</p>
<p>Given data processing commonly involves human efforts, mistakes in the use and configuration of software can lead to incorrect data processing, as can bugs in or the use of different or outdated versions of software. The use of unsuitable tools for a given project (such as spreadsheet software instead of a proper database management system or statistical analysis software) can furthermore results in mistakes when data are being processed. It has been reported [<xref ref-type="bibr" rid="ref-50">50</xref>] that on the 2nd October 2020 a total of 15,841 positive COVID-19 cases (around 20%) in England were missed because when recording daily cases an old file format of the Microsoft Excel spreadsheet software was used which allowed a maximum of 65,536 rows. Software features such as auto-completion and automatic spelling correction can furthermore lead to the wrong correction of unusual but valid data values that are not available in a dictionary.</p>
<p>As low quality data (possibly due to mistakes in data processing) are identified over time, improvements can be made to data processing methods that result in improved data quality. While this generally means that data quality can improve over time, changes in the actual data (data trends) as well as in data capturing might also mean that data processing again becomes less efficient, leading to lower data quality. Examples include data cleaning rules that have been correct in the past but might generate wrong results, such as when postcode boundaries are changing, or coding systems are revised.</p>
</sec>
<sec>
<title>(24) Aggregated data are sufficient for research</title>
<p>Highly aggregated data, for example at the level of states, counties, or large geographical units, are hardly of use for scientific research intending causal statements. Although results based on aggregated data might seem to be interesting, the number of possible alternative explanations for the same set of facts based on aggregated data is usually so large that no definitive conclusions are possible.</p>
<p>A major problem here is the <italic>ecological fallacy</italic>, describing the mistake of an aggregate relationship implying the same relationship for individuals [<xref ref-type="bibr" rid="ref-51">51</xref>]. For example, if increased mortality rates are observed in regions where vaccination rates are high, the false conclusion would be that vaccinated people have a higher probability to die. But actually the reverse might be true: People observing other people dying might be more willing to get vaccinated.</p>
<p>How data are aggregated depends on how aggregation functions are defined and interpreted. Weekly counts can, for example, be summed Monday to Sunday or alternatively Sunday to Saturday. Data that are aggregated inconsistently, or at different levels of aggregation, will unlikely be of use for any research studies (or only after additional data processing has been conducted).</p>
</sec>
<sec>
<title>(25) Metadata are correct, complete, and up-to-date</title>
<p>Metadata (also known as <italic>data dictionaries</italic>) are describing a database, how it has been created, populated, and its content captured and processed. Metadata include aspects such as the source, ownership, and provenance of a database, licensing and access limitations, costs, description of all attributes including their domains and any coding systems used, summary statistics and descriptions of data quality dimensions [<xref ref-type="bibr" rid="ref-15">15</xref>], as well as any data cleaning, imputation, editing, processing, transformation, aggregation, and linkage that was conducted on the source database(s) to obtain a given population database. Relevant documentation should be provided, including who conducted any data processing using what software (and which version of it), and containing a revision history of that documentation. Metadata are crucial to understand the actual structure, content, and quality of a database at hand.</p>
<p>Unfortunately, metadata are often not available, or they are incomplete, out of date, they need to be purchased, or can only be obtained through time-consuming approval processes [<xref ref-type="bibr" rid="ref-5">5</xref>]. A lack of metadata can lead to misunderstandings during data processing, linking and analysis, wasted time, misreporting of results, or can make a population database altogether useless [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
</sec>
</sec>
<sec>
<title>Misconceptions due to data linkage</title>
<p>Linking databases is generally based upon comparing the QID values of individuals, such as people&#x2019;s names, addresses, and other personal details (as illustrated in <xref ref-type="fig" rid="fig-1">Figure 1</xref>), to find records that refer to the same person [<xref ref-type="bibr" rid="ref-25">25</xref>, <xref ref-type="bibr" rid="ref-41">41</xref>]. These QID values, however, can contain errors, be missing, and they can change over time. This can lead to incorrect linkage results even when modern linkage methods are employed [<xref ref-type="bibr" rid="ref-52">52</xref>]. Linking databases can therefore be the source of various misconceptions about a linked data set. While all of the following misconceptions can occur when data from two sources are being linked (or even when duplicate records need to be identified within a single database [<xref ref-type="bibr" rid="ref-22">22</xref>]), in situations where records from multiple (more than two) sources have to be linked any of these misconceptions can become even more challenging and more difficult to deal with.</p>
<sec>
<title>(26) A linked data set corresponds to an actual population</title>
<p>Due to data quality issues and the record linkage technique(s) employed [<xref ref-type="bibr" rid="ref-22">22</xref>], a linked data set likely contains wrong links (Type I errors, two records referring to two different individuals were linked wrongly) while some true links have been missed (Type II errors, two records referring to the same person were not linked) [<xref ref-type="bibr" rid="ref-53">53</xref>]. The performance of most record linkage techniques can be controlled through parameters [<xref ref-type="bibr" rid="ref-52">52</xref>], allowing a trade-off between these two types of errors. Many linked data sets with different error characteristics can therefore be generated by changing parameter settings, where each generated data set provides an approximation of the actual population it is supposed to represent.</p>
</sec>
<sec>
<title>(27) Population databases represent the conditions of people at the same time</title>
<p>Data updates on individuals often occur at different points in time, usually when an event such as a medical condition occurs, or a data error is detected during a triggered data transaction such as a payment. In the German Social Security database, for example, education is entered when a record for a given person is newly created, but is not regularly updated afterwards. Therefore, highly trained professionals might have a record stating a low educational level because they were pupils at the time of their first paid job. Similar issues have been identified in Sweden for the education data of migrants, where for different subpopulations their educational levels are updated at varying rates [<xref ref-type="bibr" rid="ref-45">45</xref>]. As a result, records that represent the same individual can have different values (in both QID and microdata attributes) across the databases being linked. The assumption that the QID values of all records in the databases being linked are up-to-date might therefore not be correct, and outdated information can lead to wrong linkage results [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>Data corrections and updates can furthermore occur when incorrect historical data are being discovered and errors rectified [<xref ref-type="bibr" rid="ref-1">1</xref>]. Unless it is possible to re-conduct a linkage, which is unlikely for many research studies due to the efforts and costs involved in such a process, a linked data set might contain errors which have influenced the conclusions of the original study.</p>
</sec>
<sec>
<title>(28) A linked data set contains no duplicates</title>
<p>When linking databases, pairs or groups of records that refer to the same individual might not be linked correctly (missed true links, Type II errors). One reason for this to occur is if a wrong entity identifier has been assigned to an individual (by accident or on purpose), as has been reported even in voter registration databases [<xref ref-type="bibr" rid="ref-35">35</xref>]. If a linkage requires agreement on such unique identifier values, then two records with different values in the unique identifier will not be linked even if they have many highly similar (or even agreeing) QID values. Another reason is if crucial QID values of an individual have changed over time, such as both their name and address details, or are missing, resulting in two records that are not similar with each other [<xref ref-type="bibr" rid="ref-22">22</xref>]. Therefore, many linked data sets do contain more than one record for some individuals in a population.</p>
</sec>
<sec>
<title>(29) A linked data set is unbiased</title>
<p>Linkage errors generally do not occur at random [<xref ref-type="bibr" rid="ref-5">5</xref>], rather they depend upon the characteristics of the actual QID values of individuals, which can differ in diverse subpopulations. Examples include name structures of migrants that are different from the traditional Western standard of first, middle, and last name formats [<xref ref-type="bibr" rid="ref-54">54</xref>], or different rates of mobility (address changes) for young versus older people. As a result, there can be structural bias in a linked data set in subpopulations defined by ethnic or social categories, age, or gender (for example if women are more likely to change their names compared to men when they get married) [<xref ref-type="bibr" rid="ref-55">55</xref>].</p>
<p>Recent work has also shown that even small amounts of linkage error can result in large effects on false negative (Type II) error rates in research studies. This is especially the case with small sample sizes that can occur with the rare effects that are often sought to be identified via record linkage from large population databases [<xref ref-type="bibr" rid="ref-56">56</xref>]. If the aim of a study is to analyse certain (potentially small) subpopulations, or compare, for example, health aspects between subpopulations, then a careful assessment of the potential bias introduced via record linkage is of crucial importance [<xref ref-type="bibr" rid="ref-53">53</xref>].</p>
</sec>
<sec>
<title>(30) Attribute values in linked records are correct</title>
<p>Given a supposedly correct link, there might be contradicting attribute values in the corresponding records. Data fusion is the process of resolving such inconsistencies, where often a decision needs to be made which of many available fusion operations to apply [<xref ref-type="bibr" rid="ref-57">57</xref>]. Even if the links made between records are correct, how records are fused or merged can therefore introduce errors both into QID values as well as microdata.</p>
<p>For example, assume three records that refer to the same person have been linked correctly, where each record contains a different salary value. Should the average, median, minimum, maximum, or the most recent of these three salary values be used for the fused record of this individual? How data fusion is conducted needs to be discussed with the researchers who will be analysing a linked data set because depending upon the fusion operation applied substantially different outcomes will potentially be obtained.</p>
</sec>
<sec>
<title>(31) Linkage error rates are independent of database size</title>
<p>The QID values used to link records can be shared by multiple individuals, potentially thousands in the case of city and town names or with popular first and last names. Therefore, when larger databases are being linked, the number of record pairs with the same QID values likely increases, resulting in more highly similar pairs. Making correct classification becomes increasingly challenging as there are more possibly matching pairs. Generalising linkage quality results obtained on small data sets in published studies to much larger population sized real-world databases can therefore be dangerous.</p>
</sec>
<sec>
<title>(32) Modern record linkage techniques can handle databases of any sizes</title>
<p>Many researchers, especially in the computer science and statistical domains, who develop record linkage techniques do not have access to large real-world databases due to the sensitive nature of population data. As a result, novel linkage techniques are often evaluated on small public benchmark data sets or on synthetically generated data [<xref ref-type="bibr" rid="ref-15">15</xref>]. While error rates for linkages obtained on such data sets can provide evidence of the superiority of a novel technique over existing methods, assuming that this new technique will produce comparable high quality linkage results on larger real-world databases is not guaranteed.</p>
</sec>
<sec>
<title>(33) Linkage techniques and their settings are easily transferrable</title>
<p>If a linkage method together with its parameter settings (for example how blocking is conducted, how values are compared, and how a classification threshold is set) has been successful deployed in a given linkage project, this does not mean that the same method and settings will provide comparable high linkage quality results on a different linkage project. For each linkage project, different methods and corresponding parameter settings will need to be established. Furthermore, the same holds even when linking large disparate population databases, where for different subpopulations different optimal parameter settings (such as classification thresholds) will need to be identified. Finally, repeated linkages over time, for example a yearly update, may also require different parameter settings.</p>
</sec>
</sec>
</sec>
<sec>
<title>Conclusions and recommendations</title>
<p>Due to misconceptions such as the ones we have discussed, the much hyped promise of big data requires some careful considerations when personal data at the level of populations are used for research studies or decision making. Given population data are increasingly used in many domains of science, as illustrated in <xref ref-type="fig" rid="fig-2">Figure 2</xref>, researchers will potentially have less and less control over the quality of the data they are using for their studies and any processing done on these data [<xref ref-type="bibr" rid="ref-19">19</xref>]. They likely will also have only limited information about the provenance and other metadata that is needed to fully understand the characteristics and quality of their data. Because population data are commonly sourced from organisations other than where they are being analysed [<xref ref-type="bibr" rid="ref-10">10</xref>, <xref ref-type="bibr" rid="ref-18">18</xref>], these limitations are inherent to this kind of data.</p>
<p>There are no (simple) technical solutions to detect and correct many of the misconceptions we have discussed. What is required is heightened awareness by anybody working with population data. While our list of misconceptions is unlikely to be exhaustive, our aim was to show that there is a broad range of issues that can lead to misconceptions. The following recommendations might help to recognise and overcome such potential misconceptions.</p>
<list list-type="simple">
<list-item><label>&#x2013;</label><p>If possible, data scientists and researchers should aim to get involved in the capturing, processing, and linking of any data they plan to use for their research. This involves discussions with database owners about what data to collect in what format, how to ensure high quality of these data, and that adequate metadata are collected. It also means proper planning and designing of the information systems that are required for data capture, processing and linkage, and their adequate support as well as updates over time. It is vital to have the involvement of data scientists and data policy managers in these processes.</p></list-item>
<list-item><label>&#x2013;</label><p>If at all possible, data scientists and IT personnel who are processing and linking population data need to work in close collaboration with the researchers who will conduct the actual analysis of these data. Both technical and strategic aspects of a project that involves population data should ideally be discussed with the analysts, data scientists, database managers, developers, project managers, as well as the owners of the population databases being used, processed, and linked. Forming multi-disciplinary teams with members skilled in data science, statistics, domain expertise, as well as &#x2018;business&#x2019; aspects of research [<xref ref-type="bibr" rid="ref-5">5</xref>], is crucial for successful projects that rely upon population data. Interaction between data and domain experts might mean that a project based on population data becomes an iterative endeavour where data might have to be recaptured, reprocessed, and relinked until they are suitable for a research study.</p></list-item>
<list-item><label>&#x2013;</label><p>Cross-disciplinary training should be aimed at improving complementary skills [<xref ref-type="bibr" rid="ref-5">5</xref>]. Having data scientists who also have domain specific expertise will be highly valuable in any project involving population data. Equally crucial is for any researcher, no matter what their domain, to understand how modern data processing, record linkage, and data analytics methods work, and how these methods might introduce bias and errors into the data they are using for their research studies. Training in data exploration and data cleaning methods as well as data quality issues should be part of any degree that deals with data, including statistics, quantitative social science, computer science, and public health.</p></list-item>
<list-item><label>&#x2013;</label><p>While extensive methodologies about how to deal with uncertainties, bias, and data quality in surveys have been developed [<xref ref-type="bibr" rid="ref-58">58</xref>, <xref ref-type="bibr" rid="ref-59">59</xref>], there is a lack of corresponding rigorous methods that can be employed on large population databases. The Big data paradox [<xref ref-type="bibr" rid="ref-60">60</xref>], the illusion that large databases automatically mean valid results, requires new statistical techniques to be developed. While certain data quality issues can be identified (and potentially corrected) automatically [<xref ref-type="bibr" rid="ref-27">27</xref>], novel data exploration methods are needed to identify more subtle data issues where traditional methods are inadequate.</p></list-item>
<list-item><label>&#x2013;</label><p>A crucial aspect is to have detailed metadata about a population database available, including how the database was captured, and any processing and linkage applied to it. All relevant data definitions need to be described, and information about all sources and types of uncertainties need to be collected [<xref ref-type="bibr" rid="ref-10">10</xref>]. Detailed data profiling and exploration should be conducted by researchers before a population database is being analysed so any unexpected characteristics in their data can be identified.</p></list-item>
<list-item><label>&#x2013;</label><p>Existing guidelines and checklists, such as RECORD [<xref ref-type="bibr" rid="ref-61">61</xref>] and GUILD [<xref ref-type="bibr" rid="ref-62">62</xref>], should be employed and adapted to other research domains. Frameworks such as the Big Data Total Error method [<xref ref-type="bibr" rid="ref-18">18</xref>] can be adapted for population data to better characterise errors in such data. Furthermore, data management principles such as FAIR (Findable, Accessible, Interoperable, Reusable) [<xref ref-type="bibr" rid="ref-63">63</xref>] should be adhered to, although in some situations the sensitive nature of personal data [<xref ref-type="bibr" rid="ref-15">15</xref>] might limit or prevent such principles from being applied. In such situations, at least metadata and any software used in a study should be made public in an open research repository. Following these principles, guidelines, and checklists will allow data scientists and researchers to highlight to data custodians that having access to metadata would be highly beneficial for their work.</p></list-item>
<list-item><label>&#x2013;</label><p>The lack of publications that describe practical challenges when dealing with population data can result in the misconceptions we have discussed here. We therefore encourage increased publication of data issues and the sharing of experiences with the scientific community about lessons learnt, as well as best practice approaches being implemented when dealing with population data.</p></list-item>
</list>
<p>We have discussed some aspects in modern scientific processes that are rarely considered when population data are being used for research studies or decision making. Since good data management is a key aspect of good science [<xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-63">63</xref>], it is vital for anybody who uses population data to be aware of underlying assumptions concerning this kind of data. We hope the misconceptions and recommendations given here will help to identify and prevent misleading conclusions and poor real-world decisions, making population data the new oil of the big data era.</p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>We like to thank S. Bender, S. Redlich, J. Reinhold, and C. Nanayakkara for their critical and helpful comments, A. Pl&#x00F6;ger for help with producing the figures, and S. Weiand for technical support. P. Christen likes to acknowledge the support of the University of Leipzig and ScaDS.AI, Germany, where parts of this work was conducted while he was funded by the Leibniz Visiting Professorship. P. Christen also gratefully acknowledges the support of the UK Economic and Social Research Council (ESRC), ES/W010321/1. The work by R. Schnell was supported by the Deutsche Forschungsgemeinschaft Grant 407023611. We finally like to thank the two anonymous reviewers whose comments have helped to improve our work.</p>
</ack>
<sec>
<title>Ethics statement</title>
<p>No ethics approval was required for this study because no actual data was involved.</p>
</sec>
<fn-group>
<fn id="fn1"><label>1</label><p>According to Merriam Webster (<uri>https://www.merriam-webster.com</uri>), a myth is a &#x201C;popular belief or tradition that has grown up around something or someone&#x201D;, while a misconception is &#x201C;a wrong or inaccurate idea or conception&#x201D;. For brevity, throughout the paper we will only use misconception.</p></fn>
<fn id="fn2"><label>2</label><p><uri>https://www.census.gov/about/our-research/race-ethnicity.html</uri>.</p></fn>
<fn id="fn3"><label>3</label><p><uri>https://www.ilo.org/public/english/bureau/stat/isco/</uri>.</p></fn>
<fn id="fn4"><label>4</label><p><uri>https://www.who.int/standards/classifications/classification-of-diseases</uri>.</p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="ref-1"><label>1</label><mixed-citation publication-type="journal"><string-name><given-names>Valerie C</given-names> <surname>Bradley</surname></string-name>, <string-name><given-names>Shiro</given-names> <surname>Kuriwaki</surname></string-name>, <string-name><given-names>Michael</given-names> <surname>Isakov</surname></string-name>, <string-name><given-names>Dino</given-names> <surname>Sejdinovic</surname></string-name>, <string-name><given-names>Xiao-Li</given-names> <surname>Meng</surname></string-name>, and <string-name><given-names>Seth</given-names> <surname>Flaxman</surname></string-name>. <article-title>Unrepresentative big surveys significantly overestimated US vaccine uptake</article-title>. <source><italic>Nature</italic></source> <volume>600</volume>, <fpage>695</fpage>&#x2013;<lpage>700</lpage> (<year>2021</year>). <pub-id pub-id-type="doi">10.1038/s41586-021-04198-4</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>2</label><mixed-citation publication-type="journal"><string-name><given-names>Liran</given-names> <surname>Einav</surname></string-name> and <string-name><given-names>Jonathan</given-names> <surname>Levin</surname></string-name>. <article-title>Economics in the age of Big data</article-title>. <source><italic>Science</italic></source>, <volume>346</volume>(<issue>6210</issue>), <year>2014</year>. <pub-id pub-id-type="doi">10.1126/science.1243089</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>3</label><mixed-citation publication-type="book"><string-name><given-names>Ian</given-names> <surname>Foster</surname></string-name>, <string-name><given-names>Rayid</given-names> <surname>Ghani</surname></string-name>, <string-name><given-names>Ron S</given-names> <surname>Jarmin</surname></string-name>, <string-name><given-names>Frauke</given-names> <surname>Kreuter</surname></string-name>, and <string-name><given-names>Julia</given-names> <surname>Lane</surname></string-name>, editors. <source><italic>Big Data and Social Science</italic></source>. <publisher-loc>CRC Press</publisher-loc>, <publisher-name>Boca Raton</publisher-name>, <year>2017</year>. <pub-id pub-id-type="doi">10.1201/9781315368238</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>4</label><mixed-citation publication-type="journal"><string-name><given-names>Roxanne</given-names> <surname>Connelly</surname></string-name>, <string-name><given-names>Christopher J</given-names> <surname>Playford</surname></string-name>, <string-name><given-names>Vernon</given-names> <surname>Gayle</surname></string-name>, and <string-name><given-names>Chris</given-names> <surname>Dibben</surname></string-name>. <article-title>The role of administrative data in the Big data revolution in social science research</article-title>. <source><italic>Social Science Research</italic></source>, <volume>59</volume>(Supplement <supplement>C</supplement>):<fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2016</year>. <pub-id pub-id-type="doi">10.1016/j.ssresearch.2016.04.015</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>5</label><mixed-citation publication-type="journal"><string-name><given-names>Louisa</given-names> <surname>Jorm</surname></string-name>. <article-title>Routinely collected data as a strategic resource for research: priorities for methods and workforce</article-title>. <source><italic>Public Health Research Practice</italic></source>, <volume>25</volume> (<issue>4</issue>), <year>2015</year>. <pub-id pub-id-type="doi">10.17061/phrp2541540</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>6</label><mixed-citation publication-type="journal"><string-name><given-names>Nathaniel D</given-names> <surname>Porter</surname></string-name>, <string-name><given-names>Ashton M</given-names> <surname>Verdery</surname></string-name>, and <string-name><given-names>S Michael</given-names> <surname>Gaddis</surname></string-name>. <article-title>Enhancing Big data in the social sciences with crowdsourcing: Data augmentation practices, techniques, and opportunities</article-title>. <source><italic>PloS one</italic></source>, <volume>15</volume>(<issue>6</issue>):<fpage>e0233154</fpage>, <year>2020</year>. <pub-id pub-id-type="doi">10.1371/journal.pone.0233154</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>7</label><mixed-citation publication-type="journal"><string-name><given-names>Susan</given-names> <surname>Athey</surname></string-name>. <article-title>Beyond prediction: Using Big data for policy problems</article-title>. <source><italic>Science</italic></source>, <volume>355</volume>(<issue>6324</issue>):<fpage>483</fpage>&#x2013;<lpage>485</lpage>, <year>2017</year>. <pub-id pub-id-type="doi">10.1126/science.aal4321</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>8</label><mixed-citation publication-type="journal"><string-name><given-names>Jim</given-names> <surname>Isaak</surname></string-name> and <string-name><given-names>Mina J</given-names> <surname>Hanna</surname></string-name>. <article-title>User data privacy: Facebook, Cambridge Analytica, and privacy protection</article-title>. <source><italic>IEEE Computer</italic></source>, <volume>51</volume>(<issue>8</issue>):<fpage>56</fpage>&#x2013;<lpage>59</lpage>, <year>2018</year>. <pub-id pub-id-type="doi">10.1109/MC.2018.3191268</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>9</label><mixed-citation publication-type="book"><string-name><given-names>Foster</given-names> <surname>Provost</surname></string-name> and <string-name><given-names>Tom</given-names> <surname>Fawcett</surname></string-name>. <source><italic>Data Science for Business: What you need to know about Data Mining and Data-Analytic Thinking</italic></source>. <publisher-loc>O&#x2019;Reilly Media, Inc.</publisher-loc>, <year>2013</year>. <uri>https://learning.oreilly.com/library/view/data-science-for/9781449374273/</uri>.</mixed-citation></ref>
<ref id="ref-10"><label>10</label><mixed-citation publication-type="journal"><string-name><given-names>David J</given-names> <surname>Hand</surname></string-name>. <article-title>Statistical challenges of administrative and transaction data</article-title>. <source><italic>Journal of the Royal Statistical Society: Series A (Statistics in Society)</italic></source>, <volume>181</volume> (<issue>3</issue>):<fpage>555</fpage>&#x2013;<lpage>605</lpage>, <year>2018</year>. <pub-id pub-id-type="doi">10.1111/rssa.12315</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>11</label><mixed-citation publication-type="journal"><string-name><given-names>Valerie</given-names> <surname>Braithwaite</surname></string-name>. <article-title>Beyond the bubble that is Robodebt: How governments that lose integrity threaten democracy</article-title>. <source><italic>Australian Journal of Social Issues</italic></source>, <volume>55</volume>(<issue>3</issue>):<fpage>242</fpage>&#x2013;<lpage>259</lpage>, <year>2020</year>. <pub-id pub-id-type="doi">10.1002/ajs4.122</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12</label><mixed-citation publication-type="journal"><string-name><given-names>Stephanie E</given-names> <surname>Galaitsi</surname></string-name>, <string-name><given-names>Jeffrey C</given-names> <surname>Cegan</surname></string-name>, <string-name><given-names>Kaitlin</given-names> <surname>Volk</surname></string-name>, <string-name><given-names>Matthew</given-names> <surname>Joyner</surname></string-name>, <string-name><given-names>Benjamin D</given-names> <surname>Trump</surname></string-name>, and <string-name><given-names>Igor</given-names> <surname>Linkov</surname></string-name>. <article-title>The challenges of data usage for the United States&#x2019; COVID-19 response</article-title>. <source><italic>International Journal of Information Management</italic></source>, <volume>59</volume>:<issue>102352</issue>, <year>2021</year>. <pub-id pub-id-type="doi">10.1016/j.ijinfomgt.2021.102352</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>13</label><mixed-citation publication-type="journal"><string-name><given-names>Sarah</given-names> <surname>Giest</surname></string-name> and <string-name><given-names>Annemarie</given-names> <surname>Samuels</surname></string-name>. <article-title>&#x2018;For good measure&#x2019;: data gaps in a Big data world</article-title>. <source><italic>Policy Sciences</italic></source>, <volume>53</volume>(<issue>3</issue>):<fpage>559</fpage>&#x2013;<lpage>569</lpage>, <year>2020</year>. <pub-id pub-id-type="doi">10.1007/s11077-020-09384-1</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14</label><mixed-citation publication-type="journal"><string-name><given-names>Kimberlyn M</given-names> <surname>McGrail</surname></string-name>, <string-name><given-names>Kerina</given-names> <surname>Jones</surname></string-name>, <string-name><given-names>Ashley</given-names> <surname>Akbari</surname></string-name>, <string-name><given-names>Tellen D</given-names> <surname>Bennett</surname></string-name>, <string-name><given-names>Andy</given-names> <surname>Boyd</surname></string-name>, <etal>et al</etal>. <article-title>A position statement on population data science: The science of data about people</article-title>. <source><italic>International Journal of Population Data Science</italic></source>, <volume>3</volume>(<issue>1</issue>), <year>2018</year>. <pub-id pub-id-type="doi">10.23889/ijpds.v3i1.415</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>15</label><mixed-citation publication-type="book"><string-name><given-names>Peter</given-names> <surname>Christen</surname></string-name>, <string-name><given-names>Thilina</given-names> <surname>Ranbaduge</surname></string-name>, and <string-name><given-names>Rainer</given-names> <surname>Schnell</surname></string-name>. <source><italic>Linking Sensitive Data</italic></source>. <publisher-name>Springer</publisher-name>, <publisher-loc>Heidelberg</publisher-loc>, <year>2020</year>. <pub-id pub-id-type="doi">10.1007/978-3-030-59706-1</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>16</label><mixed-citation publication-type="journal"><string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Dibben</surname></string-name>, <string-name><given-names>James</given-names> <surname>Boyd</surname></string-name>, <string-name><given-names>Anders</given-names> <surname>Hjern</surname></string-name>, <string-name><given-names>Mahmoud</given-names> <surname>Azimaee</surname></string-name>, <string-name><given-names>Mauricio L</given-names> <surname>Barreto</surname></string-name>, and <string-name><given-names>Harvey</given-names> <surname>Goldstein</surname></string-name>. <article-title>Challenges in administrative data linkage for research</article-title>. <source><italic>Big Data and Society</italic></source>, <volume>4</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2017</year>. <pub-id pub-id-type="doi">10.1177/2053951717745678</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>17</label><mixed-citation publication-type="book"><string-name><given-names>Florian</given-names> <surname>Keusch</surname></string-name> and <string-name><given-names>Frauke</given-names> <surname>Kreuter</surname></string-name>. <chapter-title>Digital trace data: Modes of data collection, applications, and errors at a glance</chapter-title>. In <source><italic>Handbook of Computational Social Science, Vol 1</italic></source>, pages <fpage>100</fpage>&#x2013;<lpage>118</lpage>. <publisher-name>Taylor and Francis</publisher-name>, <year>2021</year>. <pub-id pub-id-type="doi">10.4324/9781003024583-8</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>18</label><mixed-citation publication-type="book"><string-name><given-names>Paul</given-names> <surname>Biemer</surname></string-name>. <chapter-title>Errors and inference</chapter-title>. In: <string-name><given-names>Ian</given-names> <surname>Foster</surname></string-name>, <string-name><given-names>Rayid</given-names> <surname>Ghani</surname></string-name>, <string-name><given-names>Ron S</given-names> <surname>Jarmin</surname></string-name>, <string-name><given-names>Frauke</given-names> <surname>Kreuter</surname></string-name>, and <string-name><given-names>Julia</given-names> <surname>Lane</surname></string-name>, editors, <source><italic>Big Data and Social Science</italic></source>, chapter 10, pages <fpage>265</fpage>&#x2013;<lpage>297</lpage>. <publisher-loc>CRC Press</publisher-loc>, <publisher-name>Boca Raton</publisher-name>, <year>2017</year>. <pub-id pub-id-type="doi">10.1201/9781315368238</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>19</label><mixed-citation publication-type="journal"><string-name><given-names>Andrew W</given-names> <surname>Brown</surname></string-name>, <string-name><given-names>Kathryn A</given-names> <surname>Kaiser</surname></string-name>, and <string-name><given-names>David B</given-names> <surname>Allison</surname></string-name>. <article-title>Issues with data and analyses: Errors, underlying themes, and potential solutions</article-title>. <source><italic>Proceedings of the National Academy of Sciences</italic></source>, <volume>115</volume>(<issue>11</issue>):<fpage>2563</fpage>&#x2013;<lpage>2570</lpage>, <year>2018</year>. <pub-id pub-id-type="doi">10.1073/pnas.1708279115</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>20</label><mixed-citation publication-type="journal"><string-name><given-names>Alessandro</given-names> <surname>Acquisti</surname></string-name>, <string-name><given-names>Laura</given-names> <surname>Brandimarte</surname></string-name>, and <string-name><given-names>George</given-names> <surname>Loewenstein</surname></string-name>. <article-title>Privacy and human behavior in the age of information</article-title>. <source><italic>Science</italic></source>, <volume>347</volume>(<issue>6221</issue>):<fpage>509</fpage>&#x2013;<lpage>514</lpage>, <year>2015</year>. <pub-id pub-id-type="doi">10.1126/science.aaa1465</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21</label><mixed-citation publication-type="journal"><string-name><given-names>Eric</given-names> <surname>Horvitz</surname></string-name> and <string-name><given-names>Deirdre</given-names> <surname>Mulligan</surname></string-name>. <article-title>Data, privacy, and the greater good</article-title>. <source><italic>Science</italic></source>, <volume>349</volume>(<issue>6245</issue>):<fpage>253</fpage>&#x2013;<lpage>255</lpage>, <year>2015</year>. <pub-id pub-id-type="doi">10.1126/science.aac4520</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>22</label><mixed-citation publication-type="book"><string-name><given-names>Peter</given-names> <surname>Christen</surname></string-name>. <source><italic>Data Matching</italic></source>. <publisher-name>Springer</publisher-name>, <publisher-loc>Heidelberg</publisher-loc>, <year>2012</year>. <pub-id pub-id-type="doi">10.1007/978-3-642-31164-2</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>23</label><mixed-citation publication-type="book"><string-name><given-names>George</given-names> <surname>Duncan</surname></string-name>, <string-name><given-names>Mark</given-names> <surname>Elliot</surname></string-name>, and <string-name><given-names>Juan-Jos&#x00E9;</given-names> <surname>Salazar-Gonz&#x00E1;lez</surname></string-name>. <source><italic>Statistical Confidentiality: Principles and Practice</italic></source>. <publisher-name>Springer</publisher-name>, <publisher-loc>New York</publisher-loc>, <year>2011</year>. <pub-id pub-id-type="doi">10.1007/978-1-4419-7802-8</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>24</label><mixed-citation publication-type="book"><string-name><given-names>Mark</given-names> <surname>Elliot</surname></string-name>, <string-name><given-names>Elaine</given-names> <surname>Mackey</surname></string-name>, and <string-name><given-names>Kieron</given-names> <surname>O&#x2019;Hara</surname></string-name>. <source><italic>The Anonymisation Decision-making Framework 2nd Edition: European Practitioners&#x2019; Guide</italic></source>. <publisher-name>UKAN Manchester</publisher-name>, <year>2020</year>. <uri>https://msrbcel.files.wordpress.com/2020/11/adf-2nd-edition-1.pdf</uri>.</mixed-citation></ref>
<ref id="ref-25"><label>25</label><mixed-citation publication-type="book"><string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>, <string-name><given-names>Harvey</given-names> <surname>Goldstein</surname></string-name>, and <string-name><given-names>Chris</given-names> <surname>Dibben</surname></string-name>. <source><italic>Methodological Developments in Data Linkage</italic></source>. <publisher-name>John Wiley and Sons</publisher-name>, <year>2015</year>. <pub-id pub-id-type="doi">10.1002/9781119072454</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>26</label><mixed-citation publication-type="book"><string-name><given-names>Carlo</given-names> <surname>Batini</surname></string-name> and <string-name><given-names>Monica</given-names> <surname>Scannapieco</surname></string-name>. <source><italic>Data and Information Quality</italic></source>. <publisher-name>Springer</publisher-name>, <publisher-loc>Heidelberg</publisher-loc>, <year>2016</year>. <pub-id pub-id-type="doi">10.1007/978-3-319-24106-7</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>27</label><mixed-citation publication-type="journal"><string-name><given-names>Won</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>Byoung-Ju</given-names> <surname>Choi</surname></string-name>, <string-name><given-names>Eui-Kyeong</given-names> <surname>Hong</surname></string-name>, <string-name><given-names>Soo-Kyung</given-names> <surname>Kim</surname></string-name>, and <string-name><given-names>Doheon</given-names> <surname>Lee</surname></string-name>. <article-title>A taxonomy of dirty data</article-title>. <source><italic>Data Mining and Knowledge Discovery</italic></source>, <volume>7</volume>(<issue>1</issue>):<fpage>81</fpage>&#x2013;<lpage>99</lpage>, <year>2003</year>. <pub-id pub-id-type="doi">10.1023/A:1021564703268</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>28</label><mixed-citation publication-type="journal"><string-name><given-names>Mark</given-names> <surname>Smith</surname></string-name>, <string-name><given-names>Lisa M</given-names> <surname>Lix</surname></string-name>, <string-name><given-names>Mahmoud</given-names> <surname>Azimaee</surname></string-name>, <string-name><given-names>Jennifer E</given-names> <surname>Enns</surname></string-name>, <string-name><given-names>Justine</given-names> <surname>Orr</surname></string-name>, <string-name><given-names>Say</given-names> <surname>Hong</surname></string-name>, and <string-name><given-names>Leslie L</given-names> <surname>Roos</surname></string-name>. <article-title>Assessing the quality of administrative data for research: a framework from the Manitoba Centre for Health Policy</article-title>. <source><italic>Journal of the American Medical Informatics Association</italic></source>, <volume>25</volume>(<issue>3</issue>):<fpage>224</fpage>&#x2013;<lpage>229</lpage>, <year>2018</year>. <pub-id pub-id-type="doi">10.1093/jamia/ocx078</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>29</label><mixed-citation publication-type="journal"><string-name><given-names>Mihnea</given-names> <surname>Tufi&#x015F;</surname></string-name> and <string-name><given-names>Ludovico</given-names> <surname>Boratto</surname></string-name>. <article-title>Toward a complete data valuation process. challenges of personal data</article-title>. <source><italic>ACM Journal of Data and Information Quality</italic></source>, <volume>13</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2021</year>. <pub-id pub-id-type="doi">10.1145/344726</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>30</label><mixed-citation publication-type="book"><string-name><given-names>Trevor</given-names> <surname>Hastie</surname></string-name>, <string-name><given-names>Robert</given-names> <surname>Tibshirani</surname></string-name>, and <string-name><given-names>Jerome</given-names> <surname>Friedman</surname></string-name>. <source><italic>The Elements of Statistical Learning</italic></source>. <publisher-name>Springer</publisher-name>, <publisher-loc>New York</publisher-loc>, <edition>2</edition> edition, <year>2009</year>. <pub-id pub-id-type="doi">10.1007/978-0-387-84858-7</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>31</label><mixed-citation publication-type="journal"><string-name><given-names>Patrick</given-names> <surname>Riley</surname></string-name>. <article-title>Three pitfalls to avoid in machine learning</article-title>. <source><italic>Nature</italic></source>, <volume>572</volume> (<issue>7767</issue>):<fpage>27</fpage>&#x2013;<lpage>29</lpage>, <year>2019</year>. <pub-id pub-id-type="doi">10.1038/d41586-019-02307-y</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>32</label><mixed-citation publication-type="book"><string-name><given-names>Jan Van</given-names> <surname>Dijk</surname></string-name>. <source><italic>The Digital Divide</italic></source>. <publisher-name>John Wiley and Sons</publisher-name>, <publisher-loc>Cambridge, UK</publisher-loc>, <year>2020</year>. <uri>https://www.wiley.com/en-us/The+Digital+Divide-p-9781509534463</uri>.</mixed-citation></ref>
<ref id="ref-33"><label>33</label><mixed-citation publication-type="journal"><string-name><given-names>Richard</given-names> <surname>Shaw</surname></string-name>, <string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>, <string-name><given-names>Julia</given-names> <surname>Pescarini</surname></string-name>, <string-name><given-names>Elzo Pereira Pinto</given-names> <surname>Junior</surname></string-name>, <string-name><given-names>Mirjam</given-names> <surname>Allik</surname></string-name>, <string-name><given-names>Andressa</given-names> <surname>Siroky</surname></string-name>, <string-name><given-names>Desmond</given-names> <surname>Campbell</surname></string-name> <etal>et al</etal>. <article-title>Biases arising from linked administrative data for epidemiological research: a conceptual framework from registration to analyses</article-title>. <source><italic>European Journal of Epidemiology</italic></source>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2022</year>. <pub-id pub-id-type="doi">10.1007/s10654-022-00934-w</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>34</label><mixed-citation publication-type="journal"><string-name><given-names>Rinette</given-names> <surname>Badker</surname></string-name>, <string-name><given-names>Kierste</given-names> <surname>Miller</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Pardee</surname></string-name>, <string-name><given-names>Ben</given-names> <surname>Oppenheim</surname></string-name>, <string-name><given-names>Nicole</given-names> <surname>Stephenson</surname></string-name>, <string-name><given-names>Benjamin</given-names> <surname>Ash</surname></string-name>, <string-name><given-names>Tanya</given-names> <surname>Philippsen</surname></string-name>, <string-name><given-names>Christopher</given-names> <surname>Ngoon</surname></string-name>, <string-name><given-names>Partrick</given-names> <surname>Savage</surname></string-name>, <string-name><given-names>Cathine</given-names> <surname>Lam</surname></string-name>, <etal>et al</etal>. <article-title>Challenges in reported COVID-19 data: best practices and recommendations for future epidemics</article-title>. <source><italic>BMJ Global Health</italic></source>, <volume>6</volume> (<issue>5</issue>):<fpage>e005542</fpage>, <year>2021</year>. <pub-id pub-id-type="doi">10.1136/bmjgh-2021-005542</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>35</label><mixed-citation publication-type="journal"><string-name><given-names>Fabian</given-names> <surname>Panse</surname></string-name>, <string-name><given-names>Andr&#x00E9;</given-names> <surname>D&#x00FC;jon</surname></string-name>, <string-name><given-names>Wolfram</given-names> <surname>Wingerath</surname></string-name>, and <string-name><given-names>Benjamin</given-names> <surname>Wollmer</surname></string-name>. <article-title>Generating realistic test datasets for duplicate detection at scale using historical voter data</article-title>. In: <source><italic>International Conference on Extending Database Technology</italic></source>, <fpage>570</fpage>&#x2013;<lpage>581</lpage>, <year>2021</year>. <pub-id pub-id-type="doi">10.5441/002/edbt.2021.67</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>36</label><mixed-citation publication-type="journal"><string-name><given-names>Peter</given-names> <surname>Christen</surname></string-name>, <string-name><given-names>Ross W</given-names> <surname>Gayler</surname></string-name>, <string-name><given-names>Khoi-Nguyen</given-names> <surname>Tran</surname></string-name>, <string-name><given-names>Jeffrey</given-names> <surname>Fisher</surname></string-name>, and <string-name><given-names>Dinusha</given-names> <surname>Vatsalan</surname></string-name>. <article-title>Automatic discovery of abnormal values in large textual databases</article-title>. <source><italic>ACM Journal of Data and Information Quality</italic></source>, <volume>7</volume>(<issue>1-2</issue>):<fpage>1</fpage>&#x2013;<lpage>31</lpage>, <year>2016</year>. <pub-id pub-id-type="doi">10.1145/2889311</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>37</label><mixed-citation publication-type="journal"><string-name><given-names>Servick</given-names>, <surname>Kelly</surname></string-name>, and <string-name><given-names>Martin</given-names> <surname>Enserink</surname></string-name>. <article-title>The pandemic&#x2019;s first major research scandal erupts</article-title>. <source>Science</source>, <volume>368</volume>(<issue>6495</issue>), <fpage>1041</fpage>-<lpage>1042</lpage>, <year>2020</year>. <pub-id pub-id-type="doi">10.1126/science.368.6495.1041</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>38</label><mixed-citation publication-type="book"><string-name><given-names>Eric</given-names> <surname>Siegel</surname></string-name>. <source><italic>Predictive Analytics: The Power to Predict who Will Click, Buy, Lie, Or Die</italic></source>. <publisher-name>John Wiley and Sons</publisher-name>, <year>2013</year>. <uri>https://learning.oreilly.com/library/view/predictive-analytics-the/9781118416853/?ar=</uri>.</mixed-citation></ref>
<ref id="ref-39"><label>39</label><mixed-citation publication-type="journal"><string-name><given-names>Sarah</given-names> <surname>Larney</surname></string-name> and <string-name><given-names>Lucy</given-names> <surname>Burns</surname></string-name>. <article-title>Evaluating health outcomes of criminal justice populations using record linkage: the importance of aliases</article-title>. <source><italic>Evaluation Review</italic></source>, <volume>35</volume>(<issue>2</issue>):<fpage>118</fpage>&#x2013;<lpage>128</lpage>, <year>2011</year>. <pub-id pub-id-type="doi">10.1177/0193841X11401695</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>40</label><mixed-citation publication-type="book"><string-name><given-names>Finn</given-names> <surname>Brunson</surname></string-name> and <string-name><given-names>Helen</given-names> <surname>Nissenbaum</surname></string-name>. <source><italic>Obfuscation: A User&#x2019;s Guide for Privacy and Protest</italic></source>. <publisher-name>MIT Press</publisher-name>, <year>2015</year>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262029735.001.0001</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>41</label><mixed-citation publication-type="book"><string-name><given-names>Thomas N</given-names> <surname>Herzog</surname></string-name>, <string-name><given-names>Fritz J</given-names> <surname>Scheuren</surname></string-name>, and <string-name><given-names>William E</given-names> <surname>Winkler</surname></string-name>. <source><italic>Data Quality and Record Linkage Techniques</italic></source>. <publisher-name>Springer Verlag</publisher-name>, <year>2007</year>. <pub-id pub-id-type="doi">10.1007/0-387-69505-2</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>42</label><mixed-citation publication-type="website"><string-name><given-names>Patrick</given-names> <surname>McKenzie</surname></string-name>. <article-title>Falsehoods programmers believe about names</article-title>, <uri>https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-believe-about-names/</uri>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-43"><label>43</label><mixed-citation publication-type="journal"><string-name><given-names>Leigh</given-names> <surname>Mellish</surname></string-name>, <string-name><given-names>Emily A</given-names> <surname>Karanges</surname></string-name>, <string-name><given-names>Melisa J</given-names> <surname>Litchfield</surname></string-name>, <string-name><given-names>Andrea L</given-names> <surname>Schaffer</surname></string-name>, <string-name><given-names>Bianca</given-names> <surname>Blanch</surname></string-name>, <string-name><given-names>Benjamin J</given-names> <surname>Daniels</surname></string-name>, <string-name><given-names>Alicia</given-names> <surname>Segrave</surname></string-name>, and <string-name><given-names>Sallie-Anne</given-names> <surname>Pearon</surname></string-name>. <article-title>The Australian Pharmaceutical Benefits Scheme data collection: a practical guide for researchers</article-title>. <source><italic>BMC Research Notes</italic></source>, <volume>8</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2015</year>. <pub-id pub-id-type="doi">10.1186/s13104-015-1616-8</pub-id></mixed-citation></ref>
<ref id="ref-44"><label>44</label><mixed-citation publication-type="website"><string-name><given-names>Alex</given-names> <surname>Berry</surname></string-name>. <article-title>Germany&#x2019;s vaccination rate could be higher than previously thought</article-title>, <uri>https://p.dw.com/p/41oi7. Deutsche Welle</uri>, <day>7</day> <month>Oct</month>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-45"><label>45</label><mixed-citation publication-type="journal"><string-name><given-names>Samaneh</given-names> <surname>Khaef</surname></string-name>. <article-title>Registration of immigrants&#x2019; educational attainment in Sweden: an analysis of sources and time to registration</article-title>. <source><italic>Genus</italic></source>, <volume>78</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>20</lpage>, <year>2022</year>. <pub-id pub-id-type="doi">10.1186/s41118-022-00159-5</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>46</label><mixed-citation publication-type="book"><string-name><given-names>David J</given-names> <surname>Hand</surname></string-name>. <source><italic>Dark Data: Why What You Don&#x2019;t Know Matters</italic></source>. <publisher-name>Princeton University Press</publisher-name>, <year>2020</year>. <pub-id pub-id-type="doi">10.2307/j.ctvmd85db</pub-id></mixed-citation></ref>
<ref id="ref-47"><label>47</label><mixed-citation publication-type="book"><string-name><given-names>Roderick J</given-names> <surname>Little</surname></string-name> and <string-name><given-names>Donald B</given-names> <surname>Rubin</surname></string-name>. <source><italic>Statistical Analysis with Missing Data</italic></source>. <publisher-name>Wiley</publisher-name>, <publisher-loc>Hoboken</publisher-loc>, <edition>3</edition> edition, <year>2020</year>. <pub-id pub-id-type="doi">10.1002/9781119482260</pub-id></mixed-citation></ref>
<ref id="ref-48"><label>48</label><mixed-citation publication-type="journal"><string-name><given-names>Giampaolo P</given-names> <surname>Velo</surname></string-name> and <string-name><given-names>Pietro</given-names> <surname>Minuz</surname></string-name>. <article-title>Medication errors: Prescribing faults and prescription errors</article-title>. <source><italic>British Journal of Clinical Pharmacology</italic></source>, <volume>67</volume>(<issue>6</issue>): <fpage>624</fpage>&#x2013;<lpage>628</lpage>, <year>2009</year>. <pub-id pub-id-type="doi">10.1111/j.1365-2125.2009.03425.x</pub-id></mixed-citation></ref>
<ref id="ref-49"><label>49</label><mixed-citation publication-type="journal"><string-name><given-names>Latanya</given-names> <surname>Sweeney</surname></string-name>. <article-title><italic>K</italic>-anonymity: A model for protecting privacy</article-title>. <source><italic>International Journal of Uncertainty Fuzziness and Knowledge Based Systems</italic></source>, <volume>10</volume> (<issue>5</issue>):<fpage>557</fpage>&#x2013;<lpage>570</lpage>, <year>2002</year>. <pub-id pub-id-type="doi">10.1142/S0218488502001648</pub-id></mixed-citation></ref>
<ref id="ref-50"><label>50</label><mixed-citation publication-type="journal"><string-name><given-names>Thiemo</given-names> <surname>Fetzer</surname></string-name> and <string-name><given-names>Thomas</given-names> <surname>Graeber</surname></string-name>. <article-title>Measuring the scientific effectiveness of contact tracing: Evidence from a natural experiment</article-title>. <source><italic>Proceedings of the National Academy of Sciences</italic></source>, <volume>118</volume>(<issue>33</issue>), <year>2021</year>. <pub-id pub-id-type="doi">10.1073/pnas.2100814118</pub-id></mixed-citation></ref>
<ref id="ref-51"><label>51</label><mixed-citation publication-type="book"><string-name><given-names>Glenn</given-names> <surname>Firebaugh</surname></string-name>. <chapter-title>Statistics of ecological fallacy</chapter-title>. In: <string-name><given-names>Neil J</given-names> <surname>Smelser</surname></string-name> and <string-name><given-names>Paul B</given-names> <surname>Baltes</surname></string-name>, editors, <source><italic>International Encyclopedia of the Social and Behavioral Sciences</italic></source>, pages <fpage>4023</fpage>&#x2013;<lpage>4026</lpage>. <publisher-name>Pergamon</publisher-name>, <publisher-loc>Oxford</publisher-loc>, <year>2001</year>. <pub-id pub-id-type="doi">10.1016/B978-0-08-097086-8.44017-1</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>52</label><mixed-citation publication-type="journal"><string-name><given-names>Olivier</given-names> <surname>Binette</surname></string-name> and <string-name><given-names>Rebecca</given-names> <surname>Steorts</surname></string-name>. <article-title>(Almost) all of entity resolution</article-title>. <source><italic>Science Advances</italic></source>, <volume>8</volume>(<issue>12</issue>):<fpage>eabi8021</fpage>, <year>2022</year>. <pub-id pub-id-type="doi">10.1126/sciadv.abi8021</pub-id></mixed-citation></ref>
<ref id="ref-53"><label>53</label><mixed-citation publication-type="journal"><string-name><given-names>James</given-names> <surname>Doidge</surname></string-name> and <string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>. <article-title>Reflections on modern methods: linkage error bias</article-title>. <source><italic>International Journal of Epidemiology</italic></source>, <volume>48</volume>(<issue>6</issue>):<fpage>2050</fpage>&#x2013;<lpage>2060</lpage>, <year>2019</year>. <pub-id pub-id-type="doi">10.1093/ije/dyz203</pub-id></mixed-citation></ref>
<ref id="ref-54"><label>54</label><mixed-citation publication-type="journal"><string-name><given-names>Louise Mc</given-names> <surname>Grath-Lone</surname></string-name>, <string-name><given-names>Nicolas</given-names> <surname>Libuy</surname></string-name>, <string-name><given-names>David</given-names> <surname>Etoori</surname></string-name>, <string-name><given-names>Ruth</given-names> <surname>Blackburn</surname></string-name>, <string-name><given-names>Ruth</given-names> <surname>Gilbert</surname></string-name>, and <string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>. <article-title>Ethnic bias in data linkage</article-title>. <source><italic>The Lancet Digital Health</italic></source>, <volume>3</volume>(<issue>6</issue>):<fpage>e339</fpage>, <year>2021</year>. <pub-id pub-id-type="doi">10.1016/S2589-7500(21)00081-9</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>55</label><mixed-citation publication-type="journal"><string-name><given-names>Megan A</given-names> <surname>Bohensky</surname></string-name>, <string-name><given-names>Damien</given-names> <surname>Jolley</surname></string-name>, <string-name><given-names>Vijaya</given-names> <surname>Sundararajan</surname></string-name>, <string-name><given-names>Sue</given-names> <surname>Evans</surname></string-name>, <string-name><given-names>David V</given-names> <surname>Pilcher</surname></string-name>, <string-name><given-names>Ian</given-names> <surname>Scott</surname></string-name>, and <string-name><given-names>Caroline A</given-names> <surname>Brand</surname></string-name>. <article-title>Data linkage: a powerful research tool with potential problems</article-title>. <source><italic>BMC Health Services Research</italic></source>, <volume>10</volume>(<issue>346</issue>):<fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2010</year>. <pub-id pub-id-type="doi">10.1186/1472-6963-10-346</pub-id></mixed-citation></ref>
<ref id="ref-56"><label>56</label><mixed-citation publication-type="journal"><string-name><given-names>Sarah</given-names> <surname>Tahamont</surname></string-name>, <string-name><given-names>Zubin</given-names> <surname>Jelveh</surname></string-name>, <string-name><given-names>Aaron</given-names> <surname>Chalfin</surname></string-name>, <string-name><given-names>Shi</given-names> <surname>Yan</surname></string-name>, and <string-name><given-names>Benjamin</given-names> <surname>Hansen</surname></string-name>. <article-title>Dude, where&#x2019;s my treatment effect? Errors in administrative data linking and the destruction of statistical power in randomized experiments</article-title>. <source><italic>Journal of Quantitative Criminology</italic></source>, <volume>37</volume>(<issue>3</issue>):<fpage>715</fpage>&#x2013;<lpage>749</lpage>, <year>2021</year>. <pub-id pub-id-type="doi">10.1007/s10940-020-09461-x</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>57</label><mixed-citation publication-type="journal"><string-name><given-names>Jens</given-names> <surname>Bleiholder</surname></string-name> and <string-name><given-names>Felix</given-names> <surname>Naumann</surname></string-name>. <article-title>Data fusion</article-title>. <source><italic>ACM Computing Surveys</italic></source>, <volume>41</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>, <year>2008</year>. <pub-id pub-id-type="doi">10.1145/1456650.1456651</pub-id></mixed-citation></ref>
<ref id="ref-58"><label>58</label><mixed-citation publication-type="book"><string-name><given-names>Johnny</given-names> <surname>Blair</surname></string-name>, <string-name><given-names>Ronald F</given-names> <surname>Czaja</surname></string-name>, and <string-name><given-names>Edward A</given-names> <surname>Blair</surname></string-name>. <source><italic>Designing Surveys: A Guide to Decisions and Procedures</italic></source>. <publisher-name>Sage</publisher-name>, <publisher-loc>Thousand Oaks</publisher-loc>, <edition>3</edition> edition, <year>2014</year>. <uri>https://us.sagepub.com/en-us/nam/designing-surveys/book235701</uri></mixed-citation></ref>
<ref id="ref-59"><label>59</label><mixed-citation publication-type="journal"><string-name><given-names>Giles</given-names> <surname>Reid</surname></string-name>, <string-name><given-names>Felipa</given-names> <surname>Zabala</surname></string-name>, and <string-name><given-names>Anders</given-names> <surname>Holmberg</surname></string-name>. <article-title>Extending TSE to administrative data: A quality framework and case studies from Stats NZ</article-title>. <source><italic>Journal of Official Statistics</italic></source>, <volume>33</volume>(<issue>2</issue>), <year>2017</year>. <pub-id pub-id-type="doi">10.1515/jos-2017-0023</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>60</label><mixed-citation publication-type="journal"><string-name><given-names>Xiao-Li</given-names> <surname>Meng</surname></string-name>. <article-title>Statistical paradises and paradoxes in Big data (I): Law of large populations, Big data paradox, and the 2016 US presidential election</article-title>. <source><italic>The Annals of Applied Statistics</italic></source>, <volume>12</volume>(<issue>2</issue>):<fpage>685</fpage>&#x2013;<lpage>726</lpage>, <year>2018</year>. <pub-id pub-id-type="doi">10.1214/18-AOAS1161SF</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>61</label><mixed-citation publication-type="journal"><string-name><given-names>Eric I</given-names> <surname>Benchimol</surname></string-name>, <string-name><given-names>Liam</given-names> <surname>Smeeth</surname></string-name>, <string-name><given-names>Astrid</given-names> <surname>Guttmann</surname></string-name>, <string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>, <etal>et al</etal>. <article-title>The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) statement</article-title>. <source><italic>PLOS Med</italic></source>, <volume>12</volume>(<issue>10</issue>):<fpage>e1001885</fpage>, <year>2015</year>. <pub-id pub-id-type="doi">10.1371/journal.pmed.1001885</pub-id></mixed-citation></ref>
<ref id="ref-62"><label>62</label><mixed-citation publication-type="journal"><string-name><given-names>Ruth</given-names> <surname>Gilbert</surname></string-name>, <string-name><given-names>Rosemary</given-names> <surname>Lafferty</surname></string-name>, <string-name><given-names>Gareth</given-names> <surname>Hagger-Johnson</surname></string-name>, <string-name><given-names>Katie</given-names> <surname>Harron</surname></string-name>, <string-name><given-names>Li-Chun</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Peter</given-names> <surname>Smith</surname></string-name>, <string-name><given-names>Chris</given-names> <surname>Dibben</surname></string-name>, and <string-name><given-names>Harvey</given-names> <surname>Goldstein</surname></string-name>. <article-title>GUILD: Guidance for information about linking data sets</article-title>. <source><italic>Journal of Public Health</italic></source>, <volume>40</volume>(<issue>1</issue>):<fpage>191</fpage>&#x2013;<lpage>198</lpage>, <year>2017</year>. <pub-id pub-id-type="doi">10.1093/pubmed/fdx037</pub-id></mixed-citation></ref>
<ref id="ref-63"><label>63</label><mixed-citation publication-type="journal"><string-name><given-names>Mark D</given-names> <surname>Wilkinson</surname></string-name>, <string-name><given-names>Michel</given-names> <surname>Dumontier</surname></string-name>, <string-name><given-names>IJsbrand J</given-names> <surname>Aalbersberg</surname></string-name>, <string-name><given-names>Gabrielle</given-names> <surname>Appleton</surname></string-name>, <string-name><given-names>Myles</given-names> <surname>Axton</surname></string-name>, <string-name><given-names>Arie</given-names> <surname>Baak</surname></string-name>, <string-name><given-names>Niklas</given-names> <surname>Blomberg</surname></string-name>, <string-name><given-names>Jan-Willem</given-names> <surname>Boiten</surname></string-name>, <string-name><given-names>Luiz Bonino da Silva</given-names> <surname>Santos</surname></string-name>, <string-name><given-names>Philip E</given-names> <surname>Bourne</surname></string-name>, <etal>et al</etal>. <article-title>The FAIR guiding principles for scientific data management and stewardship</article-title>. <source><italic>Scientific Data</italic></source>, <volume>3</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2016</year>. <pub-id pub-id-type="doi">10.1038/sdata.2016.18</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>