<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v9i1.2389</article-id>
<article-id pub-id-type="publisher-id">9:1:18</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Generating synthetic identifiers to support development and evaluation of data linkage methods</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Lam</surname><given-names initials="J">Joseph</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="corresp" rid="correspondingAurthor">*</xref></contrib>
<contrib contrib-type="author"><name><surname>Boyd</surname><given-names initials="A">Andy</given-names></name><xref ref-type="aff" rid="affil-2">2</xref></contrib>
<contrib contrib-type="author"><name><surname>Linacre</surname><given-names initials="R">Robin</given-names></name><xref ref-type="aff" rid="affil-3">3</xref></contrib>
<contrib contrib-type="author"><name><surname>Blackburn</surname><given-names initials="R">Ruth</given-names></name><xref ref-type="aff" rid="affil-1">1</xref></contrib>
<contrib contrib-type="author"><name><surname>Harron</surname><given-names initials="K">Katie</given-names></name><xref ref-type="aff" rid="affil-1">1</xref></contrib>
<aff id="affil-1"><label>1</label><institution>Population, Policy &#x0026; Practice Research and Teaching Department, UCL Great Ormond Street Institute of Child Health, London, United Kingdom</institution></aff>
<aff id="affil-2"><label>2</label><institution>Population Health Sciences, Bristol Medical School, University of Bristol, Bristol, United Kingdom</institution></aff>
<aff id="affil-3"><label>3</label><institution>UK Ministry of Justice, London, United Kingdom</institution></aff>
</contrib-group>
<author-notes>
<corresp id="correspondingAurthor"><label>*</label>Corresponding author: Joseph Lam <email>joseph.lam.18@ucl.ac.uk</email></corresp>
</author-notes>
<pub-date date-type="pub" publication-format="electronic"><day>01</day><month>07</month><year>2024</year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year>2024</year></pub-date>
<volume>9</volume>
<issue>1</issue>
<elocation-id>2389</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/2389">This article is available from the IJPDS website at: https://ijpds.org/article/view/2389</self-uri>
<abstract>
<title>Abstract</title>
<sec>
<title>Introduction</title>
<p>Careful development and evaluation of data linkage methods is limited by researcher access to personal identifiers. One solution is to generate synthetic identifiers, which do not pose equivalent privacy concerns, but can form a &#x2018;gold-standard&#x2019; linkage algorithm training dataset. Such data could help inform choices about appropriate linkage strategies in different settings.</p>
</sec>
<sec>
<title>Objectives</title>
<p>We aimed to develop and demonstrate a framework for generating synthetic identifier datasets to support development and evaluation of data linkage methods. We evaluated whether replicating associations between attributes and identifiers improved the utility of the synthetic data for assessing linkage error.</p>
</sec>
<sec>
<title>Methods</title>
<p>We determined the steps required to generate synthetic identifiers that replicate the properties of real-world data collection. We then generated synthetic versions of a large UK cohort study (the Avon Longitudinal Study of Parents and Children; ALSPAC), according to the quality and completeness of identifiers recorded over several waves of the cohort. We evaluated the utility of the synthetic identifier data in terms of assessing linkage quality (false matches and missed matches).</p>
</sec>
<sec>
<title>Results</title>
<p>Comparing data from two collection points in ALSPAC, we found within-person disagreement in identifiers (differences in recording due to both natural change and non-valid entries) in 18% of surnames and 12% of forenames. Rates of disagreement varied by maternal age and ethnic group. Synthetic data provided accurate estimates of linkage quality metrics compared with the original data (within 0.13-0.55% for missed matches and 0.00-0.04% for false matches). Incorporating associations between identifier errors and maternal age/ethnicity improved synthetic data utility.</p>
</sec>
<sec>
<title>Conclusions</title>
<p>We show that replicating dependencies between attribute values (e.g. ethnicity), values of identifiers (e.g. name), identifier disagreements (e.g. missing values, errors or changes over time), and their patterns and distribution structure enables generation of realistic synthetic data that can be used for robust evaluation of linkage methods.</p>
</sec>
</abstract>
<kwd-group>
<kwd>Keywords record linkage</kwd>
<kwd>data linkage</kwd>
<kwd>synthetic data</kwd>
<kwd>synthetic identifiers</kwd>
<kwd>linkage evaluation</kwd>
<kwd>ALSPAC</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>Introduction</title>
<p>Data linkage facilitates the combination of detailed information on individuals captured in disparate data sources, without the need for new data collection. Linkage is increasingly used as an efficient approach, particularly with existing administrative datasets, and has great potential for social good. Access to identifiable information is crucial when linking multiple datasets, as linkage depends on either the availability of unique identifiers (e.g. a social security number) or a set of individually non-unique variables such as name, sex and date of birth, which in combination, can identify an individual. This is the case for both linkage using identifiers in their natural form or for privacy preserving techniques which mask the identifiers in some form. The level of completeness, uniqueness and accuracy of identifiers recorded in administrative data pose a challenge for linkage, particularly when linking across multiple sectors in countries where unique citizen identifiers are unavailable [<xref ref-type="bibr" rid="ref-1">1</xref>]. Careful development and evaluation of linkage methods is therefore required in order to achieve high quality linkage and robust results [<xref ref-type="bibr" rid="ref-2">2</xref>, <xref ref-type="bibr" rid="ref-3">3</xref>]. However, methodological development has been constrained by confidentiality concerns and legislative restrictions governing access to personal information for research.</p>
<p>In practice, access to identifiers is usually limited to either the data owners or trusted third parties, who may be unwilling or unable to make use of these identifiers for methodological purposes. Conversely, analysts will typically only have access to the de-identified linked data, with limited information about any uncertainty in linkage, or information with which to assess the quality of linkage [<xref ref-type="bibr" rid="ref-4">4</xref>]. This separation limits opportunities for the development of advanced linkage methods and the assessment of computational performance and/or linkage quality, as the researchers cannot access the identifiable data needed for these evaluations [<xref ref-type="bibr" rid="ref-5">5</xref>]. Even when it is possible to access identifiers, lack of a &#x201C;ground truth&#x201D; makes it difficult to evaluate different linkage strategies, as there is no gold standard against which results can be compared.</p>
<p>One solution to this problem is to generate synthetic datasets of identifiers that mimic the characteristics of real identifiers (and so can be used for methodological work) but that do not pose any confidentiality issues [<xref ref-type="bibr" rid="ref-6">6</xref>, <xref ref-type="bibr" rid="ref-7">7</xref>]. Such synthetic datasets of identifiers would include a &#x201C;ground truth&#x201D; to enable evaluation of different linkage methods. Synthetic data generators have been developed in the context of providing realistic <italic>research</italic> datasets, where the aim is to mimic the underlying statistical properties of the original data whilst minimising disclosure risk [<xref ref-type="bibr" rid="ref-8">8</xref>, <xref ref-type="bibr" rid="ref-9">9</xref>]. A summary of existing synthetic data generation methods is described by Kokosi [<xref ref-type="bibr" rid="ref-9">9</xref>]. Synthetic data that retain the relationships between variables in the original data can provide an accurate representation of the original data and be used for a range of purposes, including evaluation of different methodological approaches [<xref ref-type="bibr" rid="ref-10">10</xref>]. However, these approaches have mainly been developed in the context of &#x2018;attribute&#x2019; data, i.e. variables typically used within an analysis (e.g. social or health status, occupation). In the context of developing linkage methods, we are concerned with &#x2018;identifier&#x2019; data, i.e. variables used for linkage but not necessarily for analysis (e.g. postcode, name). In some cases, there is overlap between the two: date/year of birth and sex can be both attribute variables and personal identifiers.</p>
<p>In most applications of synthetic data, retaining the relationships between different variables helps to replicate the underlying structure of the data and enables users to test and evaluate different methodological approaches. When generating synthetic identifier data, there is also a need to ensure that the data retain dependencies between variables (for example, name might be associated with date of birth). However, there are a number of reasons why an alternative approach to generating synthetic data is required to address the idiosyncrasy of identifiers. Firstly, identifier variables do not always follow standard statistical distributions. Secondly, identifiers are affected by specific types of recording errors and changes that occur within and between datasets, and over time. We refer to these disagreements as &#x2018;errors&#x2019;, whilst recognising that in some cases these will be genuine changes (e.g. address change due to migration or surname change following marriage) rather than errors in recording. Such errors are often related to attribute variables (e.g. names may be more often misspelt for particular ethnic groups; address changes are associated with age and changes in socio-economic and potentially health status). There may also be interdependencies between identifier errors, e.g. if name and address change at the same time due to divorce. Accurate replication of identifier errors and their dependencies on attribute variables is important, since these dependencies are directly related to the impact that linkage errors have on analysis [<xref ref-type="bibr" rid="ref-11">11</xref>]. Therefore, these errors and dependencies should be replicated within any synthetic identifier datasets that are used to test linkage methods, so that an assessment of bias resulting from linkage can be conducted [<xref ref-type="bibr" rid="ref-12">12</xref>]. Existing datasets generated to facilitate the development and testing of data linkage algorithms have typically not focussed on preserving these dependencies, and have not been evaluated in terms of their utility for testing linkage algorithms. There is therefore a scientific requirement to develop more robust and realistic synthetic identifier datasets [<xref ref-type="bibr" rid="ref-13">13</xref>].</p>
<p>This paper presents a framework for generating synthetic identifier data that could be used by data owners to enable researchers to develop and test linkage methods in different settings. Use of these data could help overcome the limited capacity for linkage methodology development by providing wider access to realistic identifier data, without disclosure risk, and with a ground truth against which linkage quality can be assessed. In Section 1, we describe a motivating scenario and outline how the steps needed to generate synthetic data can be implemented. Importantly, we consider the need to preserve the dependencies between identifier values, identifier errors, and attribute values. In Section 2, we evaluate the use of synthetic data for assessing linkage quality, based on an exemplar of longitudinal linkage within a large UK cohort study (the Avon Longitudinal Study of Parents and Children; ALSPAC)</p>
</sec>
<sec>
<title>Section 1: A framework for generating synthetic identifier data</title>
<sec>
<title>Motivating scenario</title>
<p>Our motivating scenario is one in which we aim to conduct performance comparisons between different linkage approaches, in a secure manner with low ethico-legal barriers and no intrusion into personal privacy. The aim of such performance comparisons is to optimise linkage algorithms that would then be applied to specific, real-world linkage projects. We assume that those commissioning the linkage (e.g. a researcher) will not have access to identifiers and that the linkage will be conducted by a trusted third party or data owner. We will examine the utility of synthetic datasets to conduct linkage performance comparisons that would be sufficiently similar to the real data whilst not intrusive of personal privacy. Such datasets could be useful for the development of linkage methods in two settings: 1) by data owners who have access to identifiers but where there are restrictions around using these identifiers for methodological development rather than business-as-usual linkage; 2) by researchers or research infrastructure providers (such as Trusted Research Environments) who cannot access identifiers but for whom synthetic data would be useful for understanding the implications of different linkage methods on their outputs.</p>
</sec>
<sec>
<title>Types of variables</title>
<p>First, we distinguish between two types of variables: identifiers and attributes.</p>
<list list-type="roman-lower">
<list-item><p><italic>Identifier variables</italic> (e.g. name, NHS number, postcode, sex, date of birth) that are used within linkage but not necessarily the analysis (though some, e.g. sex, are also attribute variables). Some of these variables may be related to the values of other identifiers and/or attribute variables (e.g. values of name might be associated with sex and ethnicity). Presence of errors in one identifier might be related to errors in other identifiers (e.g. if name is mistyped, it might be more likely that date of birth is also recorded with error).</p></list-item>
<list-item><p><italic>Attribute variables</italic> (e.g. ethnicity) that are used within analysis but not necessarily the linkage. Some of these variables may be associated with patterns in identifier values and/or identifier errors (e.g. ethnicity might be associated with <italic>values</italic> and also <italic>errors</italic> in name).</p></list-item>
</list>
<p>We make this distinction because under our motivating scenario, we are mostly interested in generating identifier variables. However, to ensure that the synthetic data are realistic, we need to consider i) the dependencies between identifier values and attributes, and ii) how errors in identifiers are distributed in relation to attribute variables. We often find that errors in linkage (and by implication, in identifiers) are related to differences in the underlying data quality for particular subgroups or to particular events and circumstances. For example, family name may have a higher probability of being typed incorrectly for individuals from minority ethnic groups compared to a majority ethnic group given that the family name may be unfamiliar to the operative recording the data, or that there are cultural differences in the length or structural complexity of names; linkage may be less likely to be successful for individuals following family separation given the tendency for this to result in changes in both address and names. This can therefore lead to dependencies between errors in identifiers &#x2013; i.e. a change in name may be more likely for individuals who have also changed address. Evidence from the literature suggests that age, ethnic group, sex, deprivation and measures of health and social status may all be related to the risk of linkage error [<xref ref-type="bibr" rid="ref-14">14</xref>, <xref ref-type="bibr" rid="ref-15">15</xref>]. In order to generate a dataset that is realistic for testing linkage methods, it is therefore crucial to consider whether identifier errors are likely to be related to attribute values. An example of the possible dependencies between identifier values, attribute values, and identifier errors is presented in <xref ref-type="fig" rid="fig-1">Figure 1</xref>.</p>
<fig id="fig-1"><label>Figure 1: Dependencies between identifier values, attribute values, and identifier errors</label>
<graphic xlink:href="ijpds-09-2389-g001.tif"/>
<attrib>The examples given are not exhaustive but suggestive of the dependencies that might exist between values and errors in different variables.</attrib>
</fig>
</sec>
<sec>
<title>Steps in the process of generating synthetic identifier data</title>
<p>This section outlines five steps that are required to generate a realistic set of synthetic identifiers. In summary, the objective of these steps is to generate a &#x2018;gold-standard&#x2019; dataset, i.e. the correct identifiers recorded in the absence of errors or changes over time (Steps 1&#x2013;3). We then need to customise types and patterns of errors to be introduced to the gold-standard dataset (Step 4). Finally, we create multiple versions of the corrupted data (Step 5). A workflow for this process is presented in <xref ref-type="fig" rid="fig-2">Figure 2</xref>.</p>
<fig id="fig-2"><label>Figure 2: Workflow for generating synthetic identifier datasets</label>
<graphic xlink:href="ijpds-09-2389-g002.tif"/>
</fig>
<sec>
<title>Step 1: Elicit Information</title>
<p>Data linkers or data owners should elicit information on the set of identifiers in each file that are available for linkage, the rates of missingness (percentage of records with missing values for each identifier) and the characteristics of these identifiers (e.g. the range of dates of birth, the percentage of records that have a unique name).</p>
<p>They also need to elicit information about likely rates of errors, types of errors and their patterns of co-occurrence in identifiers, and how these errors are associated with attribute variables. In practice, information on errors may be difficult to obtain, and may need to be based on knowledge about identifier errors (or linkage errors) from other similar data sources or the literature [<xref ref-type="bibr" rid="ref-16">16</xref>]. For the purposes of this paper, we use information on the rates, types and distribution of identifier errors based on analysis of data collected over different waves of the Avon Longitudinal Study of Parents and Children (ALSPAC) birth cohort study (see Section 2 for details on ALSPAC) [<xref ref-type="bibr" rid="ref-17">17</xref>, <xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Information should be obtained on the number of records in each dataset and the joint distribution of key attribute variables (e.g. age and ethnicity), and whether individuals are likely to be recorded multiple times within a dataset (e.g., as they would in hospital admission records).</p>
</sec>
<sec>
<title>Step 2: Generate attribute variables</title>
<p>With access to the gold-standard identifiers and attribute data, we can use then use the <italic>Synthpop</italic> package in R to synthesise attribute variables (such as age or ethnicity) [<xref ref-type="bibr" rid="ref-19">19</xref>]. <italic>Synthpop</italic> uses a series of conditional models based on the original data to sequentially predict and impute values of each variable in the synthetic data. This process preserves variable inter-dependency by utilising classification and regression tree models. The output of this step is a &#x2018;gold-standard&#x2019; synthetic dataset of attribute variables replicating those found in the original data.</p>
</sec>
<sec>
<title>Step 3: Generate identifiers</title>
<p>Two types of identifiers can be generated: those that are dependent on attribute variables, and those that are independent. Independent identifiers, e.g. NHS number (or social security number, etc.) can be generated according to predefined rules. Identifiers that are dependent on attribute variables will be generated according to the attribute values generated in Step 2. For example, date of birth can be generated according to the distribution of age. <xref ref-type="table" rid="table-1">Table 1</xref> describes how different types of identifiers might be generated, according to whether or not they are dependent on attribute variables.</p>
<table-wrap id="table-1">
<label>Table 1</label><caption><title>Generating identifier values</title></caption>
<table frame="hsides" rules="groups">
<col width="20%"/>
<col width="20%"/>
<col width="30%"/>
<col width="30%"/>
<tbody>
<tr>
<td valign="middle" align="left" style="border-top: solid 1pt; border-bottom: solid 1pt"><bold>Identifier</bold></td>
<td valign="middle" align="left" style="border-top: solid 1pt; border-bottom: solid 1pt"></td>
<td valign="middle" align="left" style="border-top: solid 1pt; border-bottom: solid 1pt"><bold>Generation process</bold></td>
<td valign="middle" align="left" style="border-top: solid 1pt; border-bottom: solid 1pt"><bold>Example</bold></td>
</tr>
<tr>
<td align="left" valign="top">Identifiers that are independent of attribute variables</td>
<td align="left" valign="top">Date of birth and sex</td>
<td align="left" valign="top"><p>Given aggregate information on the distribution of these identifiers, or elements of these identifiers (i.e. year of birth) within the original data, values can be sampled directly from the relevant distributions.</p>
<p>In some cases, we might want to reflect dependency between records, for example, there will be a minimum distance between the date of birth of a baby and their mother.</p>
<p>In other cases, we might want to allow date of birth to depend on attribute variables, such as place of residence.</p></td>
<td align="left" valign="top">Date of birth can be generated from the distribution of year of birth, by assuming a uniform distribution over all possible (or eligible) dates within each year. We could also allow for variations according to day of the week or month of the year (e.g. those born on 31st of December of any year may have a higher probability of being recorded as being born on 1st January of the next year, rather than another random date).</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">Unique identifiers</td>
<td align="left" valign="top">Values of unique identifiers such as a social security number or NHS number can be randomly generated following defined rules. The assumption that unique identifiers are independent does not hold where another identifier is included within in the unique identifiers (e.g. the Community Health Index number in Scotland, which is derived from date of birth and sex).</td>
<td align="left" valign="top">NHS number is assigned at birth in England and is unrelated to any other personal information [<xref ref-type="bibr" rid="ref-22">22</xref>]. It comprises ten digits, of which the majority are random numbers and the tenth is a check digit to confirm validity: it can therefore be generated using a simple algorithm. If there are multiple unique identifiers (e.g. NHS number and hospital number), these can be generated independently.</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">Other identifiers</td>
<td align="left" valign="top">Personal identifiers such as email addresses, telephone numbers, and social media handles can be generated according to rules. In some cases, we might also want to allow these identifiers to depend on attribute variables: e.g. generating random telephone numbers based on the country and area of residence, or generating random email addresses based on names, date and country of birth.</td>
<td align="left" valign="top">Fake Mail Generator (<uri>https://fakedetail.com/fake-mail-generator</uri>) allows the generation of random email addresses given real domains. We can also allow these identifiers to depend on other identifiers or attributes: for example, Fake Number (<uri>https://fakenumber.org/united-kingdom</uri>) can generate random telephone numbers based on the country and area of residence.</td>
</tr>
<tr>
<td align="left" valign="top">Identifiers that are dependent on attribute variables</td>
<td align="left" valign="top">Names</td>
<td align="left" valign="top">First names may be related to age, ethnicity, sex and geography; surnames may also be related to ethnicity. Frequency look-up tables provide a useful tool for sampling names and mapping them to predictor attribute variables. Names can be directly sampled from such frequency tables, and can be allowed to depend on attribute variables such as sex and ethnicity, where these are available.</td>
<td align="left" valign="top">The Office for National Statistics (ONS) publishes the rank and count of the baby birth names in England and Wales every year, which can be used as the forename frequency table for the England and Wales population [<xref ref-type="bibr" rid="ref-23">23</xref>]. National Records of Scotland also publish popular baby forenames depending on year of birth and gender [<xref ref-type="bibr" rid="ref-24">24</xref>]. Another example is data on forename, gender and ethnicity extracted from the US census and implemented in the R package &#x2018;randomNames&#x2019; [<xref ref-type="bibr" rid="ref-25">25</xref>]. Similar frequency tables, including for surnames, are published in many countries [<xref ref-type="bibr" rid="ref-26">26</xref>].</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">Addresses</td>
<td align="left" valign="top">Addresses may be related to personal social status, income and ethnic background [<xref ref-type="bibr" rid="ref-27">27</xref>]. For example, in 2018, 41% of residents in the London borough of Tower Hamlets were of Asian ethnic background, compared with 5% in the borough of Bromley.</td>
<td align="left" valign="top">To represent these dependencies in synthetic data, we can start by sampling postcodes from a relevant list. Levels of deprivation can then be assigned to each postcode using the English indices of deprivation (Index of Multiple Deprivation; IMD), and ethnic group distributions can be assigned using ethnic group statistics by geography [<xref ref-type="bibr" rid="ref-28">28</xref>, <xref ref-type="bibr" rid="ref-29">29</xref>]. Given information on the distribution of ethnic group in the original data, addresses can then be sampled from a frequency table.</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">Indirect identifiers</td>
<td align="left" valign="top">Other non-traditional identifiers used for linkage might include &#x2018;indirect&#x2019; identifiers such as clinical variables or dates [<xref ref-type="bibr" rid="ref-30">30</xref>]. Given sufficient aggregate data on the distributions of these variables, and assumptions about their dependence on attribute variables, these could be generated in a similar way to date of birth and sex (i.e. according to specified distributions).</td>
<td align="left" valign="top"></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Identifiers that are dependent on attribute variables and have high cardinality, such as names, are more challenging to synthesise. There is no existing library that readily generates names and maintains dependencies with other variables. Generation of names should consider the following factors:</p>
<list list-type="order">
<list-item><p>Privacy and Disclosure Risk: ensuring none of the unique forename-surname combination in the original data appear in the synthesised data</p></list-item>
<list-item><p>Uniqueness</p></list-item>
<list-item><p>Frequency: common names should be synthesised for common names in the original data. For example: &#x201C;John&#x201D; (White, Male, common) in ALSPAC surnames or forename could be replaced with &#x201C;Peter&#x201D; (White, Male, common) from the name dictionary/look up table.</p></list-item>
<list-item><p>Sharing of surnames between siblings, parents and children and partners.</p></list-item>
</list>
<p>It is helpful to consider the frequency or uniqueness of different identifier values, as well as their distribution with respect to attribute variables. For example, a male has a higher probability than a female of having a forename of &#x2018;Patrick&#x2019;, and the distribution or uniqueness of names may vary according to ethnic group, levels of deprivation and by age (reflecting changing fashions for names).</p>
</sec>
<sec>
<title>Step 4: Data corruption</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> provides a summary of different types of errors that may be found in real data and can be introduced during the synthetic data generation.</p>
<table-wrap id="table-2">
<label>Table 2</label><caption><title>Types of identifier errors that can be introduced to synthetic data</title></caption>
<table frame="hsides" rules="groups">
<col width="20%"/>
<col width="40%"/>
<col width="40%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Type of error/identifier</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Description</bold></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Manifestations</bold></td>
</tr>
<tr>
<td align="left" valign="top">Typographic error (string variables)</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Occurs during manual typing, e.g. a receptionist types a patient&#x2019;s information for a general practitioner appointment booking.</p></list-item>
<list-item><p>Depending on the keyboard layout, characters may be <italic>substituted</italic> with neighbouring keyboard characters e.g. &#x2018;s&#x2019; instead of &#x2018;d&#x2019;.</p></list-item>
<list-item><p>New characters or space may be accidentally <italic>inserted</italic> into a field, random characters may be <italic>omitted</italic> from a field, or character positions may be <italic>transposed.</italic></p></list-item>
<list-item><p>Errors may result from hitting a key twice, letting eyes move faster than the hand, or misreading [<xref ref-type="bibr" rid="ref-31">31</xref>].</p></list-item>
</list></td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Typographical errors are more likely to happen in the middle or towards the end of the word, and in longer words [<xref ref-type="bibr" rid="ref-32">32</xref>, <xref ref-type="bibr" rid="ref-33">33</xref>].</p></list-item>
<list-item><p>Over 80% of typographical errors are single instances of substitution, insertion, deletion or transposition [<xref ref-type="bibr" rid="ref-31">31</xref>].</p></list-item>
<list-item><p>The likelihood of substituting neighbouring characters differs according to layout as well as personal typing habit, e.g. it is more likely that &#x2018;d&#x2019; is replaced with &#x2018;s&#x2019; than with &#x2018;x&#x2019; [<xref ref-type="bibr" rid="ref-34">34</xref>].</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">Phonetic error (string variables &#x2013; particularly name)</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Occurs during dictation, where letters may be substituted with letters that are phonetically the same but orthographically incorrect for the intended word, e.g. when a receptionist records information given by a patient, (s)he may mishear information due to the accent of the patient or the pronunciation of similar words or characters, such as &#x2018;F&#x2019; instead of &#x2018;Ph&#x2019; [<xref ref-type="bibr" rid="ref-34">34</xref>].</p></list-item>
</list></td>
<td align="left" valign="top">
<list list-type="bullet"><list-item><p>Information on phonetic errors can be derived from phonetic algorithms, which apply a range of rules and exceptions to encode words by their pronunciation, instead of spellings. These algorithms have been widely used in applications such as spell checkers and search engines and algorithms have been transformed into look-up tables and rules to group similar-sounding words together [<xref ref-type="bibr" rid="ref-35">35</xref>].</p></list-item>
<list-item><p>Soundex is one of the most widely known phonetic algorithms for Anglo-Saxon surname encoding [<xref ref-type="bibr" rid="ref-36">36</xref>]. Extensions to Soundex overcome limitations in recognising different languages and dialects that may have different pronunciations for the same names [<xref ref-type="bibr" rid="ref-37">37</xref>].</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">Optical Character Recognition error (OCR, any identifiers)</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>The OCR system is used to process scanned handwritten documents into electronic versions.</p></list-item>
<list-item><p>OCR errors occur when the system fails to distinguish two characters that have similar shapes, such as &#x2018;l&#x2019; and &#x2018;1&#x2019; or &#x2018;m&#x2019; and &#x2018;rn&#x2019;.</p></list-item>
</list></td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Error rates in OCR systems can be high if the scanned documents are poorly handwritten, in bad physical condition or have a complex layout [<xref ref-type="bibr" rid="ref-38">38</xref>].</p></list-item>
<list-item><p>Error rates in OCR systems are impacted by configuration settings (such as where the threshold is set for manual review).</p></list-item>
<list-item><p>Look up tables are available that provide around 80 pairs of OCR errors where letters, digits, symbols and combinations of these appear to be similar.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">Naming convention inconsistencies</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Some people have two first names (with or without a hyphen), or middle names that are used as first names or vice versa.</p></list-item>
<list-item><p>Double-barrel surnames may be recorded differently in different datasets (e.g. with or without hyphens) and may include abbreviations (e.g. Saint John as St. John).</p></list-item>
<list-item><p>First names and surnames may be swapped.</p></list-item>
<list-item><p>Migrant groups might &#x2018;adopt&#x2019; localised versions of names</p></list-item>
<list-item><p>Nicknames and diminutives might be provided</p></list-item>
</list></td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Look up tables of common name variants are available. The software &#x2018;Febrl&#x2019; provides around 350 rules and name variants (e.g. &#x2018;Edward&#x2019; for &#x2018;Ted&#x2019;, &#x2018;Edwin&#x2019; and Edwards&#x2019;) [<xref ref-type="bibr" rid="ref-13">13</xref>]. Database of common English diminutives of formal given names are available on Wiktionary.</p></list-item>
<list-item><p>Table of common surnames with different Romanised representation of the same character are available on Wiktionary (Mandarin Chinese, Cantonese, Hakkan, Korean, Vietnamese, Japanese)</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">Date errors (date of birth or other date identifiers)</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Format differences, i.e. between countries or people. In the UK, people usually record their date of births in Day-Month-Year format, while in the US it is more often written in the format of Month-Day-Year and in China is Year-Month-Day.</p></list-item>
<list-item><p>Default/generic values. Some systems have a default value for the date of birth, resulting in those people with missing date of birth automatically being given a default date.</p></list-item>
<list-item><p>Accidental input of &#x2018;today&#x2019;s date&#x2019;</p></list-item>
</list></td>
<td align="left" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top">Changes over time (e.g. name, sex, postcode)</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Postcode changes occur as people move and if addresses are not updated on a system (e.g. postcodes in healthcare data might only be updated when a patient registers with a new general practitioner, which might be some time after an address change).</p></list-item>
<list-item><p>Children may have multiple genuine postcodes if they have more than one residence, e.g. mother&#x2019;s or father&#x2019;s address.</p></list-item>
<list-item><p>Surnames may change following marriage or divorce; recorded sex may change over time.</p></list-item>
<list-item><p>Postcodes change over time for the same property to reflect changes in the postal system.</p></list-item>
</list></td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>In the UK, evidence suggests that 40% of children move home in the first 5 years of life; 5% move 3 or more times within this time period [<xref ref-type="bibr" rid="ref-39">39</xref>].</p></list-item>
</list></td>
</tr>
<tr>
<td align="left" valign="top">Unique identifier errors</td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Checksums or other validation methods may be used to prevent invalid identifiers from being recorded.</p></list-item>
<list-item><p>Intentional use of another person&#x2019;s identifier may lead to errors.</p></list-item>
<list-item><p>Changes to unique identifiers may occur over time, and some identifiers might be reused, resulting in multiple individuals with the same identifier [<xref ref-type="bibr" rid="ref-40">40</xref>].</p></list-item>
<list-item><p>Individuals may be issued many unique IDs (e.g. a pupil moving from one school to another)</p></list-item>
</list></td>
<td align="left" valign="top">
<list list-type="bullet">
<list-item><p>Accurate recording of unique identifiers that depend on interactions with services may be related to how different individuals access those services. For example, completeness of NHS number is often lower for young males [<xref ref-type="bibr" rid="ref-41">41</xref>].</p></list-item>
</list></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Data corruption is split into the following steps:</p>
<p>Firstly, error rates, types and co-occurrence patterns are defined and pre-specified.</p>
<p>Secondly, for each row of synthetic data, a corrupted version is generated. There are several approaches available for this data corruption. One approach is to generate multiple rows of corrupted data capturing all combinations of expected errors and patterns. This method retains all pre-specified error type combinations but could be computationally expensive for large datasets. Alternatively, the Splink synthetic data corruptor adapts a likelihood approach to introducing errors, generating multiple rows of corrupted data probabilistically [<xref ref-type="bibr" rid="ref-20">20</xref>]. In Splink&#x2019;s synthetic data corruptor, a baseline probability is assigned for each type of error, and a multiplier is applied based on attribute variables. For example, by following the Zipf distribution, up to 20 rows with varying error types and combinations can be generated for each row of data [<xref ref-type="bibr" rid="ref-21">21</xref>]. This method is less computationally expensive and has the capability to introduce some error-attribute variable dependency. However, this method does not necessarily capture all pre-specified error type combinations and co-occurrence patterns.</p>
<p>The final stage is to draw samples from the corrupted data that satisfy the pre-specified error types, co-occurrence patterns, and error-attribute characteristics.</p>
</sec>
<sec>
<title>Step 5: Generate linkage files</title>
<p>Since the errors selection in Step 4 is probabilistic, we can generate multiple sets of corrupted data files by repeating the step. This gives us several (e.g. 5) different corrupted versions of the same gold standard file, which represent multiple versions of a &#x2018;linkage&#x2019; file. Generating multiple versions of the linkage file is appropriate as it reflects the uncertainty in the process of replicating the original data, in line with the logic of using multiple imputation to model uncertainty.</p>
</sec>
</sec>
</sec>
<sec>
<title>Section 2: Evaluating synthetic data</title>
<sec>
<title>Motivating scenario</title>
<p>The following section describes an evaluation of the utility of the data we have generated under our framework. We use an exemplar of data linkage within the ALSPAC birth cohort. In ALSPAC, identifiers for each participant were recorded at multiple time points or data collection waves. For the purposes of evaluating the synthetic data, we used data from a gold-standard list of identifiers held within the ALSPAC administrative database (called ARCADIA, see <xref ref-type="supplementary-material" rid="sup-a">Appendix Table 1</xref>) which contains the &#x2018;live&#x2019; best understanding of participants current details, and raw records from one data collection wave (the Child Health Database; CHDB), collected when participants were aged 6 years. A unique ALSPAC ID identifies the same individual within ARCADIA and CHDB, but the identifiers collected in each dataset differ. This gives us a gold-standard database that can be used to assess how well synthetic data performs at evaluating different linkage approaches.</p>
<p>We first generate synthetic versions of the identifier data held within ALSPAC, creating a number of &#x2018;linkage files&#x2019; to represent ARCADIA and CHDB. Next, we link the synthetic versions of ARCADIA with synthetic versions of CHDB, and derive metrics of linkage quality. Finally, we compare the linkage quality metrics derived from the synthetic data to the metrics derived from the gold-standard ALSPAC data.</p>
</sec>
<sec>
<title>Source data</title>
<p>The Avon Longitudinal Study of Parents and Children (ALSPAC) is a prospective population-based study [<xref ref-type="bibr" rid="ref-17">17</xref>, <xref ref-type="bibr" rid="ref-18">18</xref>]. Initial recruitment of pregnant women took place in 1990-1992 and the health and development of the children from these pregnancies and their family members have been followed ever since. For this study, we focus on the original parents/carers (Generation 0, G0) and the index children (Generation 1, G1). ALSPAC recruited 14,541 pregnancies by women (G0) who were resident in and around the City of Bristol (South West UK) with expected dates of delivery 1st April 1991 to 31st December 1992. Of these initial pregnancies, there were a total of 14,676 foetuses, resulting in 14,062 live births and 13,988 children who were alive at 1 year of age. The eligible sampling frame was constructed retrospectively using linked recruitment and health service records. Additional offspring that were eligible to enrol in the study have been welcomed through major recruitment drives at the ages of 7 and 18 years; and through opportunistic contacts since the age of 7. A total of 913 additional G1 participants have been enrolled in the study since the age of 7 years with 195 of these joining since the age of 18. This additional enrolment provides a baseline sample of 14,901 G1 participants who were alive at 1 year of age.</p>
</sec>
<sec>
<title>Linkage methods</title>
<p>Our aim was to determine whether we could use synthetic data to evaluate the quality of different linkage algorithms. Therefore, we used three different linkage strategies to link data for 13,281 individuals in ARCADIA who also had a record in CHDB. We conducted the linkage based on child&#x2019;s forename, surname, date of birth and gender, plus mother&#x2019;s surname, using the following methods, with further details in <xref ref-type="supplementary-material" rid="sup-a">Appendix 4</xref>. Linkage strategies were compared and probabilistic linkage thresholds were chosen to align with the deterministic linkage model, to enable a fair comparison. We estimated false match and missed match rates for each method.</p>
<list list-type="order">
<list-item><p>Deterministic linkage. We classified records as belonging to the same individual if at least 4 of the 5 identifiers matched exactly.</p></list-item>
<list-item><p>Probabilistic linkage with similarity scores. We calculated probabilistic match weights for agreement/disagreement using the Fellegi-Sunter approach [<xref ref-type="bibr" rid="ref-42">42</xref>]. To allow for typographical errors in names, we calculated probabilistic match weights using the Jaro-Winkler similarity score [<xref ref-type="bibr" rid="ref-33">33</xref>]. Similarity scores were categorised as little agreement (a score of 0&#x2013;0.8), moderate agreement (0.8&#x2013;&lt;1), or full agreement (a score of 1). Linkages were accepted at or above the weight threshold of 3.</p></list-item>
<list-item><p>Probabilistic linkage with similarity scores and term frequency adjustments for forenames and surnames. On top of method 2, we accounted for name frequencies by proportionally adjusting u-probabilities for agreement or disagreement on less common names. Linkages were accepted at or above the weight threshold of 2.</p></list-item>
</list>
</sec>
<sec>
<title>Generating synthetic ALSPAC data</title>
<p>In order to generate realistic synthetic data, we first needed to understand the levels of errors observed in identifiers within ALSPAC. Since we had access to the gold-standard ALSPAC data, we could directly estimate the error rates for each identifier (see <xref ref-type="supplementary-material" rid="sup-a">Appendix 1</xref>, <xref ref-type="supplementary-material" rid="sup-a">Appendix Tables 2</xref>&#x2013;<xref ref-type="supplementary-material" rid="sup-a">4</xref>).</p>
<p>We generated synthetic data to replicate the two ALSPAC datasets described above (ARCADIA and CHDB). Using &#x201C;<italic>Synthpop</italic>&#x201D; in <italic>R</italic>, we generated a &#x2018;gold-standard&#x2019; dataset of identifier and attribute variables (apart from forenames and surnames) to replicate ARCADIA [<xref ref-type="bibr" rid="ref-19">19</xref>]. The dataset contained 13,281 records and was generated using sequential regression modelling based on the original ALSPAC data, using date of birth, gender, maternal age category, ethnic group, and quintile of the Index of Multiple Deprivation (<xref ref-type="supplementary-material" rid="sup-a">Appendix 3</xref>, <xref ref-type="supplementary-material" rid="sup-a">Appendix Tables 5</xref>, <xref ref-type="supplementary-material" rid="sup-a">6</xref>). Given the small number of people with non-white ethnicity, not all combinations of maternal age and ethnicity exist in the original data. We used a rejection sampling mechanism to ensure synthesised dataset did not generate combinations of attribute variables that did not appear in the original study [<xref ref-type="bibr" rid="ref-43">43</xref>]. Detailed methodology used to synthesise attributes, identifiers and names are described in <xref ref-type="supplementary-material" rid="sup-a">Appendix 2</xref> and <xref ref-type="supplementary-material" rid="sup-a">3</xref>.</p>
</sec>
<sec>
<title>Data corruption</title>
<p>Four different data corruption approaches were used to examine how results were affected by differences in the types, co-occurrences and dependencies of errors that were introduced to the synthetic data:</p>
<list list-type="order">
<list-item><p>Error types: We varied whether or not the synthetic data had the same types of errors as original data.</p></list-item>
<list-item><p>Error Field co-occurrence pattern: We varied whether or not the synthetic data had the same pattern of co-occurrence at the field level (e.g. 5% of errors co-occur in G1 forename and surname).</p></list-item>
<list-item><p>Error type co-occurrence pattern: We varied whether synthetic data had the same pattern of co-occurrence of errors at field and type level. For example, 5% of errors co-occurred in G1 forename and G1 surname; 30% of the error co-occurrence is a random name replacement, and 70% of the co-occurrence is a forename variant error with random surname replacement).</p></list-item>
<list-item><p>Error-attribute variable dependency: We varied whether or not the error rates were dependent on attribute variables (in our case, maternal age and ethnic group).</p></list-item>
</list>
</sec>
<sec>
<title>Generating linkage files</title>
<p>Under each of the four scenarios below, we created five synthetic datasets to examine differential impact of error distribution and characteristics on linkage.</p>
<list list-type="order">
<list-item><p>Scenario 1: Error rates were based on known values derived directly from the original data source. We specified the error rate for each identifier. We allowed identifier error rates to vary according to maternal age and ethnic group. Identifier errors were of the same types as in the original (e.g. 95% surname errors were random replacements). We used the same error co-occurrence patterns as the original data.</p></list-item>
<list-item><p>Scenario 2: Error rates were assumed to be unknown but were assumed to be dependent on maternal age and ethnic group. Identifier errors were restricted to random replacements. We did not allow errors to co-occur in this scenario.</p></list-item>
<list-item><p>Scenario 3: Error rates were assumed to be unknown and were assumed to be independent of attribute characteristics (i.e. constant across maternal age and ethnicity). Identifier errors were of the same types as in the original data but the error co-occurrence pattern was assumed to be unknown.</p></list-item>
<list-item><p>Scenario 4: Error rates were assumed to be unknown but were assumed to be independent of attribute characteristics. In this scenario, we assumed that identifier error rates were constant across maternal age and ethnicity. Identifier error types were randomly assigned. We did not allow errors to co-occur in this scenario.</p></list-item>
</list>
</sec>
<sec>
<title>Deriving linkage quality metrics</title>
<p>Using the three linkage methods described, we linked the five synthetic gold-standard datasets to each of the corrupted synthetic datasets in the four scenarios. Since we had generated these data ourselves, we knew the true match status of each record pair. We were therefore able to evaluate the quality of each linkage method by deriving the rates of missed matches (true links that were matched) and false matches (records that were linked to the wrong individual) for each linkage method. Estimates were averaged over the five synthetic datasets. We then compared these results with linkage error rates derived from the original source data.</p>
</sec>
</sec>
<sec>
<title>Results</title>
<sec>
<title>Linkage results</title>
<p>There were 13,281 records that linked between the ARCADIA and CHDB datasets based on the gold-standard ALSPAC data. Using deterministic linkage, 12,673 individuals were linked. The number of linked records ranged from 12,920 with probabilistic linkage using similarity scores for comparing names, to 12,962 with probabilistic linkage using term frequency adjustments for comparing names (<xref ref-type="table" rid="table-4">Table 4</xref>). Rates of errors (both missed matches and false matches) were lower using probabilistic compared with deterministic linkage, and lowest with the addition of term frequency adjustment. All results presented for the synthetic linkages were averaged over 5 synthetic datasets.</p>
<table-wrap id="table-3">
<label>Table 3</label><caption><title>Identifier error rates introduced to synthetic linking files. No errors were introduced to sex or date of birth, and no missing values were introduced</title></caption>
<table frame="hsides" rules="groups">
<col width="20%"/>
<col width="30%"/>
<col width="20%"/>
<col width="10%"/>
<col width="10%"/>
<col width="10%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"></td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>G1 Surname</bold>^</td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>G1 Forename</bold>*</td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>G0 Surname</bold>^</td>
</tr>
<tr>
<td align="left" valign="top">Scenario 1: Known error rates</td>
<td align="left" valign="top"><bold>Error Co-occurring Patterns (% of all records)</bold></td>
<td align="center" valign="top"><bold>Maternal Age</bold></td>
<td align="center" valign="top">% error</td>
<td align="center" valign="top">% error</td>
<td align="center" valign="top">% error</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">G0 surname &#x0026; G1 surname &#x0026; G1 forename (0.30%)</td>
<td align="center" valign="top">&#x003c;20</td>
<td align="center" valign="top">11.1</td>
<td align="center" valign="top">6.6</td>
<td align="center" valign="top">36.7</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">G1 surname &#x0026; G1 forename (0.43%)</td>
<td align="center" valign="top">20-29</td>
<td align="center" valign="top">6.0</td>
<td align="center" valign="top">11.5</td>
<td align="center" valign="top">19.7</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">G1 surname &#x0026; G0 surname (1.90%)</td>
<td align="center" valign="top">30-39</td>
<td align="center" valign="top">4.2</td>
<td align="center" valign="top">14.5</td>
<td align="center" valign="top">12.0</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">G1 forename &#x0026; G0 surname (1.90%)</td>
<td align="center" valign="top">40+</td>
<td align="center" valign="top">5.8</td>
<td align="center" valign="top">15.6</td>
<td align="center" valign="top">13.6</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Missing</td>
<td align="center" valign="top">7.5</td>
<td align="center" valign="top">11.5</td>
<td align="center" valign="top">16.5</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top"><bold>Ethnic group</bold></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">White</td>
<td align="center" valign="top">5.6</td>
<td align="center" valign="top">13.0</td>
<td align="center" valign="top">17.5</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Black</td>
<td align="center" valign="top">8.8</td>
<td align="center" valign="top">13.6</td>
<td align="center" valign="top">22.4</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Asian</td>
<td align="center" valign="top">1.9</td>
<td align="center" valign="top">13.3</td>
<td align="center" valign="top">8.6</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Other</td>
<td align="center" valign="top">12.2</td>
<td align="center" valign="top">14.9</td>
<td align="center" valign="top">18.9</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Missing</td>
<td align="center" valign="top">5.7</td>
<td align="center" valign="top">9.2</td>
<td align="center" valign="top">17.7</td>
</tr>
<tr>
<td align="left" valign="top">Scenario 2: Estimated error rates#</td>
<td align="left" valign="top">No co-occurring errors</td>
<td align="center" valign="top"><bold>Maternal Age</bold></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">&#x003c;20</td>
<td align="center" valign="top">7.2</td>
<td align="center" valign="top">13.1</td>
<td align="center" valign="top">4.7</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">20-29</td>
<td align="center" valign="top">4.6</td>
<td align="center" valign="top">9.4</td>
<td align="center" valign="top">5.7</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">30-39</td>
<td align="center" valign="top">4.0</td>
<td align="center" valign="top">8.4</td>
<td align="center" valign="top">6.2</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">40+</td>
<td align="center" valign="top">6.1</td>
<td align="center" valign="top">5.7</td>
<td align="center" valign="top">6.6</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Missing</td>
<td align="center" valign="top">6.1</td>
<td align="center" valign="top">7.7</td>
<td align="center" valign="top">5.4</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top"><bold>Ethnic group</bold></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">White</td>
<td align="center" valign="top">4.4</td>
<td align="center" valign="top">9.2</td>
<td align="center" valign="top">5.9</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Black</td>
<td align="center" valign="top">7.2</td>
<td align="center" valign="top">14.3</td>
<td align="center" valign="top">9.2</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Asian</td>
<td align="center" valign="top">4.1</td>
<td align="center" valign="top">6.8</td>
<td align="center" valign="top">7.4</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Other</td>
<td align="center" valign="top">16.4</td>
<td align="center" valign="top">10.5</td>
<td align="center" valign="top">7.3</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top"></td>
<td align="center" valign="top">Missing</td>
<td align="center" valign="top">5.0</td>
<td align="center" valign="top">8.5</td>
<td align="center" valign="top">5.2</td>
</tr>
<tr>
<td align="left" valign="top">Scenario 3: Independent error rates</td>
<td align="left" valign="top">G0 surname &#x0026; G1 surname (2.30%)</td>
<td align="center" valign="top"></td>
<td align="center" valign="top">5.0</td>
<td align="center" valign="top">10.0</td>
<td align="center" valign="top">15.0</td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">G1 forename &#x0026; G1 surname (1.00%)</td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top"></td>
<td align="left" valign="top">G1 forename &#x0026; G0 surname (0.75%)</td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
<td align="center" valign="top"></td>
</tr>
<tr>
<td align="left" valign="top">Scenario 4: Independent error rates</td>
<td align="left" valign="top">No co-occurring errors</td>
<td align="center" valign="top"></td>
<td align="center" valign="top">5.0</td>
<td align="center" valign="top">10.0</td>
<td align="center" valign="top">15.0</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>*73% of errors were name variants (e.g. Sam for Samuel, Becky for Rebecca); 14% were typographical errors (e.g. insertions/deletions); 7% were due to one dataset recording multiple first names (e.g. Lisa Marie versus Lisa), 6% were completely different names.</p>
<p>^3% of the errors were due to the gold-standard dataset having two surnames (e.g. Harron Kent) and the linking file only having the first name (Harron); 2% were where the gold-standard had two surnames but the linking file only has the second name (Kent).</p>
<p># Error rates presented are based on the relative risk of identifier errors according to attribute variables in <xref ref-type="supplementary-material" rid="sup-a">Appendix 1</xref>, <xref ref-type="supplementary-material" rid="sup-a">Appendix Table 4</xref>, with estimated baseline likelihood of error of 0.1 (G1 Surname), 0.15 (G1 Forename), 0.2 (G0 Surname).</p>
</table-wrap-foot>
</table-wrap>
<table-wrap id="table-4">
<label>Table 4</label><caption><title>Comparison of linkage quality metrics based on the original ALSPAC data, and synthetic data generated under three scenarios</title></caption>
<table frame="hsides" rules="groups">
<col width="30%"/>
<col width="20%"/>
<col width="10%"/>
<col width="20%"/>
<col width="20%"/>
<tbody>
<tr>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"></td>
<td align="left" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"></td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Deterministic linkage</bold></td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Probabilistic linkage with similarity scores</bold></td>
<td align="center" style="border-top: solid 1pt; border-bottom: solid 1pt;" valign="middle"><bold>Probabilistic linkage with similarity scores and term frequency adjustment</bold></td>
</tr>
<tr>
<td rowspan="3" align="left" valign="top"><italic><bold>Original data</bold></italic></td>
<td align="center" valign="top"><italic>n linked records</italic></td>
<td align="center" valign="top"><italic>12,673</italic></td>
<td align="center" valign="top"><italic>12,920</italic></td>
<td align="center" valign="top"><italic>12,962</italic></td>
</tr>
<tr>
<td align="center" valign="top"><italic>Missed match rate</italic>*</td>
<td align="center" valign="top"><italic>4.59%</italic></td>
<td align="center" valign="top"><italic>2.61%</italic></td>
<td align="center" valign="top"><italic>2.40%</italic></td>
</tr>
<tr>
<td align="center" valign="top"><italic>False match rate</italic>*</td>
<td align="center" valign="top"><italic>0.23%</italic></td>
<td align="center" valign="top"><italic>0.12%</italic></td>
<td align="center" valign="top"><italic>0.05%</italic></td>
</tr>
<tr>
<td rowspan="3" align="left" valign="top"><bold>Synthetic data &#x2013; known error rates, dependent on attributes, original error co-occurence</bold><sup>1</sup></td>
<td align="center" valign="top">n linked records</td>
<td align="center" valign="top">12,656</td>
<td align="center" valign="top">12,718</td>
<td align="center" valign="top">12,712</td>
</tr>
<tr>
<td align="center" valign="top">Missed match rate</td>
<td align="center" valign="top">4.72%</td>
<td align="center" valign="top">4.26%</td>
<td align="center" valign="top">4.29%</td>
</tr>
<tr>
<td align="center" valign="top">False match rate</td>
<td align="center" valign="top">0.29%</td>
<td align="center" valign="top">0.32%</td>
<td align="center" valign="top">0.16%</td>
</tr>
<tr>
<td rowspan="3" align="left" valign="top"><bold>Synthetic data &#x2013; guessed error rates, dependent on attributes, no error co-occurence</bold><sup>2</sup></td>
<td align="center" valign="top">n linked records</td>
<td align="center" valign="top">13,274</td>
<td align="center" valign="top">13,276</td>
<td align="center" valign="top">13,279</td>
</tr>
<tr>
<td align="center" valign="top">Missed match rate</td>
<td align="center" valign="top">0.05%</td>
<td align="center" valign="top">0.04%</td>
<td align="center" valign="top">0.02%</td>
</tr>
<tr>
<td align="center" valign="top">False match rate</td>
<td align="center" valign="top">0.09%</td>
<td align="center" valign="top">0.12%</td>
<td align="center" valign="top">0.07%</td>
</tr>
<tr>
<td rowspan="3" align="left" valign="top"><bold>Synthetic data &#x2013; guessed error rates, independent of attributes, assumed pattern of error co-occurence</bold><sup>3</sup></td>
<td align="center" valign="top">n linked records</td>
<td align="center" valign="top">12,746</td>
<td align="center" valign="top">12,817</td>
<td align="center" valign="top">12,809</td>
</tr>
<tr>
<td align="center" valign="top">Missed match rate</td>
<td align="center" valign="top">4.04%</td>
<td align="center" valign="top">3.50%</td>
<td align="center" valign="top">3.56%</td>
</tr>
<tr>
<td align="center" valign="top">False match rate</td>
<td align="center" valign="top">0.25%</td>
<td align="center" valign="top">0.24%</td>
<td align="center" valign="top">0.13%</td>
</tr>
<tr>
<td rowspan="3" align="left" valign="top"><bold>Synthetic data &#x2013; guessed error rates, independent of attributes, no error co-occurence</bold><sup>4</sup></td>
<td align="center" valign="top">n linked records</td>
<td align="center" valign="top">13,266</td>
<td align="center" valign="top">13,277</td>
<td align="center" valign="top">13,279</td>
</tr>
<tr>
<td align="center" valign="top">Missed match rate</td>
<td align="center" valign="top">0.12%</td>
<td align="center" valign="top">0.03%</td>
<td align="center" valign="top">0.02%</td>
</tr>
<tr>
<td align="center" valign="top">False match rate</td>
<td align="center" valign="top">0.10%</td>
<td align="center" valign="top">0.10%</td>
<td align="center" valign="top">0.04%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><sup>1</sup>Scenario 1: Error rates were specified correctly, based on the original data source (<xref ref-type="table" rid="table-3">Table 3</xref>).</p>
<p><sup>2</sup>Scenario 2: Error rates were guessed, and were allowed to vary according to maternal age and ethnicity (<xref ref-type="table" rid="table-3">Table 3</xref>).</p>
<p><sup>3</sup>Scenario 3: Error rates were guessed and were assumed to be unrelated to attribute characteristics, with estimated error co-occurrence patterns: 15% G0 surname-G1 surname, 10% G1 forename-G1 surname, 5% G1 surname-G1 forename. Types of error when errors co-occur: G0 surname and G1 surname errors = random replacement, G1 forename and G1 surname errors = random replacement (surname) <italic>+</italic> 30% typo, 70% forename variant (<xref ref-type="table" rid="table-3">Table 3</xref>).</p>
<p><sup>4</sup>Scenario 4: Error rates were guessed and were assumed to be unrelated to attribute characteristics, and no restriction on error co-occurring patterns (<xref ref-type="table" rid="table-3">Table 3</xref>).</p>
<p><sup>*</sup>missed match rate = % of true matches that were not identified, i.e. 1-sensitivity.</p>
<p><sup>*</sup>false match rate = % of linked records that were not true matches, i.e. 1-positive predictive value.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Linkage quality metrics</title>
<p>All of the synthetic datasets broadly replicated the same pattern seen in the original data linkage, i.e. that rates of missed matches were lower than rates of false matches, and that probabilistic linkage with similarity scores and term frequency adjustment had the best performance (<xref ref-type="table" rid="table-4">Table 4</xref>). Scenarios 1 and 3 result in comparable linkage error rates compared to the original linkage, and successfully demonstrated that probabilistic linkage was able to reduce both false-matches and missed-matches compared with deterministic linkage.</p>
<p>Across the deterministic linkages, scenarios 1 and 3 had more comparable linkage error rates, with an absolute difference of 0.13&#x2013;0.55% for missed matches, and 0.00-0.04% for false matches. Linkages for scenarios 2 and 4 had larger variations of linkage errors compared to the original, with a difference of 4.05&#x2013;4.12% for missed matches, and 0.15&#x2013;0.16% for false matches.</p>
<p>In scenario 1, linkage error rates were slightly over-estimated in the synthetic data, by 0.55&#x2013;2.02% for missed matches and 0.04&#x2013;0.11% for false matches (<xref ref-type="table" rid="table-4">Table 4</xref>). The difference in estimation was similar in scenario 3, at 0.13&#x2013;1.26% for missed matches and 0.00-0.12% for false matches.</p>
<p>In scenario 2, linkage error rates were under-estimated in the synthetic data, by 2.20&#x2013;4.12% for missed matches and 0.00&#x2013;0.16% for false matches (<xref ref-type="table" rid="table-4">Table 4</xref>). Under-estimation was found to a similar extent in scenario 4, by 2.21&#x2013;4.05% for missed matches and 0.01&#x2013;0.15% for false matches.</p>
</sec>
<sec>
<title>Missed match and false match characteristics</title>
<p>In the original linkage, of the 346 true matches missed by probabilistic linkages with similarity scores, 65.0% were those where there was agreement on forename, date of birth and gender, but disagreement on surname and mother&#x2019;s surname. These missed matches affected people of different ethnicities and genders similarly and affected younger mothers more than older mothers. Use of term frequency adjustment further reduced missed matches to 319. These missed matches appeared to correspond to cases in which both the mother and the child changed their surname between data collection waves (rather than being due to typographical errors). The second most common missed match pattern occurred for records with disagreement on forename and mother&#x2019;s surname, with 13.6% in probabilistic linkage with similarity scores, and 13.8% with term frequency adjustments. These missed matches appeared to correspond to mothers changing their surnames, and children providing alternative names or derivatives at different data collection waves.</p>
<p>Missed match rates were comparable to the original linkage in scenarios 1 and 3. The disagreement pattern of missed matches were also similar to the original linkage (<xref ref-type="supplementary-material" rid="sup-a">Appendix Table 8</xref>).</p>
<p>Compared to the original linkage, missed matches in scenario 3 had similar distributions of disagreement patterns. With probabilistic linkage with similarity scores, 61% of missed matches disagreed on surname and mother&#x2019;s surname; with term frequency adjustment, 58% missed matches disagreed on surname and mother&#x2019;s surname. In both probabilistic linkage with similarity scores and term frequency adjustment, 19.3% missed matches disagreed on forename and mother&#x2019;s surname.</p>
<p>Comparing to the original linkage, missed matches in scenario 1 had a lower proportion of disagreements on surname and mother&#x2019;s surname with 42.3% for probabilistic linkage and 40.2% for term frequency adjustments. Higher proportions of missed matches disagreed on forename and mother&#x2019;s surname, with 37.9% for probabilistic linkage, and 37.8% for term frequency adjustments.</p>
<p>False-match rates were low in the original linkage and synthetic linkages. The higher rate of false-matches with deterministic linkage was predominantly explained by the 54% of record pairs that agreed on surname (both mother and child), sex and date of birth, but disagreed on forename. This was followed by 39% of false-matches in pairs that disagreed on date of birth only (<xref ref-type="supplementary-material" rid="sup-a">Appendix Table 9</xref>). In the deterministic linkages using synthetic datasets, similar patterns and proportions of false-matches records were replicated, where 60.0% of false-matches disagreed only on forename, and 34.9% disagreed on date of birth only. As date of birth was recorded with high accuracy in the ALSPAC data, these pairs were (correctly) not accepted as links by the probabilistic strategies (since disagreement on date of birth conferred a large penalty to the match weight).</p>
<p>In terms of missed match rates, false match rates, and characteristics of missed matches, we were able to best produce linkages most similar to original linkage in scenario 3. This demonstrates that replicating error types and co-occurrence patterns (even if the co-occurrence patterns are estimated) without incorporating dependencies between error and attribute is sufficient to produce realistic synthetic data. Further incorporating information about dependencies between errors and attribute (with true error rates), and error co-occurrence patterns (scenario 1) did not produce substantially more realistic linkages.</p>
<p>Conversely, retaining error and attribute dependency without incorporating error types and co-occurrence (scenario 2) performed similarly to when identifier error rates were assumed to be independent (scenario 4).</p>
</sec>
</sec>
<sec>
<title>Discussion</title>
<p>We provide a generalisable and open-source framework for generating synthetic identifier data that can be used to facilitate development and evaluation of improved linkage methodologies. We show how this framework can be implemented and provide a means of producing corrupted datasets that can be used for linkage development and a complete &#x2018;gold standard&#x2019; file that can be used for linkage validation. We generated synthetic ALSPAC identifier datasets, which are freely available for legitimate users on request to the authors: the intention is that these data, with known characteristics, can be used for the development and comparative benchmarking of different linkage approaches.</p>
<p>Our framework builds on previous methodological work aiming to generate synthetic identifier data for use in data linkage [<xref ref-type="bibr" rid="ref-3">3</xref>, <xref ref-type="bibr" rid="ref-13">13</xref>]. We extended previous work by overcoming the assumption of independence of identifier errors through explicitly incorporating the associations between identifier errors and attribute variables. If accurate information on the joint distribution of identifiers and identifier errors were available, there would be no need to include information on their dependencies with attributes. However, evaluating linkage quality according to attributes such as age, sex and ethnicity is convenient and intuitive, and knowledge of how linkage errors are typically distributed amongst these subgroups can be easily incorporated into synthetic data generators [<xref ref-type="bibr" rid="ref-44">44</xref>]. Our findings comparing linkage quality metrics for synthetic data generated under different scenarios highlight that preserving error types and co-occurrence patterns is vital for generating a dataset that accurately represents real-world data and that can be meaningfully used to evaluate linkage algorithms, and is useful when incorporating the dependencies between identifier errors and attributes is not easily achievable [<xref ref-type="bibr" rid="ref-16">16</xref>]. This framework can be used to test linkages between more than 2 datasets.</p>
<p>The strengths of our study include the use of gold-standard data from a large cohort study that was used to assess the performance of synthetic data for deriving linkage quality metrics. We compared a range of linkage methods and different scenarios under which the synthetic data were generated. It is likely that the errors observed in these data are representative of those occurring in other administrative and research datasets. We acknowledge that the exemplar ALSPAC dataset is predominately of a White UK population, and recommend that other cultural, geographic and time-point specific alternatives are generated in order to avoid any unintended bias in linkage algorithm development (i.e. to factor in error patterns that exist yet were not observed in the ALSPAC data). However, synthetic data generators such as this one should give users the ability to alter the identifier error rates, types and co-occurrence patterns according to their particular data context. This allows for any uncertainty to be explored, by using a range of error rates and patterns to investigate how results may vary. This could be used to help inform choice of linkage strategy: for example, it could tell us that a simple deterministic approach might generate results of sufficiently high quality if identifier error rates are low, whilst a more sophisticated and resource intensive probabilistic approach might be more suitable in settings where identifier error rates are high. It could also point to possible improvements in algorithms: in our example, all three linkage algorithms failed to identify true matches where there was a disagreement on surname and mother&#x2019;s surname: better handling of name specific characteristics, such as double-barrel surnames, could go some way to mitigating this problem. Using synthetic data could also be used to provide a plausible range of linkage error rates that are likely to arise, given different assumptions about the levels of identifier errors. Under these assumptions, researchers can explore the sensitivity of their linkage approach by assessing the impact of including or excluding certain error-prone identifiers on linkage rates. This is particularly relevant for longitudinal population data, where richer insight in the variation of identifier errors is more observable, researchers could demonstrate with which data sets the original data could be best linked. Researchers can then use different methods to account for linkage error rates within analysis, e.g. quantitative bias analysis to explore the extent to which results of analyses might be affected by linkage error rates [<xref ref-type="bibr" rid="ref-45">45</xref>]. Synthetic identifier data would be particularly useful for evaluating the quality of privacy preserving linkage techniques, where access to identifiers in the clear is not permitted. Currently, access to real data is needed generate synthetic identifier data. Alternative approaches, such as estimating parameters from existing publications, could provide information sufficient to assess linkage quality to a certain extent (such as Scenario 3, where error rates for each variable were educated guesses). However, this approach might be blind to error characteristics, error co-occurring patterns, and error inter-dependencies that may underlie specific data sources. As these synthetic data would be used to evaluate the validity and utility of the linkages, using mis-specified models, or multiple proposed synthetic models would confer to challenges in data governance. Our proposed framework, while seemingly relying on higher involvement of the data owners, has the advantage of giving more control to data owners, and presents as a more pragmatic approach to drive change.</p>
<p>Limitations of our study are that we only had one gold-standard dataset with which to evaluate the performance of the synthetic data and therefore our testing of dependencies is based on information about a specific population group; further evaluations should be conducted on other datasets with varying proportions of missingness in identifiers. Our name generation mechanism takes advantage of the small sample size and low cardinality of name distributions in ALSPAC (4,000&#x2013;7,000 distinct forename and surname terms). Replication using the same method would require a more diverse name dictionary. The key advantage of generating realistic names with name dictionaries, (versus string or number sequences), is the potential to better reflect dimensions of name characteristics that are non-metricized and may associate with error distributions by attributes. The current name generation mechanism did not fully preserve name clusters and name-specific characteristics, such as word length, hyphens or number of terms per name [<xref ref-type="bibr" rid="ref-46">46</xref>]. Our framework could be extended in several ways, including by adding in additional variable types, error types and error co-occurrence patterns, by allowing the generation of data at the household level, or for multiple generations to capture between record dependencies. More sophisticated synthetic identifier data might include more nuanced errors (i.e. specifying the most likely letter transpositions based on keyboard strokes, or introducing date-specific errors such as recording today&#x2019;s date). However, these nuances would only be required if the linkage algorithm that was being evaluated was tailored towards resolving these specific sorts of errors. A further problem is on assessing how accurate the error type and co-occurrence pattern has to be for the generated synthetic data to be considered similar enough to reliably test the proposed linkage methods. Our study offers an approach to start investigating this idea more structurally, by contrasting multiple data corruption scenarios. Further investigations on this direction would allow us to be more confident in our comparisons.</p>
<p>Our framework provides a novel and generalisable mechanism for developing and benchmarking record linkage algorithms, which is protective of public privacy and avoids assumptions that errors in personal identifiers are independent of the other identifiers and attribute data. Our findings show that replicating dependencies between attribute values (e.g. ethnicity), values of identifiers (e.g. name), and errors in identifiers (e.g. missing values, typographical errors or changes over time) and its patterns enables generation of realistic synthetic data that can be used to evaluate different linkage methods.</p>
</sec>
<sec sec-type="supplementary-material">
<title>Supplementary Files</title>
<supplementary-material id="sup-a">
<label>Supplementary Appendices</label> 
<media mimetype="application" mime-subtype="pdf" xlink:href="ijpds-09-2389-s001.pdf"/>
</supplementary-material>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<p>Harvey Goldstein had a key role in developing this study, but sadly died prior to publication. We are very grateful to his input to this work. We would also like to thank Haoyuan Zhang for his early input to this work.</p>
<p>We are extremely grateful to all the families who took part in this study, the midwives for their help in recruiting them, and the whole ALSPAC team, which includes interviewers, computer and laboratory technicians, clerical workers, research scientists, volunteers, managers, receptionists and nurses. We particularly thank Mark Mumme for his time in producing extracts of data for this study. We thank Ruth Gilbert and James Doidge for their input to and feedback on this work.</p>
</ack>
<sec>
<title>Ethics</title>
<p>Ethical approval for the ALSPAC cohort study was obtained from the ALSPAC Ethics and Law Committee (a University of Bristol Faculty Ethics Committee) and NHS Local Research Ethics Committee(s). Informed consent for the use of data collected via questionnaires and clinics was obtained from participants following the recommendations of the ALSPAC Ethics and Law Committee at the time. This study was approved through the ALSPAC data access mechanism (ALSPAC Reference: B3002, <uri>https://proposals.epi.bristol.ac.uk/?q=node/127384</uri>). Please note that the study website contains details of all the data that is available through a fully searchable data dictionary and variable search tool (<uri>http://www.bristol.ac.uk/alspac/researchers/our-data</uri>).</p>
</sec>
<sec>
<title>Funding</title>
<p>The UK Medical Research Council and Wellcome (Grant ref: 217065/Z/19/Z) and the University of Bristol provide core support for ALSPAC. This publication is the work of the authors and Katie Harron and Andy Boyd will serve as guarantors for the contents of this paper. This research was funded in whole, or in part, by the Wellcome Trust [212953/Z/18/Z]. For the purpose of Open Access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission.</p>
</sec>
<sec>
<title>Data availability</title>
<p>Synthetic ALSPAC data can be downloaded on UCL Research Data Repository (doi:10.5522/04/25921408). Original ALSPAC data can be requested from ALSPAC website.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="ref-1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Dibben</surname> <given-names>C</given-names></string-name>, <string-name><surname>Boyd</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hjern</surname> <given-names>A</given-names></string-name>, <string-name><surname>Azimaee</surname> <given-names>M</given-names></string-name>, <string-name><surname>Barreto</surname> <given-names>ML</given-names></string-name>, <etal>et al</etal>. <article-title>Challenges in administrative data linkage for research</article-title>. <source>Big Data &#x0026; Society</source>. <year>2017</year> <month>Dec</month>;<volume>4</volume>(<issue>2</issue>):<fpage>205395171774567</fpage>. <pub-id pub-id-type="doi">10.1177/2053951717745678</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Jorm</surname> <given-names>L</given-names></string-name>. <article-title>Routinely collected data as a strategic resource for research: priorities for methods and workforce</article-title>. <source>Public Health Res Pr</source> [Internet]. <year>2015</year> [cited <year>2023</year> <month>Dec</month> <day>7</day>];<volume>25</volume>(<issue>4</issue>). Available from: <uri>http://www.phrp.com.au/issues/september-2015-volume-25-issue-4/routinely-collected-data-as-a-strategic-resource-for-research-priorities-for-methods-and-workforce/</uri> <pub-id pub-id-type="doi">10.17061/phrp2541540</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>3</label><mixed-citation publication-type="book"><string-name><surname>Christen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Vatsalan</surname> <given-names>D</given-names></string-name>. <chapter-title>Flexible and extensible generation and corruption of personal data</chapter-title>. In: <source>Proceedings of the 22nd ACM international conference on Conference on information &#x0026; knowledge management - CIKM &#x2019;13 [Internet]</source>. <publisher-loc>San Francisco, California, USA</publisher-loc>: <publisher-name>ACM Press</publisher-name>; <year>2013</year> [cited <year>2023</year> <month>Dec</month> <day>7</day>]. p. <fpage>1165</fpage>&#x2013;<lpage>8</lpage>. Available from: <uri>http://dl.acm.org/citation.cfm?doid=2505515.2507815</uri>. <pub-id pub-id-type="doi">10.1145/2505515.2507815</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>4</label><mixed-citation publication-type="journal"><string-name><surname>Kelman</surname> <given-names>CW</given-names></string-name>, <string-name><surname>Bass</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Holman</surname> <given-names>CDJ</given-names></string-name>. <article-title>Research use of linked health data&#x2013;a best practice protocol</article-title>. <source>Aust N Z J Public Health</source>. <year>2002</year>;<volume>26</volume>(<issue>3</issue>):<fpage>251</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>5</label><mixed-citation publication-type="journal"><string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wade</surname> <given-names>A</given-names></string-name>, <string-name><surname>Muller-Pebody</surname> <given-names>B</given-names></string-name>, <string-name><surname>Goldstein</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gilbert</surname> <given-names>R</given-names></string-name>. <article-title>Opening the black box of record linkage</article-title>. <source>J Epidemiol Community Health</source>. <year>2012 Dec</year>;<volume>66</volume>(<issue>12</issue>):<fpage>1198</fpage>. <pub-id pub-id-type="doi">10.1136/jech-2012-201376</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>6</label><mixed-citation publication-type="book"><string-name><surname>Christen</surname> <given-names>P</given-names></string-name>. <chapter-title>Probabilistic Data Generation for Deduplication and Data Linkage</chapter-title>. In: <string-name><surname>Gallagher</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hogan</surname> <given-names>JP</given-names></string-name>, <string-name><surname>Maire</surname> <given-names>F</given-names></string-name>, editors. <source>Intelligent Data Engineering and Automated Learning - IDEAL 2005 [Internet]</source>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer Berlin Heidelberg</publisher-name>; <year>2005</year> [cited <year>2023</year> <month>Dec</month> <day>7</day>]. p. <fpage>109</fpage>&#x2013;<lpage>16</lpage>. (<string-name><surname>Hutchison</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kanade</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kittler</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kleinberg</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Mattern</surname> <given-names>F</given-names></string-name>, <string-name><surname>Mitchell</surname> <given-names>JC</given-names></string-name>, <etal>et al</etal>., editors. <collab>Lecture Notes in Computer Science</collab>; vol. <volume>3578</volume>). Available from: <uri>http://link.springer.com/10.1007/11508069_15</uri>. <pub-id pub-id-type="doi">10.1007/11508069_15</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>7</label><mixed-citation publication-type="journal"><string-name><surname>Ferrante</surname> <given-names>A</given-names></string-name>, <string-name><surname>Boyd</surname> <given-names>J</given-names></string-name>. <article-title>A transparent and transportable methodology for evaluating Data Linkage software</article-title>. <source>J Biomed Inform</source>. <year>2012</year> <month>Feb</month>;<volume>45</volume>(<issue>1</issue>):<fpage>165</fpage>&#x2013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2011.10.006</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>8</label><mixed-citation publication-type="journal"><string-name><surname>Nowok</surname> <given-names>B</given-names></string-name>, <string-name><surname>Raab</surname> <given-names>GM</given-names></string-name>, <string-name><surname>Dibben</surname> <given-names>C</given-names></string-name>. <article-title>Providing bespoke synthetic data for the UK Longitudinal Studies and other sensitive data with the synthpop package for R 1</article-title>. <source>Statistical Journal of the IAOS</source>. <year>2017</year> <month>Jan</month> <day>1</day>;<volume>33</volume>(<issue>3</issue>):<fpage>785</fpage>&#x2013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.3233/SJI-150153</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>9</label><mixed-citation publication-type="journal"><string-name><surname>Kokosi</surname> <given-names>T</given-names></string-name>, <string-name><surname>De Stavola</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mitra</surname> <given-names>R</given-names></string-name>, <string-name><surname>Frayling</surname> <given-names>L</given-names></string-name>, <string-name><surname>Doherty</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dove</surname> <given-names>I</given-names></string-name>, <etal>et al</etal>. <article-title>An overview on synthetic administrative data for research</article-title>. <source>IJPDS</source> [Internet]. <year>2022</year> <month>May</month> <day>23</day> [cited <year>2022</year> <month>Jul</month> <day>7</day>];<volume>7</volume>(<issue>1</issue>). Available from: <uri>https://ijpds.org/article/view/1727</uri>. <pub-id pub-id-type="doi">10.23889/ijpds.v7i1.1727</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>10</label><mixed-citation publication-type="journal"><string-name><surname>Raghunathan</surname> <given-names>TE</given-names></string-name>. <article-title>Annual Review of Statistics and Its Application Synthetic Data</article-title>. <source>Annual Review of Statistics and Its Application</source>. <year>2021</year>;<volume>8</volume>(<issue>1</issue>):<fpage>129</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1146/annurev-statistics-040720-031848</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Doidge</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Harron</surname> <given-names>KL</given-names></string-name>. <article-title>Reflections on modern methods: linkage error bias</article-title>. <source>International Journal of Epidemiology</source>. <year>2019</year> <month>Dec</month> <day>1</day>;<volume>48</volume>(<issue>6</issue>):<fpage>2050</fpage>&#x2013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1093/ije/dyz203</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Harron</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Doidge</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Knight</surname> <given-names>HE</given-names></string-name>, <string-name><surname>Gilbert</surname> <given-names>RE</given-names></string-name>, <string-name><surname>Goldstein</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cromwell</surname> <given-names>DA</given-names></string-name>, <etal>et al</etal>. <article-title>A guide to evaluating linkage quality for the analysis of linked data</article-title>. <source>International Journal of Epidemiology</source>. <year>2017</year> <month>Oct</month> <day>1</day>;<volume>46</volume>(<issue>5</issue>):<fpage>1699</fpage>&#x2013;<lpage>710</lpage>. <pub-id pub-id-type="doi">10.1093/ije/dyx177</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>13</label><mixed-citation publication-type="book"><string-name><surname>Christen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Pudjijono</surname> <given-names>A</given-names></string-name>. <chapter-title>Accurate Synthetic Generation of Realistic Personal Information</chapter-title>. In: <string-name><surname>Theeramunkong</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kijsirikul</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cercone</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>TB</given-names></string-name>, editors. <source>Advances in Knowledge Discovery and Data Mining</source>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2009</year>. p. <fpage>507</fpage>&#x2013;<lpage>14</lpage>. (Lecture Notes in Computer Science). <pub-id pub-id-type="doi">10.1007/978-3-642-01307-2_47</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14</label><mixed-citation publication-type="book"><string-name><surname>Bohensky</surname> <given-names>M</given-names></string-name>. <chapter-title>Chapter 4: Bias in data linkage studies</chapter-title>. In: In: <string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Dibben</surname> <given-names>C</given-names></string-name>, <string-name><surname>Goldstein</surname> <given-names>H</given-names></string-name>, editors <source>Methodological Developments in Data Linkage [Internet]</source>. <publisher-name>John Wiley &#x0026; Sons, Ltd</publisher-name>; <year>2015</year> [cited <year>2023</year> <month>Dec</month> <day>7</day>]. p. <fpage>63</fpage>&#x2013;<lpage>82</lpage>. Available from: <uri>https://onlinelibrary.wiley.com/doi/abs/10.1002/9781119072454.ch4</uri>. <pub-id pub-id-type="doi">10.1002/9781119072454.ch4</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>15</label><mixed-citation publication-type="journal"><string-name><surname>Bohensky</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Jolley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Sundararajan</surname> <given-names>V</given-names></string-name>, <string-name><surname>Evans</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pilcher</surname> <given-names>DV</given-names></string-name>, <string-name><surname>Scott</surname> <given-names>I</given-names></string-name>, <etal>et al</etal>. <article-title>Data Linkage: A powerful research tool with potential problems</article-title>. <source>BMC Health Services Research</source>. <year>2010</year> <month>Dec</month> <day>22</day>;<volume>10</volume>(<issue>1</issue>):<fpage>346</fpage>. <pub-id pub-id-type="doi">10.1186/1472-6963-10-346</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>16</label><mixed-citation publication-type="journal"><string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hagger-Johnson</surname> <given-names>G</given-names></string-name>, <string-name><surname>Gilbert</surname> <given-names>R</given-names></string-name>, <string-name><surname>Goldstein</surname> <given-names>H</given-names></string-name>. <article-title>Utilising identifier error variation in linkage of large administrative data sources</article-title>. <source>BMC Medical Research Methodology</source>. <year>2017</year> <month>Feb</month> <day>7</day>;<volume>17</volume>(<issue>1</issue>):<fpage>23</fpage>. <pub-id pub-id-type="doi">10.1186/s12874-017-0306-8</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>17</label><mixed-citation publication-type="journal"><string-name><surname>Fraser</surname> <given-names>A</given-names></string-name>, <string-name><surname>Macdonald-Wallis</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tilling</surname> <given-names>K</given-names></string-name>, <string-name><surname>Boyd</surname> <given-names>A</given-names></string-name>, <string-name><surname>Golding</surname> <given-names>J</given-names></string-name>, <string-name><surname>Davey Smith</surname> <given-names>G</given-names></string-name>, <etal>et al</etal>. <article-title>Cohort Profile: the Avon Longitudinal Study of Parents and Children: ALSPAC mothers cohort</article-title>. <source>Int J Epidemiol</source>. <year>2013</year> <month>Feb</month>;<volume>42</volume>(<issue>1</issue>):<fpage>97</fpage>&#x2013;<lpage>110</lpage>. <pub-id pub-id-type="doi">10.1093/ije/dys066</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>18</label><mixed-citation publication-type="journal"><string-name><surname>Boyd</surname> <given-names>A</given-names></string-name>, <string-name><surname>Golding</surname> <given-names>J</given-names></string-name>, <string-name><surname>Macleod</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lawlor</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Fraser</surname> <given-names>A</given-names></string-name>, <string-name><surname>Henderson</surname> <given-names>J</given-names></string-name>, <etal>et al</etal>. <article-title>Cohort Profile: the &#x2019;children of the 90s&#x2019;&#x2013;the index offspring of the Avon Longitudinal Study of Parents and Children</article-title>. <source>Int J Epidemiol</source>. <year>2013</year>;<volume>42</volume>(<issue>1</issue>):<fpage>111</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1093/ije/dys064</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>19</label><mixed-citation publication-type="journal"><string-name><surname>Nowok</surname> <given-names>B</given-names></string-name>, <string-name><surname>Raab</surname> <given-names>GM</given-names></string-name>, <string-name><surname>Dibben</surname> <given-names>C</given-names></string-name>. <article-title>synthpop: Bespoke Creation of Synthetic Data in R</article-title>. <source>Journal of Statistical Software</source>. <year>2016</year> <month>Oct</month> <day>28</day>;<volume>74</volume>:<fpage>1</fpage>&#x2013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v074.i11</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>20</label><mixed-citation publication-type="journal"><string-name><surname>Linacre</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lindsay</surname> <given-names>S</given-names></string-name>, <string-name><surname>Manassis</surname> <given-names>T</given-names></string-name>, <string-name><surname>Slade</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hepworth</surname> <given-names>T</given-names></string-name>. <article-title>Splink: Free software for probabilistic record linkage at scale</article-title>. <source>International Journal of Population Data Science</source> [Internet]. <year>2022</year> <month>Aug</month> <day>25</day> [cited 2023 Jun 5];<volume>7</volume>(<issue>3</issue>). Available from: <uri>https://ijpds.org/article/view/1794</uri>. <pub-id pub-id-type="doi">10.23889/ijpds.v7i3.1794</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Piantadosi</surname> <given-names>ST</given-names></string-name>. <article-title>Zipf&#x2019;s word frequency law in natural language: A critical review and future directions</article-title>. <source>Psychon Bull Rev</source>. <year>2014</year> <month>Oct</month> <day>1</day>;<volume>21</volume>(<issue>5</issue>):<fpage>1112</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.3758/s13423-014-0585-6</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>22</label><mixed-citation publication-type="website"><article-title>CLOSER-resource-NHS-Numbers-and-their-management-systems.pdf [Internet]</article-title>. [cited <year>2024</year> <month>Jan</month> <day>3</day>]. Available from: <uri>https://www.closer.ac.uk/wp-content/uploads/CLOSER-resource-NHS-Numbers-and-their-management-systems.pdf</uri>.</mixed-citation></ref>
<ref id="ref-23"><label>23</label><mixed-citation publication-type="website"><collab>Office for National Statistics</collab>. <article-title>Baby names in England and Wales statistical bulletins</article-title>. [cited <year>2024</year> <month>Jan</month> <day>3</day>]. <source>Office for National Statistics</source>. Available from: <uri>https://www.ons.gov.uk/peoplepopulationandcommunity/birthsdeathsandmarriages/livebirths/bulletins/babynamesenglandandwales/previousReleases</uri>.</mixed-citation></ref>
<ref id="ref-24"><label>24</label><mixed-citation publication-type="website"><collab>National Records of Scotland. National Records of Scotland</collab>. <article-title>National Records of Scotland</article-title>; <year>2013</year> [cited <year>2024</year> <month>Jan</month> <day>3</day>]. <source>National Records of Scotland</source>. Available from: <uri>https://www.nrscotland.gov.uk/statistics-and-data/statistics/statistics-by-theme/vital-events/names/babies-first-names/</uri></mixed-citation></ref>
<ref id="ref-25"><label>25</label><mixed-citation publication-type="website"><string-name><surname>Betebenner</surname> <given-names>DW</given-names></string-name>. <article-title>randomNames: Function for Generating Random Names and a Dataset</article-title>. [Internet]. <year>2021</year>. Available from: <uri>https://cran.r-project.org/package=randomNames</uri>.</mixed-citation></ref>
<ref id="ref-26"><label>26</label><mixed-citation publication-type="journal"><string-name><surname>McElduff</surname> <given-names>F</given-names></string-name>, <string-name><surname>Mateos</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wade</surname> <given-names>A</given-names></string-name>, <string-name><surname>Borja</surname> <given-names>MC</given-names></string-name>. <article-title>What&#x2019;s in a name? The frequency and geographic distributions of UK surnames</article-title>. <source>Significance</source>. <year>2008</year>;<volume>5</volume>(<issue>4</issue>):<fpage>189</fpage>&#x2013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1111/j.1740-9713.2008.00332.x</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>27</label><mixed-citation publication-type="journal"><string-name><surname>Danesh</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gault</surname> <given-names>S</given-names></string-name>, <string-name><surname>Semmence</surname> <given-names>J</given-names></string-name>, <string-name><surname>Appleby</surname> <given-names>P</given-names></string-name>, <string-name><surname>Peto</surname> <given-names>R</given-names></string-name>. <article-title>Postcodes as useful markers of social class: population based study in 26 000 British households</article-title>. <source>BMJ</source>. <year>1999</year> <month>Mar</month> <day>27</day>;<volume>318</volume>(<issue>7187</issue>):<fpage>843</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1136%2Fbmj.318.7187.843</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>28</label><mixed-citation publication-type="website"><collab>Ministry of Housing</collab>, <article-title>Communities &#x0026; Local Government. GOV.UK</article-title>. <year>2019</year> [cited <year>2024</year> <month>Jan</month> <day>3</day>]. <source>English indices of deprivation</source>. Available from: <uri>https://www.gov.uk/government/collections/english-indices-of-deprivation</uri></mixed-citation></ref>
<ref id="ref-29"><label>29</label><mixed-citation publication-type="website"><collab>Office for National Statistics</collab>. <article-title>Ethnic Groups by Borough - London Datastore [Internet]</article-title>. [cited <year>2024</year> <month>Jan</month> <day>3</day>]. Available from: <uri>https://data.london.gov.uk/dataset/ethnic-groups-borough</uri></mixed-citation></ref>
<ref id="ref-30"><label>30</label><mixed-citation publication-type="journal"><string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gilbert</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cromwell</surname> <given-names>D</given-names></string-name>, <string-name><surname>van der Meulen</surname> <given-names>J</given-names></string-name>. <article-title>Linking Data for Mothers and Babies in De-Identified Electronic Health Data</article-title>. <string-name><surname>Gebhardt</surname> <given-names>S</given-names></string-name>, editor. <source>PLoS ONE</source>. <year>2016</year> <month>Oct</month> <day>20</day>;<volume>11</volume>(<issue>10</issue>):<fpage>e0164667</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0164667</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>31</label><mixed-citation publication-type="journal"><string-name><surname>Damerau</surname> <given-names>FJ</given-names></string-name>. <article-title>A technique for computer detection and correction of spelling errors</article-title>. <source>Commun ACM</source>. <year>1964</year> <month>Mar</month> <day>1</day>;<volume>7</volume>(<issue>3</issue>):<fpage>171</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1145/363958.363994</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>32</label><mixed-citation publication-type="journal"><string-name><surname>Pollock</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Zamora</surname> <given-names>A</given-names></string-name>. <article-title>Automatic spelling correction in scientific and scholarly text</article-title>. <source>Commun ACM</source>. <year>1984</year> <month>Apr</month>;<volume>27</volume>(<issue>4</issue>):<fpage>358</fpage>&#x2013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1145/358027.358048</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>33</label><mixed-citation publication-type="book"><string-name><surname>Thomas</surname> <given-names>N</given-names></string-name>. <string-name><surname>Herzog</surname></string-name>, <string-name><surname>Fritz</surname> <given-names>J</given-names></string-name>. <string-name><surname>Scheuren</surname></string-name>, <string-name><surname>William</surname> <given-names>E</given-names></string-name>. <string-name><surname>Winkler</surname></string-name>. <chapter-title>Data Quality and Record Linkage Techniques [Internet]</chapter-title>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2007</year> [cited <year>2024</year> <month>Jan</month> <day>3</day>]. Available from: <uri>http://link.springer.com/10.1007/0-387-69505-2</uri>. <pub-id pub-id-type="doi">10.1007/0-387-69505-2</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>34</label><mixed-citation publication-type="journal"><string-name><surname>Kukich</surname> <given-names>K</given-names></string-name>. <article-title>Techniques for automatically correcting words in text</article-title>. <source>ACM Comput Surv</source>. <year>1992</year> <month>Dec</month>;<volume>24</volume>(<issue>4</issue>):<fpage>377</fpage>&#x2013;<lpage>439</lpage>. <pub-id pub-id-type="doi">10.1145/146370.146380</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>35</label><mixed-citation publication-type="website"><string-name><surname>Black</surname> <given-names>PE</given-names></string-name>. <article-title>Dictionary of Algorithms and Data Structures</article-title>. <source>NIST [Internet]</source>. <year>1998</year> <month>Oct</month> <day>1</day> [cited <year>2024</year> <month>Jan</month> <day>3</day>]; Available from: <uri>https://www.nist.gov/publications/dictionary-algorithms-and-data-structures</uri>.</mixed-citation></ref>
<ref id="ref-36"><label>36</label><mixed-citation publication-type="journal"><string-name><surname>Odell</surname> <given-names>M.K</given-names></string-name>. <article-title>The profit in records management</article-title>. <source>Systems (New York)</source>. <year>1956</year>; <volume>20</volume>(<issue>20</issue>).</mixed-citation></ref>
<ref id="ref-37"><label>37</label><mixed-citation publication-type="book"><string-name><surname>Holmes</surname> <given-names>D</given-names></string-name>, <string-name><surname>McCabe</surname> <given-names>MC</given-names></string-name>. <chapter-title>Improving precision and recall for Soundex retrieval</chapter-title>. In: <source>Proceedings International Conference on Information Technology: Coding and Computing [Internet]</source>. <publisher-loc>Las Vegas, NV, USA</publisher-loc>: <publisher-name>IEEE Comput. Soc</publisher-name>; <year>2002</year> [cited <year>2024</year> <month>Jan</month> <day>3</day>]. p. <fpage>22</fpage>&#x2013;<lpage>6</lpage>. Available from: <uri>http://ieeexplore.ieee.org/document/1000354/</uri>. <pub-id pub-id-type="doi">10.1109/ITCC.2002.1000354</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>38</label><mixed-citation publication-type="journal"><string-name><surname>Cheriet</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kharma</surname> <given-names>N</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>CL</given-names></string-name>, <string-name><surname>Suen</surname> <given-names>C</given-names></string-name>. <article-title>Character Recognition Systems: A Guide for Students and Practitioners</article-title>. <source>John Wiley &#x0026; Sons</source>; <year>2007</year>.</mixed-citation></ref>
<ref id="ref-39"><label>39</label><mixed-citation publication-type="journal"><string-name><surname>Gambaro</surname> <given-names>L</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>H</given-names></string-name>. <article-title>Moving home in the early years: what happens to children in the UK? Longitudinal and Life Course Studies</article-title>. <year>2016</year> <month>Jul</month> <day>18</day>;<volume>7</volume>(<issue>3</issue>):<fpage>265</fpage>&#x2013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.14301/llcs.v7i3.375</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>40</label><mixed-citation publication-type="journal"><string-name><surname>Ludvigsson</surname> <given-names>JF</given-names></string-name>, <string-name><surname>Otterblad-Olausson</surname> <given-names>P</given-names></string-name>, <string-name><surname>Pettersson</surname> <given-names>BU</given-names></string-name>, <string-name><surname>Ekbom</surname> <given-names>A</given-names></string-name>. <article-title>The Swedish personal identity number: possibilities and pitfalls in healthcare and medical research</article-title>. <source>Eur J Epidemiol</source>. <year>2009</year> <month>Nov</month> <day>1</day>;<volume>24</volume>(<issue>11</issue>):<fpage>659</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1007/s10654-009-9350-y</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>41</label><mixed-citation publication-type="journal"><string-name><surname>Aldridge</surname> <given-names>RW</given-names></string-name>, <string-name><surname>Shaji</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hayward</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Abubakar</surname> <given-names>I</given-names></string-name>. <article-title>Accuracy of Probabilistic Linkage Using the Enhanced Matching System for Public Health and Epidemiological Studies</article-title>. <source>PLOS ONE</source>. <year>2015</year> <month>Aug</month> <day>24</day>;<volume>10</volume>(<issue>8</issue>):<fpage>e0136179</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0136179</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>42</label><mixed-citation publication-type="journal"><string-name><surname>Fellegi</surname> <given-names>IP</given-names></string-name>, <string-name><surname>Sunter</surname> <given-names>AB</given-names></string-name>. <article-title>A Theory for Record Linkage</article-title>. <source>Journal of the American Statistical Association</source>. <year>1969</year>;<volume>64</volume>(<issue>328</issue>):<fpage>1183</fpage>&#x2013;<lpage>210</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>43</label><mixed-citation publication-type="journal"><collab>Roger Eckhardt</collab>. <article-title>Stan Ulam, John von Neumann and the Monte Carlo Method</article-title>. <source>Los Alamos Science</source>. <year>1987</year>;<volume>100</volume>(<issue>15</issue>):<fpage>131</fpage>.</mixed-citation></ref>
<ref id="ref-44"><label>44</label><mixed-citation publication-type="journal"><string-name><surname>Harron</surname> <given-names>K</given-names></string-name>, <string-name><surname>Doidge</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Goldstein</surname> <given-names>H</given-names></string-name>. <article-title>Assessing data linkage quality in cohort studies. Annals of Human Biology</article-title>. <year>2020</year> <month>Feb</month> <day>17</day>;<volume>47</volume>(<issue>2</issue>):<fpage>218</fpage>&#x2013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1080%2F03014460.2020.1742379</pub-id></mixed-citation></ref>
<ref id="ref-45"><label>45</label><mixed-citation publication-type="journal"><string-name><surname>Doidge</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Morris</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Harron</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Stevens</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gilbert</surname> <given-names>R</given-names></string-name>. <article-title>Prevalence of Down&#x2019;s Syndrome in England, 1998&#x2013;2013: Comparison of linked surveillance data and electronic health records</article-title>. <source>International Journal of Population Data Science [Internet]</source>. <year>2020</year> <month>Mar</month> <day>19</day> [cited <year>2023</year> <month>Jun</month> <day>11</day>];<volume>5</volume>(<issue>1</issue>). Available from: <uri>https://ijpds.org/article/view/1157</uri>. <pub-id pub-id-type="doi">10.23889/ijpds.v5i1.1157</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>46</label><mixed-citation publication-type="journal"><string-name><surname>Nanayakkara</surname> <given-names>C</given-names></string-name>, <string-name><surname>Christen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ranbaduge</surname> <given-names>T</given-names></string-name>. <article-title>An Anonymiser Tool for Sensitive Graph Data</article-title>. <source>InCIKM (workshops)</source> <year>2020</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>
