<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3531</article-id>
<article-id pub-id-type="publisher-id">11:5:3531</article-id>
<article-id pub-id-type="pii">S239949082103531X</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Publishing Linked Data with Correction Weights</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Liu</surname><given-names initials="A">An-Chiao</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Lugtig</surname><given-names initials="P">Peter</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Scholtus</surname><given-names initials="S">Sander</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>de Waal</surname><given-names initials="T">Ton</given-names></name><xref ref-type="aff" rid="affil-3"><sup>3</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Utrecht University, Utrecht, Netherlands</institution></aff>
<aff id="affil-2"><label>2</label><institution>Statistics Netherlands, Den Haag, Netherlands</institution></aff>
<aff id="affil-3"><label>3</label><institution>Statistics Netherlands, Den Haag, Netherlands; Tilburg University, Tilburg, Netherlands</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3531</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3531">This article is available from the IJPDS website at: https://ijpds.org/article/view/3531</self-uri>
<abstract>
<p>Data linkage is often used to make inferences from distinct data sources. The data may be linked by a set of variables with high discrimination, for example, a social security number or some personal information. A good linking variable often implies a high disclosure risk in identifying the corresponding person. To prevent the disclosure risk, the data linker may choose to remove the linking variables before publishing the linked datasets. The design of the original data sources, the response pattern, and the unlinked part of the original data are often also unknown to the secondary user. However, the quality of the linked data is constrained by the original data sources and the linking process. Without considering the potential error or bias in linked datasets, naïvely treating them as error-free in secondary analysis may result in biased inference. To indicate the quality of the linked dataset without sacrificing privacy, we propose publishing correction weights alongside the linked datasets. The weights are generated given the information in the original data sources and the quality of the linkage. Both selection issues of the sample and measurement issues of the linking variables are addressed in the constructed weights, and we allow the possibility of having multiple potential links for a record. Secondary users may apply design-based estimators for subsequent analyses based on the correction weights, or apply sensitivity analysis given different sample inclusion criteria. An example is presented, and the option of secondary analysis given the constructed weights is discussed.</p>
</abstract>
</article-meta>
</front>
</article>