<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3694</article-id>
<article-id pub-id-type="publisher-id">11:5:3694</article-id>
<article-id pub-id-type="pii">S2399490821036946</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Data Quality Framework for Integrated Data via Statistical Matching</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Moretti</surname><given-names initials="A">Angelo</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Salvatore</surname><given-names initials="C">Camilla</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Kraemer</surname><given-names initials="F">Fabienne</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Utrecht University, Utrecht, Netherlands</institution></aff>
<aff id="affil-2"><label>2</label><institution>GESIS - Leibniz Institute for the Social Sciences, Mannheim, Germany</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3694</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3694">This article is available from the IJPDS website at: https://ijpds.org/article/view/3694</self-uri>
<abstract>
<p>Researchers often combine data from multiple sources with the goal of bringing together different types of information and creating a richer data set. With the combined data researchers can cross-validate information or explore relationships that cannot be examined using only one data source. To ensure the validity of empirical conclusions, the integrated data needs to meet certain data quality standards. However, each data source may carry specific representation or measurement errors into the integrated dataset, and the data integration process itself can introduce new biases or amplify existing ones. While quality measures have been proposed to assess data from multiple sources (De Waal et al., 2019), guidance on potential data quality issues across the stages of the integration process remains limited. We present a comprehensive data quality framework to guide researchers through integration via statistical matching. The framework outlines data quality measures and best practices at each stage of the process: pre-matching, matching, and post-matching. To illustrate its practical application, we apply the framework to a real-world case study integrating the 2021 German General Social Survey (ALLBUS) (probability-based survey) with data from the German Longitudinal Election Study in 2021 (nonprobability survey). The integrated data are used to examine how various social attitudes relate to voting behavior and political engagement. Our work contributes to understanding data quality issues in integrated data while providing practical implementation guidance for researchers and practitioners.</p>
</abstract>
</article-meta>
</front>
</article>