<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3747</article-id>
<article-id pub-id-type="publisher-id">11:5:3747</article-id>
<article-id pub-id-type="pii">S2399490821037472</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Tackling undercoverage in educational attainment registers using large-scale survey data: Statistics Netherlands’ Educational Attainment File</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Fang</surname><given-names initials="C">Christian</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>van Hedel</surname><given-names initials="K">Karen</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Hogendoorn</surname><given-names initials="B">Bram</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>van Rooijen</surname><given-names initials="J">Johan</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Statistics Netherlands, The Hague, Netherlands</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3747</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3747">This article is available from the IJPDS website at: https://ijpds.org/article/view/3747</self-uri>
<abstract>
<p>Educational attainment is a key variable in social-scientific research. In the Netherlands, most information on educational attainment is obtained from various registers of the Ministry of Education. These cover student enrollments in nearly all Dutch education institutions from the 1980s onwards. However, some residents have attained (additional) education before these registers came into existence, abroad, or at institutions not included in these registers (e.g. private institutions). This means that for 29% of the population information on education is entirely absent, whereas for some of the observed 71%, educational attainment is underestimated. Here, we present a unique method developed to deal with these deficiencies. By combining and integrating administrative data and data from the Dutch Labour Force Survey in a unique way, Statistics Netherlands can make valid inference to the population level on the basis of incomplete coverage. We compute weights for individuals not observed in any register but who were observed in the survey, with these survey observations effectively being treated as though they are a random sample from the unobserved population. Furthermore, we impute scores for individuals observed in only registers to correct for the undercoverage of education. Finally, a bootstrap approach is applied to assess uncertainty of statistical outcomes. Taken together, this method allows valid inference to the population level, generating point estimates as well as estimates of uncertainty.</p>
</abstract>
</article-meta>
</front>
</article>