<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3486</article-id>
<article-id pub-id-type="publisher-id">11:5:3486</article-id>
<article-id pub-id-type="pii">S2399490821034868</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Are births predictable with linked survey and register data? Evidence from the Predicting Fertility data challenge</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Sivak</surname><given-names initials="E">Elizaveta</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Stulp</surname><given-names initials="G">Gert</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>University of Groningen, Groningen, Netherlands</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3486</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3486">This article is available from the IJPDS website at: https://ijpds.org/article/view/3486</self-uri>
<abstract>
<p>Accurate predictions of life outcomes can inform social theory and policy. Yet, social science predictions often perform poorly. One common explanation is small sample sizes of social datasets. We examine the predictability of having a child within three years under conditions that are close to the best currently possible. We use full-population Dutch registers and LISS survey data, linkable to the registers, in the Predicting Fertility data challenge. LISS data includes a wider range of theoretically relevant variables, including subjective measures such as fertility intentions. Register data provides vast, high-resolution coverage of life-course trajectories but lacks subjective measures. This unique framework enables us to leverage the strengths of these datasets to assess the current predictability. Over 150 people participated in the data challenge and submitted over 70 models, both traditional machine learning and cutting-edge approaches. Predictive performance remained modest for both datasets: even the best models fell below the upper limit of predictability caused by randomness in conception and fetal survival. Survey-based predictions performed slightly better. Combining the best survey- and register-based models – by using the predicted probabilities from the best register-based model as an additional feature in the best survey-based model – slightly improved predictions. These results suggest that accurately predicting individual fertility remains highly uncertain, even in the short term and with extensive data. Modest predictive accuracy is unlikely to stem primarily from limited sample size, but may reflect the inherent unpredictability of life outcomes.</p>
</abstract>
</article-meta>
</front>
</article>