<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.2811</article-id>
      <article-id pub-id-type="publisher-id">9:5:321</article-id>
      <title-group>
        <article-title>Investigating variation in reported location of death: A comparison of administrative data sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Elliott</surname>
            <given-names initials="T">Tom</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
          <xref ref-type="aff" rid="affil-2">2</xref>
          <xref ref-type="aff" rid="affil-3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Milne</surname>
            <given-names initials="B">Barry</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Li</surname>
            <given-names initials="E">Eileen</given-names>
          </name>
          <xref ref-type="aff" rid="affil-2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Sporle</surname>
            <given-names initials="A">Andrew</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Simpson</surname>
            <given-names initials="C">Colin</given-names>
          </name>
          <xref ref-type="aff" rid="affil-3">3</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>iNZight Analytics Ltd</institution></aff>
      <aff id="affil-2"><label>2</label><institution>University of Auckland</institution></aff>
      <aff id="affil-3"><label>3</label><institution>Victoria University of Wellington</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>2811</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2811">This article is available from the IJPDS website at: https://ijpds.org/article/view/2811</self-uri>
    </article-meta>
  </front>
  <body>
    <p>How can we lower barriers to reuse of linked administrative data and improve quality at the same time? Linked data resources are expensive and complex, but accompanying readily-accessible metadata supports data reuse and quality control. However, metadata systems that are unsystematic or non-automated can lead to inconsistencies, making metadata use difficult even for experienced analysts.</p>
    <p>Our national statistics office (NSO) maintains a research database (RD) of administrative datasets linkable at the individual level. With the NSO’s agreement, we extracted schema information (e.g. tables and variables) from the RD and obtained data dictionaries. R was used to extract information (e.g., descriptions, human-friendly names, codings) from the data dictionaries, which identified errors and inconsistencies that the NSO then fixed. A metadata database was collated using schema and data dictionary information about variables and datasets in the RD. Finally, a public web app was developed to enable exploration of the meta database by searching specific terms or navigating the hierarchical relationships.</p>
    <p>We now work with the NSO data team to re-extract variables and add or update dictionaries ahead of the regular data updates. Metadata coverage is displayed on the app and used by our NSO to improve metadata quality and coverage. Our scripting workflow demonstrates the utility of automation in picking up errors and inconsistencies quickly, and before they can propagate by copy and paste. The web app, available to existing and new users, is now listed by the NSO as a “go-to” resource for using the linked administrative data resources.</p>
  </body>
</article>