<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd" [
]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML"
  dtd-version="1.2" article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v9i5.2891</article-id>
      <article-id pub-id-type="publisher-id">9:5:399</article-id>
      <title-group>
        <article-title>Progress and developments towards making historic administrative data research ready</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Williamson</surname>
            <given-names initials="L">Lee</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Dibben</surname>
            <given-names initials="C">Chris</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label><institution>University of Edinburgh</institution></aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>18</day>
        <month>09</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection" publication-format="electronic">
        <year>2024</year>
      </pub-date>
      <volume>9</volume>
      <issue>5</issue>
      <elocation-id>2891</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/2891">This article is available from the IJPDS website at: https://ijpds.org/article/view/2891</self-uri>
    </article-meta>
  </front>
  <body>
    <sec>
      <title>Objectives</title>
      <p>To make research ready data from transcribed historic civil registration records: birth, marriage and death (from mid-19th century onwards). To use historic records effectively for large-scale research they must not only be made machine-readable, but also coded in a suitable format – along with classification of transcribed information. Included on the digitised historic records are textual descriptions of occupations and causes of death. Thus, to code transcribed occupations to HISCO and causes of death to ICD.</p>
    </sec>
    <sec>
      <title>Approach</title>
      <p>It is impractical to hand-code the records manually (34 million occupations and 8 million causes of death), especially for deaths where more than one cause can be given. As such, coding is viewed as a text classification task and the process is automated. To facilitate auto-coding, a proportion of records were hand-coded as training data (90,000 occupations and 102,000 deaths). Ahead of the auto-coding, initial pre-processing, cleaning and standardising is done on both the occupations and deaths.</p>
    </sec>
    <sec>
      <title>Results</title>
      <p>Preliminary experiments undertaken obtained reasonable results from a combination of exact matching and statistical classification. Experiments using a larger pilot uncovered that since some occupations are very common, the training data set covers a very large proportion of the records (ie exact match). This proportion is not as high for deaths given the different ways causes are written.</p>
    </sec>
    <sec>
      <title>Conclusion</title>
      <p>This is work in progress to create the research ready data, and the poster will include the results from experiments using the full training data (90,000 occupations and 102,000 deaths).</p>
    </sec>
  </body>
</article>