<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2"  article-type="abstract">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJPDS</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Population Data Science</journal-title>
        <abbrev-journal-title>IJPDS</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">2399-4908</issn>
      <publisher>
        <publisher-name>Swansea University</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.23889/ijpds.v7i3.1921</article-id>
      <article-id pub-id-type="publisher-id">7:03:147</article-id>
      <title-group>
        <article-title>Ethical considerations in the use of Machine Learning for research and statistics.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Toms</surname>
            <given-names initials="A">Alice</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Whitworth</surname>
            <given-names initials="S">Simon</given-names>
          </name>
          <xref ref-type="aff" rid="affil-1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="affil-1"><label>1</label>
        <institution>Office for National Statistics</institution>
      </aff>
      <pub-date date-type="pub" publication-format="electronic"><day></day><month>09</month><year>2022</year></pub-date>
      <pub-date date-type="collection" publication-format="electronic"><year>2022</year></pub-date>
      <volume>7</volume>
      <issue>3</issue>
      <elocation-id>1921</elocation-id>
      <permissions>
        <license license-type="open-access" xlink:href="https://creativecommons.org/licences/by/4.0/">
          <license-p>This work is licenced under a Creative Commons Attribution 4.0 International License.</license-p>
        </license>
      </permissions>
      <self-uri xlink:href="https://ijpds.org/article/view/1921">This article is available from the IJPDS website at: https://ijpds.org/article/view/1921</self-uri>
    </article-meta>
  </front>
  <body>
    <p>This paper, based upon new guidance created in collaboration with researchers from several national statistical institutes, explores the main ethical considerations associated with the use of machine learning techniques for aggregate statistics. The aim of this paper is to provide applied, practical ethical guidance for researchers using machine learning techniques.</p>
    <p>Following an extensive literature review, alongside discussion and collaboration with a number of national statistical institutes, it was identified that there was a need for applied guidance on the use of machine learning for the production of official statistics by the international research and statistical community. Feedback was gathered from interested stakeholders, which found that whilst there were resources available to researchers relating to the ethical considerations of machine learning projects, these focus mainly on operational uses of machine learning, and furthermore, lacked advice on how to practically mitigate ethical issues that arise throughout the project lifecycle.</p>
    <p>The guidance focuses on four main ethical considerations, found to be prevalent within machine learning research, and offers ways to mitigate these issues should they arise. These are: the importance of minimising and mitigating social bias and discrimination within machine learning research, and clearly communicating these, and the limitations of our research; the need to consider the transparency and explainability of machine learning research, and the implications this has for reproducibility; the importance of maintaining accountability throughout machine learning processes, ensuring that models are used only for their intended purposes, and that different stakeholders are aware of their responsibilities; the need to consider the confidentiality and privacy risks arising from the data used, both in relation to training data which is fed into the machine, and outputs resulting from the machine learning’s findings.</p>
    <p>The guidance has been well-received following its release, and feedback from the wider user community to date has been positive. The IPDLN conference provides an opportunity for further feedback to be collated to ensure that the guidance continues to be valuable to its intended audience.</p>
  </body>
</article>