<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v11i5.3678</article-id>
<article-id pub-id-type="publisher-id">11:5:3678</article-id>
<article-id pub-id-type="pii">S2399490821036788</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Clean Data, Breakthrough Insights: The Importance of Standardizing Your Research Warehouse</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Youngson</surname><given-names initials="E">Erik</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Armani</surname><given-names initials="N">Nathan</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Whitten</surname><given-names initials="T">Tara</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Quan</surname><given-names initials="H">Hude</given-names></name><xref ref-type="aff" rid="affil-2"><sup>2</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Li</surname><given-names initials="B">Bing</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Wang</surname><given-names initials="T">Ting</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<contrib contrib-type="author"><name><surname>Bakal</surname><given-names initials="J">Jeffrey</given-names></name><xref ref-type="aff" rid="affil-1"><sup>1</sup></xref></contrib>
<aff id="affil-1"><label>1</label><institution>Alberta Provincial Research Data Services, Edmonton, Canada</institution></aff>
<aff id="affil-2"><label>2</label><institution>University of Calgary, Calgary, Canada</institution></aff>
</contrib-group>
<pub-date date-type="pub" publication-format="electronic"><day></day><month></month><year></year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year></year></pub-date>
<volume>11</volume>
<issue>5</issue>
<elocation-id>3678</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/3678">This article is available from the IJPDS website at: https://ijpds.org/article/view/3678</self-uri>
<abstract>
<p>Leveraging routinely collected health services data to conduct scientifically robust research requires a clear understanding and standardization of data from disparate clinical information systems which often change over time. Within Alberta Provincial Research Data Services (PRDS) a large team of analysts utilize a vast Enterprise Data Warehouse containing dozens of datasets with various update cycles and availability, generally spanning more than 20 years of data for a Canadian province of over 4 million people. We examine two recently developed standardized data marts for mortality and medical lab tests. These data marts are now updated daily using Snowflake Tasks to ensure timely data can be accessed as needed to support research studies. The mortality data are linked from three sources to ensure comprehensive coverage: Vital Statistics, Provincial Registry, and Connect Care (the provincial Epic based electronic medical record) and resolves inconsistencies in personal health numbers and discrepancies in dates of death. The Lab data mart uses test level data from several different historical lab systems and the current Connect Care system and consolidates coding and definitions over time to create a clean, easy to use data mart. Both use cases create a single, validated and linkable source of truth that is easy to use for research studies, negating the need for each analyst to re-engineer from several different source tables each time. The process used here will be applied to other subject matters internally and could be similarly applied data warehouse environments in other jurisdictions.</p>
</abstract>
</article-meta>
</front>
</article>