<?xml version="1.0"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.2 20190208//EN" "JATS-journalpublishing1.dtd"[]>
<article xml:lang="en" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" dtd-version="1.2" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">IJPDS</journal-id>
<journal-title-group>
<journal-title>International Journal of Population Data Science</journal-title>
<abbrev-journal-title>IJPDS</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2399-4908</issn>
<publisher>
<publisher-name>Swansea University</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.23889/ijpds.v6i1.1680</article-id>
<article-id pub-id-type="publisher-id">6:1:1680</article-id>
<article-id pub-id-type="pii">S2399490821016803</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Population Data Science</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Data harmonization and data pooling from cohort studies: a practical approach for data management</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Adhikari</surname><given-names initials="K">Kamala</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="corresp" rid="correspondingAurthor">*</xref></contrib>
<contrib contrib-type="author"><name><surname>Patten</surname><given-names initials="SB">Scott B</given-names></name><xref ref-type="aff" rid="affil-1">1</xref></contrib>
<contrib contrib-type="author"><name><surname>Patel</surname><given-names initials="AB">Alka B</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-2">2</xref></contrib>
<contrib contrib-type="author"><name><surname>Premji</surname><given-names initials="S">Shahirose</given-names></name><xref ref-type="aff" rid="affil-3">3</xref></contrib>
<contrib contrib-type="author"><name><surname>Tough</surname><given-names initials="S">Suzanne</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-4">4</xref></contrib>
<contrib contrib-type="author"><name><surname>Letourneau</surname><given-names initials="N">Nicole</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-4">4</xref><xref ref-type="aff" rid="affil-5">5</xref><xref ref-type="aff" rid="affil-6">6</xref></contrib>
<contrib contrib-type="author"><name><surname>Giesbrecht</surname><given-names initials="G">Gerald</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-4">4</xref></contrib>
<contrib contrib-type="author"><name><surname>Metcalfe</surname><given-names initials="A">Amy</given-names></name><xref ref-type="aff" rid="affil-1">1</xref><xref ref-type="aff" rid="affil-7">7</xref><xref ref-type="aff" rid="affil-8">8</xref></contrib>
<aff id="affil-1"><label>1</label><institution>Department of Community Health Sciences, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-2"><label>2</label><institution>Applied Research and Evaluation- Primary Health Care, Alberta Health Services, Calgary, Canada</institution></aff>
<aff id="affil-3"><label>3</label><institution>School of Nursing, Faculty of Health, York University, Calgary, Canada</institution></aff>
<aff id="affil-4"><label>4</label><institution>Department of Pediatrics, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-5"><label>5</label><institution>Faculty of Nursing University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-6"><label>6</label><institution>Deprtment of Psychiatry, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-7"><label>7</label><institution>Department of Obstetrics and Gynecology, University of Calgary, Calgary, Canada</institution></aff>
<aff id="affil-8"><label>8</label><institution>Department of Medicine, University of Calgary, Calgary, Canada</institution></aff>
</contrib-group>
<author-notes>
<corresp id="correspondingAurthor"><label>*</label>Corresponding author: Kamala Adhikari <email>kamala.adhikaridahal@ucalgary.ca</email>
</corresp>
<fn fn-type="conflict">
<label>Competing interests</label>
<p>The authors declare that they have no competing interests.</p>
</fn>
</author-notes>
<pub-date date-type="pub" publication-format="electronic"><day>30</day><month>11</month><year>2021</year></pub-date>
<pub-date date-type="collection" publication-format="electronic"><year>2021</year></pub-date>
<volume>6</volume>
<issue>1</issue>
<elocation-id>1680</elocation-id>
<permissions>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/">
<license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p>
</license>
</permissions>
<self-uri xlink:href="https://ijpds.org/article/view/1680">This article is available from the IJPDS website at: https://ijpds.org/article/view/1680</self-uri>
<abstract>
<title>Abstract</title>
<p>Data pooling from pre-existing datasets can be useful to increase study sample size and statistical power in order to answer a research question. However, individual datasets may contain variables that measure the same construct differently, posing challenges for data pooling. Variable harmonization, an approach that can generate comparable datasets from heterogeneous sources, can address this issue in some circumstances. As an illustrative example, this paper describes the data harmonization strategies that helped generate comparable datasets across two Canadian pregnancy cohort studies: All Our Families; and the Alberta Pregnancy Outcomes and Nutrition.</p>
<p>Variables were harmonized considering multiple features across the datasets: the construct measured; question asked/response options; the measurement scale used; the frequency of measurement; timing of measurement, and the data structure. Completely matching, partially matching, and completely un-matching variables across the datasets were determined based on these features. Variables that were an exact match were pooled as is. Partially matching variables were harmonized or processed under a common format across the datasets considering the frequency of measurement, the timing of measurement, the measurement scale used, and response options. Variables that were completely unmatching could not be harmonized into a single variable.</p>
<p>The variable harmonization strategies that were used to generate comparable cohort datasets for data pooling are applicable to other data sources. Future studies may employ or evaluate these strategies, which permit researchers to answer novel research questions in a statistically efficient, timely, and cost-efficient manner that could not be achieved using a single data source.</p>
</abstract>
<kwd-group>
<kwd>data harmonization</kwd>
<kwd>data pooling or combination</kwd>
<kwd>comparable dataset</kwd>
<kwd>cohort studies</kwd>
<kwd>harmonization strategies</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec>
<title>Introduction</title>
<p>Data pooling from multiple studies into a single dataset provides opportunities to increase the statistical power of a study and to answer novel research questions that could not be addressed using data from a single study [<xref ref-type="bibr" rid="ref-1">1</xref>, <xref ref-type="bibr" rid="ref-2">2</xref>]. Data pooling from existing data sources allows investigators to conduct research more rapidly and at a lower cost than primary data collection would allow, providing opportunities for timely translation of knowledge into practice.</p>
<p>Individual datasets from different studies or data sources often measure the same construct differently, which poses challenges for data pooling. These challenges are addressed by data harmonization. Data harmonization refers to efforts that provide comparability of datasets from heterogeneous sources and allows for combining, pooling, or integrating them in a coherent way [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Data harmonization can take a prospective or retrospective approach. Prospective data harmonization occurs at the initial stage of study design, or at least before data collection. For this, investigators agree on a common core set of variables or measures, compatible data collection tools, and standard operating procedures, often leading to a high degree of homogeneity [<xref ref-type="bibr" rid="ref-3">3</xref>, <xref ref-type="bibr" rid="ref-4">4</xref>]. Retrospective harmonization is a flexible approach, which targets the synthesis of already-collected information. For this, researchers define a core set of variables, and then assess the compatibility of information collected and the potential for creating single harmonized variables. If harmonization is possible, strategies for data processing are developed [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>Data harmonization is particularly valuable when the outcome and/or risk factor is rare, since examining interactions among risk factors and investigating population subgroups requires a large sample size to ensure adequate study power. It is not always feasible to accomplish this with primary data collection from a single study given the resources required. Additionally, measurement of the same construct using multiple measurement scales is generally unreasonable or unfeasible for a single study unless the primary aim of the study is to compare the results from the multiple scales.</p>
<p>The use of multiple existing datasets from the studies that were conducted in similar target populations using comparable methodologies but different measurement scales can address these issues. However, data harmonization (in this case, retrospective) involves extensive data processing or data cleaning and management and variable transformation processes. While these processes are critical [<xref ref-type="bibr" rid="ref-6">6</xref>], the literature or guidelines on how to do this remain limited [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>Our research project aimed to improve the understanding of risk factors for preterm birth using data from two pregnancy cohort studies conducted in Alberta, Canada&#x2013;All Our Families (AOF: n = 3,351) and Alberta Pregnancy Outcomes and Nutrition (APrON: n = 2,187) [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Specifically, our research intended to develop and validate a prediction model for preterm birth, to evaluate the suitability of and comparability of multiple anxiety scales to measure anxiety during pregnancy, and to examine if neighborhood socioeconomic status modified the association between anxiety and/or depression status during pregnancy and preterm birth [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>]. Achieving these goals required data harmonization.</p>
<p>This paper describes the data harmonization strategies that helped generate comparable datasets across these two studies to address our research objectives. It presents examples of data harmonization strategies that were used to generate comparable datasets. These strategies may be employed or evaluated in subsequent studies, and may serve as useful starting points for other projects.</p>
</sec>
<sec>
<title>Methods</title>
<sec>
<title>Data sources</title>
<p>We obtained two de-identified datasets from the two prospective pregnancy cohort studies (AOF: n = 3,351 and APrON: n = 2,187). Both datasets are available for secondary analysis and are housed in SAGE (Secondary Analysis to Generate Evidence), a secure data repository developed by PolicyWise for Children &#x0026; Families, which houses these datasets (<uri>https://policywise.com</uri>).</p>
<p>The AOF and APrON studies are ongoing cohort studies of mother and child dyads. Both cohort studies use quality control procedures to maintain the quality of study-specific data. To illustrate, both studies use data management standards for data storage, data entry, data dictionary and data cleaning. The data are double-entered by trained research assistants with discrepancies resolved by a master coder. All implausible or unusual values are re-entered to verify the data. In some cases, participants are contacted for clarification, and in other cases, the studies collected additional information that allowed them to correct implausible values. Where such corrections are not possible, the data are set to missing.</p>
<p>Each dataset was linked (by SAGE) with neighbourhood socioeconomic status measured by both the average household income and the Pampalon material deprivation index. Both measures were derived from 2011 Statistics Canada census data [<xref ref-type="bibr" rid="ref-14">14</xref>&#x2013;<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>The AOF and APrON studies are comparable in many ways including target population, recruitment time periods, inclusion criteria, sampling design, data collection methods, cohort characteristics (such as age, income, and parity), and participant follow-ups and retention during the perinatal period (<xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 1</xref>) [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Both studies collect data about mothers, children, and partners, using methods including questionnaires, health records, and lab samples.</p>
<p>Given the similarity between the study populations and methodologies, pooling data from these studies was justifiable [<xref ref-type="bibr" rid="ref-1">1</xref>]. However, each study measured/recorded the same construct/variables differently and therefore, data harmonization strategies were used to generate a comparable dataset across the studies.</p>
<p>Data harmonization focused only on the maternal data obtained from questionnaires. Both studies collected data using questionnaires on perinatal health, including maternal demographics, socioeconomic status, lifestyle, social support, depression, anxiety, and preterm delivery [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Details on the description and comparability of these cohort studies is available elsewhere [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>], and are summarized in <xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 1</xref>.</p>
</sec>
<sec>
<title>Variable harmonization</title>
<p>Study documentation from the AOF and APrON studies (such as study protocols and standard operating procedures, questionnaires and instrument calibration procedures, data dictionaries, and published papers) were accessed and reviewed. Conversations between our research team and the AOF/APrON research teams enabled an understanding of the level of substantive heterogeneity (i.e., study methodologies and equivalence of variables to be harmonized) and data management systems across studies [<xref ref-type="bibr" rid="ref-6">6</xref>]. Agreement on data access and intellectual property from each study and ethics approval from the Conjoint Health Research Ethics Board at the University of Calgary were obtained before data harmonization. We also performed preliminary exploration of each dataset before initiating the actual harmonization to further understand the constructs, questions, responses, variables available in the datasets, data distributions and value labels, or the data quality and comparability [<xref ref-type="bibr" rid="ref-6">6</xref>]. These strategies facilitated the identification and selection of variables to consider for harmonization and helped decide harmonization strategies to be employed.</p>
<p>Variables pertinent to address our research objectives were selected to consider for harmonization (<xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 2</xref>). These variables were harmonized in each dataset considering multiple features of the data, as recommended by previous authors [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>, <xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-17">17</xref>]. These features included whether the variables were completely or partially identical regarding: (a) the construct measured; (b) question asked and response options; (c) the measurement scale used; (d) the frequency of measurement; (e) the timing of the measurement (i.e., when in pregnancy the variable was measured); and (f) the coding features of variables. The coding features of variables considered for data harmonization included: variable name, definition, type, format, and response categories; variable value label; and missing values, including response categories &#x201C;not applicable&#x201D;, &#x201C;not stated&#x201D;, and &#x201C;don&#x2019;t know&#x201D;.</p>
<p>Multiple features of data were checked through the review of the documentations of the primary studies, the conversations with primary study research teams, and preliminary exploration of variables in the datasets. If the variables were found to have an exact match for each of these features, they were considered completely matching. If the variables were the same in terms of what construct was measured, but were different in terms of frequency of measurement, the timing of measurement, and variable response options and coding features, these variables were considered partially matching. These partially matching variables were harmonized or processed under a common format and, if needed, to the same frequency and timing of measurements across the datasets. Finally, some important variables did not match, and required a different approach (<xref ref-type="table" rid="table-1">Table 1</xref> and <xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 2</xref>).</p>
<table-wrap id="table-1">
<label>Table 1: Variable harmonization</label> 
<table frame="hsides" rules="groups">
<col width="15%"/>
<col width="20%"/>
<col width="20%"/>
<col width="20%"/>
<col width="25%"/>
<tbody>
<tr>
<th valign="top" style="border-top: solid 1pt; border-bottom: solid 1pt" align="left"><bold>Variables</bold></th>
<th valign="top" style="border-top: solid 1pt; border-bottom: solid 1pt" align="left"><bold>AOF cohort dataset</bold></th>
<th valign="top" style="border-top: solid 1pt; border-bottom: solid 1pt" align="left"><bold>APrON cohort dataset</bold></th>
<th valign="top" style="border-top: solid 1pt; border-bottom: solid 1pt" align="left"><bold>Harmonization process</bold></th>
<th valign="top" style="border-top: solid 1pt; border-bottom: solid 1pt" align="left"><bold>Variables combined (and recoded if needed)</bold></th>
</tr>
<tr>
<td valign="top" align="left">Maternal age</td>
<td valign="top" align="left">Variable name: Q1MMAGE2 Construct: Maternal age at recruitment <break/>Type of data: continuous <break/>Missing:. (period)</td>
<td valign="top" align="left">Variable name: MAQ <break/>Construct: Maternal age at recruitment <break/>Type of data: continuous <break/>Missing: 999</td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching of construct</p></list-item>
<list-item><p>Complete matching of response or data type and coding, except missing value coding (partial matching)</p></list-item>
</list> <break/>Action taken: Coded missing data on APrON as. (period) and both variables renamed with same name</td>
<td valign="top" align="left">Maternal age variables with continuous data combined and recoded as
<list list-type="bullet">
<list-item><p>&#x003C;35 years</p></list-item>
<list-item><p>&#x2265;35 years</p></list-item>
<list-item><p>. Missing</p></list-item></list></td>
</tr>
<tr>
<td valign="top" align="left">Marital status</td>
<td valign="top" align="left">Variable name: Q1MMSTAT1 <break/>Construct: Current marital status <break/>Data type: Categorical <break/>Response category and value level:
<list list-type="bullet">
<list-item><p>1 Single</p></list-item>
<list-item><p>2 Single with partner</p></list-item>
<list-item><p>3 Married</p></list-item>
<list-item><p>4 Common-law</p></list-item>
<list-item><p>5 Divorced</p></list-item>
<list-item><p>6 Separated</p></list-item>
<list-item><p>. Missing</p></list-item></list></td>
<td valign="top" align="left">Variable name: MAGB1 <break/>Construct: Current marital status <break/>Data type: Categorical <break/>Response category and value level:
<list list-type="bullet">
<list-item><p>0 Single</p></list-item>
<list-item><p>1 Married</p></list-item>
<list-item><p>2 Divorced</p></list-item>
<list-item><p>3 Common-law</p></list-item>
<list-item><p>4 Widowed</p></list-item>
<list-item><p>5 Separated</p></list-item>
<list-item><p>999 Missing</p></list-item></list></td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching construct</p></list-item>
<list-item><p>Partial matching of variable response and coding</p></list-item></list> <break/>Action taken: Recoding <break/>AOF: Combined single and single with partner response into &#x201C;single&#x201D; and combined divorced, widowed and separated response into divorced/ separated/widowed, <break/>APrON: Combined divorced, widowed and separated response into divorced/ separated/widowed. <break/>Variable in both datasets were recoded as:
<list list-type="bullet">
<list-item><p>0 Single</p></list-item>
<list-item><p>1 Married/common-law</p></list-item>
<list-item><p>2 Divorced/separated/widowed</p></list-item>
<list-item><p>. Missing</p></list-item></list> <break/>Variable renamed with same name</td>
<td valign="top" align="left">Variables with the following categories combined
<list list-type="bullet">
<list-item><p>0 Single</p></list-item>
<list-item><p>1 Married/common-law</p></list-item>
<list-item><p>2 Divorced/separated/widowed</p></list-item>
<list-item><p>. Missing</p></list-item></list></td>
</tr>
<tr>
<td valign="top" align="left">Maternal ethnicity</td>
<td valign="top" align="left">Variable name: Q1METH1_2 Construct: Ethnic origin <break/>Data type: Categorical <break/>Response category and value level:
<list list-type="bullet">
<list-item><p>0 Others</p></list-item>
<list-item><p>1 White/Caucasian</p></list-item></list></td>
<td valign="top" align="left">Variable name: MAGB16 <break/>Construct: Ethnic origin <break/>Data type: Categorical <break/>Response category and value level:
<list list-type="bullet">
<list-item><p>1 Caucasian</p></list-item>
<list-item><p>2 Chinese</p></list-item>
<list-item><p>3 Filipino</p></list-item>
<list-item><p>4 Japanese</p></list-item>
<list-item><p>5 Korean</p></list-item>
<list-item><p>6 Latin American</p></list-item>
<list-item><p>7 Aboriginal/Native</p></list-item>
<list-item><p>8 South Asian</p></list-item>
<list-item><p>9 South East Asian</p></list-item>
<list-item><p>10 Arab</p></list-item>
<list-item><p>11 West Asian</p></list-item>
<list-item><p>12 Black</p></list-item>
<list-item><p>13 Others</p></list-item></list></td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching construct</p></list-item>
<list-item><p>Partial matching of variable response and coding</p></list-item></list> <break/>Action taken: Recoding <break/>APrON: Combined coding 2-13 into &#x201C;others&#x201D; and recoded as
<list list-type="bullet">
<list-item><p>0 Others</p></list-item>
<list-item><p>1 White/Caucasian</p></list-item>
<list-item><p>Variable renamed with same name</p></list-item></list></td>
<td valign="top" align="left">Variables with the following categories combined
<list list-type="bullet">
<list-item><p>0 Others</p></list-item>
<list-item><p>1 White/Caucasian</p></list-item></list></td>
</tr>
<tr>
<td valign="top" align="left">Body mass index</td>
<td valign="top" align="left">Variable name: Q1MHW8 <break/>Construct: Pre-pregnancy weight in kg <break/>Variable name: Q1MHW5 <break/>Construct: Height in cm <break/>Data type: continuous <break/>Missing:.</td>
<td valign="top" align="left">Variable name: MAANTH2 and MBANTH2 <break/>Construct: Pre-pregnancy weight in kg <break/>Variable name: MAANTH3 and MBANTH <break/>Construct: Pre-pregnancy height in cm <break/>Data type: continuous <break/>Missing: 999</td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching construct</p></list-item>
<list-item><p>Partial matching of data coding or management system</p></list-item></list> <break/>Action taken: Variable managed and body mass index calculated <break/>AOF: Calculated body mass index <break/>APrON:
<list list-type="bullet">
<list-item><p>Combined 2 weight variables into one</p></list-item>
<list-item><p>Combined 2 height variables into one</p></list-item>
<list-item><p>Recoded missing (999) into (.)</p></list-item>
<list-item><p>Calculated body mass index</p></list-item></list></td>
<td valign="top" align="left">Combined continuous body mass index variable and recoded as 4 categories
<list list-type="bullet">
<list-item><p>0 Underweight &#x003C;18.5</p></list-item>
<list-item><p>1 Normal weight 18.5 &#x2013; 24.9</p></list-item>
<list-item><p>2 Overweight 25 &#x2013; 29.9</p></list-item>
<list-item><p>3 Obese 30+</p></list-item></list></td>
</tr>
<tr>
<td valign="top" align="left">Parity</td>
<td valign="top" align="left">Variable name: Q1MPPI1_1 <break/>Construct: Parity (birth to a fetus &#x003E;24 weeks) <break/>Data type: Categorical <break/>Response category and value level:
<list list-type="bullet">
<list-item><p>0 No previous births</p></list-item>
<list-item><p>1 Previous birth to a fetus (at least once)</p></list-item>
<list-item><p>. Missing</p></list-item></list> <break/>If previous birth to a fetus, number of live births
<list list-type="bullet">
<list-item><p>1 to 7</p></list-item>
<list-item><p>Missing (.)</p></list-item></list></td>
<td valign="top" align="left">Variable name: MAPI3 <break/>Construct: Live born children have you had <break/>Data type: Categorical <break/>Response category and value level:
<list list-type="bullet">
<list-item><p>0 to 4</p></list-item>
<list-item><p>missing (999)</p></list-item></list></td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching construct</p></list-item>
<list-item><p>Partial matching variable response and coding</p></list-item></list> <break/>Action taken: Recoding <break/>In both datasets, responses were recoded as
<list list-type="bullet">
<list-item><p>1 Primiparous</p></list-item>
<list-item><p>2 Multiparous</p></list-item>
<list-item><p>3 Grand multiparous (&#x003E;2 live births)</p></list-item>
<list-item><p>. &#x201C;missing&#x201D;</p></list-item></list></td>
<td valign="top" align="left">Variables with the following categories combined
<list list-type="bullet">
<list-item><p>1 Primiparous</p></list-item>
<list-item><p>2 Multiparous</p></list-item>
<list-item><p>3 Grand multiparous</p></list-item>
<list-item><p>. Missing</p></list-item></list>
</td>
</tr>
<tr>
<td valign="top" align="left">Depression during pregnancy</td>
<td valign="top" align="left">Variable name: Q1MEDPS <break/>Construct: EPDS score in first measurement (during recruitment: &#x003C;24 weeks of gestation) <break/>Variable name: Q2MEDPS Construct: EPDS score in second measurement (in third trimester: 34-38 weeks gestation)</td>
<td valign="top" align="left">Variable name: MAEPDS_Score <break/>Construct: EPDS score in first measurement (during recruitment: &#x003C;27 weeks of gestation) <break/>Variable name: MBEPDS_Score <break/>Construct: EPDS score in second measurement (in 14-26 weeks of gestation for those participants who were 0-13 weeks of gestation during the recruitment) <break/>Variable name: MCEPDS_Score <break/>Construct: EPDS score in third measurement (in 27-40 weeks of gestation for those who were 0-26 weeks of gestation during recruitment)</td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching construct</p></list-item>
<list-item><p>Partial matching in terms of number of measurements and measurement time during pregnancy (week of gestation)</p></list-item></list><break/>Action taken: In both datasets, using the recorded week of gestation at first, second and third measurements, 3 variables of EPDS score for each trimester were created.<break/>
<list list-type="bullet">
<list-item><p>EPDS score in first trimester</p></list-item>
<list-item><p>EPDS score in second trimester</p></list-item>
<list-item><p>EPDS score third trimester</p></list-item></list></td>
<td valign="top" align="left">Three combined variables for depression during pregnancy
<list list-type="bullet">
<list-item><p>EPDS score in first trimester</p></list-item>
<list-item><p>EPDS score in second trimester</p></list-item>
<list-item><p>EPDS score third trimester</p></list-item></list></td>
</tr>
<tr>
<td valign="top" align="left">Anxiety during pregnancy</td>
<td valign="top" align="left">Variable name: Q1MSSAI <break/>Construct: anxiety score in first measurement (during recruitment: &#x003C;24 weeks of gestation), measured by STAI-20 <break/>Variable name: Q2MSSAI Construct: anxiety score in second measurement (in third trimester: 34-38 weeks gestation), measured by STAI-20</td>
<td valign="top" align="left">Variable name: MASCL_Score <break/>Construct: anxiety score in first measurement, measured by SCL-90 (during recruitment: &#x003C;27 weeks of gestation) <break/>Variable name: MBSCL_Score <break/>Construct: anxiety score in second measurement, measured by SCL-90 (in second trimester:14-26 weeks of gestation for those participants who were 0-13 weeks of gestation during the recruitment) <break/>Variable name: MCSCL_Score <break/>Construct: anxiety score in third measurement, measured by SCL-90 (in third trimester: 27-40 weeks for those who were 0-26 weeks of gestation during recruitment)</td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Completely un-matching variable Action taken:</p></list-item>
<list-item><p>Harmonized anxiety score measured by each scale for each trimester using the same process for depression during pregnancy. Accordingly, three separate variables for anxiety during pregnancy by trimester (as for depression) for each anxiety scale were created.</p></list-item>
<list-item><p>Overlapped participants and their anxiety data measured by both scales identified.</p></list-item></list></td>
<td valign="top" align="left">Anxiety data measured by two different scales were pooled as two different variables
<list list-type="bullet">
<list-item><p>For 231 participants who participated both studies, each variable contained anxiety data.</p></list-item>
<list-item><p>For independent participants, each variable contained missing values if they did not have anxiety data measured by the same scale.</p></list-item></list>
</td>
</tr>
<tr>
<td valign="top" align="left">Anxiety during pregnancy, measured by EPDS-3A</td>
<td valign="top" align="left">Variable name: Q1MEDPS <break/>Construct: EPDS score (comprising EPDS-3A anxiety score) in first measurement (during recruitment: &#x003C;24 weeks of gestation) <break/>Variable name: Q2MEDPS Construct: EPDS score (comprising EPDS-3A anxiety score) in second measurement (in third trimester:34-38 weeks gestation)</td>
<td valign="top" align="left">Variable name: MAEPDS_Score <break/>Construct: EPDS score (comprising EPDS-3A anxiety score) in first measurement (during recruitment: &#x003C;27 weeks of gestation) <break/>Variable name: MBEPDS_Score <break/>Construct: EPDS score (comprising EPDS-3A anxiety score) in second measurement (in second trimester: 14-26 weeks of gestation for those participants who were 0-13 weeks of gestation during the recruitment) <break/>Variable name: MCEPDS_Score <break/>Construct: EPDS score (comprising EPDS-3A anxiety score) in third measurement (in third trimester: 27-40 weeks for those who were 0-26 weeks of gestation during recruitment)</td>
<td valign="top" align="left">
<list list-type="bullet">
<list-item><p>Complete matching construct</p></list-item>
<list-item><p>Partial matching in terms of number of measurements and measurement time during pregnancy (week of gestation)</p></list-item></list> <break/>Action taken: In both datasets, we created the compatible anxiety variables, by extracting the data on three items of the EPDS (i.e., anxiety items 3, 4, and 5) measured by both studies. The three items comprise the anxiety subscale (EDPS-3A) <break/>In both datasets, using the recorded week of gestation at first, second and third measurements, 3 variables of EPDS-3A score for each trimester were created.<break/>
<list list-type="bullet">
<list-item><p>EPDS-3A score in first trimester</p></list-item>
<list-item><p>EPDS-3A score in second trimester</p></list-item>
<list-item><p>EPDS- 3A score third trimester</p></list-item></list></td>
<td valign="top" align="left">Three combined variables for anxiety during pregnancy
<list list-type="bullet">
<list-item><p>EPDS-3A score in first trimester</p></list-item>
<list-item><p>EPDS-3A score in second trimester</p></list-item>
<list-item><p>EPDS-3A score third trimester</p></list-item></list></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Note: AOF: All Our Families; APrON: Alberta Pregnancy Outcomes and Nutrition; EPDS: Edinburgh Postnatal Depression Scale; STAI-20: State-Trait Anxiety Inventory-State 20-item scale; SCL-90: Symptoms Checklist-90; EPDS-3A: Edinburgh Postnatal Depression scale- anxiety subscale.</p>
</table-wrap-foot>
</table-wrap>
<p>If the construct was not measured in one of the datasets or if different measurement scales that emphasize the different components were used to measure the same construct across the datasets, the variables were deemed completely un-matching (<xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 2</xref>). In particular, the AOF dataset had data on anxiety during pregnancy measured by the State-Trait Anxiety Inventory-State 20-item scale (STAI-20), and the APrON dataset had anxiety data during pregnancy measured by the anxiety subscale of the Symptoms Checklist-90 (SCL-90). The variables comprising the anxiety data measured by these two different scales were important for our research that intended to compare the performance of multiple anxiety scales in measuring anxiety during pregnancy. Hence, we created anxiety data measured by two different scales as two different variables. We identified that there were participants who participated in both cohort studies (n <bold>=</bold> 231) and their anxiety data measured by both scales (<xref ref-type="table" rid="table-1">Table 1</xref>).</p>
<p>Anxiety data with a large sample size was critical for our research that aimed to examine effect modification between anxiety and/or depression status during pregnancy and neighborhood socioeconomic status on the risk of preterm birth. Since harmonization of direct measures of anxiety into a single variable was not feasible, we created comparable anxiety variables across studies by extracting data on three items of the Edinburgh Postnatal Depression Scale (EPDS) [<xref ref-type="bibr" rid="ref-18">18</xref>], which was used in both studies (<xref ref-type="table" rid="table-1">Table 1</xref>). Specifically, items 3, 4 and 5 of the EPDS comprise an anxiety subscale (EPDS-3A), which has been suggested by previous studies as a measure of anxiety in the obstetric population [<xref ref-type="bibr" rid="ref-19">19</xref>, <xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p>Documentation was created for variables across two datasets in terms of a variable name (a unique identity of the variable, e.g., smoking), variable definition (a short description of the variable, e.g., smoking status before pregnancy), variable value label (a short description of the response attributed to the underlying numerical values, e.g., &#x201C;no&#x201D; for 0, &#x201C;yes&#x201D; for 1), variable type (continuous or discrete), variable format (numeric or character), and missing value coding (&#x201C;.&#x201D; or &#x201C;999&#x201D;). Once the selected variables in each dataset were harmonized and documented, the datasets were organized such that the same number of appending variables appeared in the same order for both datasets. Hence, the datasets were vertically identical by appending variables. Then, the two harmonized cohort datasets were concatenated into a single dataset (n = 5,538).</p>
<p>We used quality control procedures to test and describe the quality of harmonized data. Cross-tabulation or five-number summary (as appropriate to the data type) of each harmonized variable was done in each dataset to evaluate the consistency of those variables and the distribution of participants across the datasets (<xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 3</xref>). Variable formatting and descriptive statistics or distribution of participants were also assessed on the harmonized, combined datasets to explore any discrepancies with the variables on study-specific datasets.</p>
<p>Data harmonization procedures and the descriptive statistics of study-specific and combined data were documented as described above, and discussed with our research team and the broader AOF and APrON study teams. The discussion with the teams provided a qualitative validation of the data harmonization strategies used, a key step to make sure that the data harmonization process maintained the integrity of the original data and the original data were not lost. The discussion also facilitated to fix the errors (related to original variable coding or data entry) that were observed during the data harmonization process.</p>
<p>The final, harmonized data set was then used to answer our research objectives. Analytic approaches included regression analyses, structural equation modeling, and prediction model development and evaluation [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>]. We imputed missing values, for the study variables that were not measured (thus contained missing values) in one cohort/dataset and also for those that were measured in both dataset with = &#x2265;5% missing data, from the predictive distribution based on the observed data.</p>
</sec>
</sec>
<sec>
<title>Results</title>
<p>A total of 20 variables were considered for harmonization, and of those, 18 variables (90.0%) were successfully harmonized. Of 20 variables, three variables (15.0%) were completely matching and 14 (70.0%) were partially matching. These variables were successfully harmonized across the datasets and pooled/combined (i.e., appended into a single variable). One variable (5.0%) was completely unmatching across the datasets due to the different measurement scales used to measure the same construct, this variable was harmonized across the datasets for the purpose of data merging (i.e., pooling data as two different variables). Two variables (10.0%) were only available in one dataset; thus, variable harmonization was not applicable (<xref ref-type="supplementary-material" rid="sup-a">Supplement Table 2</xref>). Characteristics (or distribution) of participants across the studies were similar in harmonized data, except drug abuse and smoking status (<xref ref-type="supplementary-material" rid="sup-a">Supplementary Table 3</xref>). There were discrepancies in the proportion of missing data for some variables, particularly body mass index and gestational age at delivery. These discrepancies also existed in the original datasets; thus, they were not related to the data harmonization process.</p>
<p>Several partially matching variables such as marital status, ethnicity, income, parity, depression, and smoking were successfully harmonized (<xref ref-type="supplementary-material" rid="sup-a">Supplement Table 2</xref>). For example, one variable, current marital status, was partially identical across the datasets as the construct measured (or question asked) was completely identical across both datasets but the variable response categories and the value level coding were different across the datasets. As the variable response categories were collapsible to identical and meaningful categories across the datasets, the variable response was re-organized into three identical categories in both datasets. Another variable, depression symptoms during pregnancy - which was measured in both datasets using the same scale, the EPDS - was not compatible in terms of frequency of measurement and gestational age at each measurement. Accordingly, the depression variables were harmonized by creating three unique variables in each dataset that indicated the depression score in each trimester of pregnancy.</p>
<p>Similarly, the EPDS-3A-based anxiety variables, which were made by extracting data on three items of the EPDS, were harmonized by creating three unique variables in each dataset that indicated the anxiety score in each trimester of pregnancy. The anxiety variables measured by two different anxiety scales (i.e., STAI-20 and SCL-90) were harmonized by creating three unique variables in each dataset that indicated the anxiety score (measured by different scales across the datasets) in each trimester of pregnancy (<xref ref-type="table" rid="table-1">Table 1</xref>).</p>
<p>The harmonized combined cohort dataset (n = 5,538) contained several important variables, including maternal age, gestational age at delivery, marital status, ethnicity, duration of stay in Canada, body mass index, parity, smoking, and anxiety (measured by EPDS-3A), and depression during pregnancy for each trimester. Additionally, variables that were important for our research but were only available in one of the datasets (previous preterm birth and prenatal care visits) or measured by different anxiety measurement scales (anxiety during pregnancy) were included in the combined dataset.</p>
<p>The anxiety data measured by two different scales across the datasets were pooled as two different variables, with missing values recorded for measures on the scale not included in the original study. Anxiety data or values were available for both anxiety-related variables for participants who participated in both cohort studies (overlapping study cohort, n = 231) (<xref ref-type="table" rid="table-1">Table 1</xref>). Similarly, the combined dataset contained missing values for the cohort with no measurement of previous preterm birth and prenatal care visits variables.</p>
</sec>
<sec>
<title>Discussion</title>
<p>This study describes data harmonization strategies, which helped create comparable datasets across two cohort studies and enabled the datasets pooling. The combined dataset created unique research opportunities to answering our clinically relevant research questions, by providing a large sample size (thus increased study power and efficiency), additional variables, and data measured by multiple different scales [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>]. The use of the harmonized, combined dataset facilitated statistical analysis to answer our research questions and added comprehensiveness to our research, which would have been less feasible using either of the datasets alone.</p>
<p>For example, the large sample size provided an opportunity to analyze the risk of preterm birth (relatively a rare outcome) across the several strata of risk factors, such as anxiety alone, depression alone, and both anxiety and depression and their stratification across socioeconomic variables [<xref ref-type="bibr" rid="ref-13">13</xref>]. Similarly, we evaluated the performance of multiple anxiety scales in measuring anxiety during pregnancy: the suitability of STAI- 20 and SCL-90 anxiety screening scales in the individual study cohort and the comparability of these scales (correlation between the anxiety scores measured by two scales) restricting our analysis in the overlapping study cohort [<xref ref-type="bibr" rid="ref-12">12</xref>]. We performed analyses including those variables that were available in both datasets [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>]. We also performed sensitivity analyses using the additional variables available in one dataset [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>].</p>
<p>The harmonized data are stored in a secure data repository (SAGE - Secondary Analysis to Generate Evidence) which also houses the cohort-specific datasets. The dataset may be available upon request from the AOF and APrON data custodians. The harmonization strategies described are applicable to generate comparable data across administrative databases, survey cycles, jurisdictions (provincial, national or international), and measures repeated over time. However, the strategies may not be necessarily directly applicable to different contexts, such as harmonizing data from a larger number of studies or data sources. Heterogeneity across studies or datasets becomes more persistent and data harmonization process becomes complex as the number of datasets or data sources increases. Nevertheless, it may be worthwhile to evaluate the utility and applicability of these strategies in subsequent studies.</p>
<p>The success or the scientific impact of any data harmonization and integration research project depends on the quality of the data harmonization process, the quality of the information collected by the primary studies, and the ability to access the data collected [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>, <xref ref-type="bibr" rid="ref-17">17</xref>]. Hence, a series of procedures should be considered as a part of the data harmonization and synthesis initiatives to ensure the quality and validity of the harmonized databases created. To illustrate, the potential to harmonize and integrate information depends on homogeneity across a range of study-specific factors. These include the study design, target population, time period, and duration of follow-up; the type of information and samples collected; the specific tools and standard operating procedures used to collect or generate data; and the data coding and data management systems employed. The incompatibility of these study-specific factors can affect whether variables recorded in different data sources are actually measuring the same construct. Access to documentation from the primary studies, dialogue with their research teams, and preliminary exploration of the dataset before the actual harmonization allow researchers to understand the level of substantive heterogeneity across studies [<xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-6">6</xref>]. These strategies ultimately facilitate the selection of variables to be harmonized or combined and helps decide harmonization strategies to be employed.</p>
<p>Additionally, agreement on data access and intellectual property from each study and ethical approval must be obtained before data harmonization. Finally, it is important that the person(s) involved in data harmonization always create new files for the harmonization and document the harmonization process [<xref ref-type="bibr" rid="ref-17">17</xref>]. This facilitates the evaluation of the data harmonization process and reproducibility.</p>
<p>While the need for additional statistical power has often led investigators to employ data harmonization and data pooling, there are several other benefits as well [<xref ref-type="bibr" rid="ref-2">2</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>]. These include increased use of existing data, strengthening the scientific impact of individual studies, and optimal return on investments. To illustrate, compared to building new studies involving thousands of participants, employing data harmonization on existing data can permit the generation of research projects relatively rapidly and at a lower cost, with timely knowledge translation opportunities. This also allows researchers to properly explore similarities and differences across time and place. Ultimately, data harmonization initiatives leverage national and international collaborations, facilitates the emergence of leading-edge collaborative and cross-disciplinary research initiatives and innovations, and thereby minimizes the duplication of research efforts [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
<p>While data harmonization is an important component in research, its application (harmonization process and harmonized data) has some challenges and limitations. Recent publications provide high-level guideline on harmonization [<xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-6">6</xref>], but literature on how to perform data processing and evaluate harmonization quality (practical approaches) is limited [<xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-6">6</xref>]. The data harmonization process is resource intensive. It involves a repetitive/iterative and time-consuming process, requires thorough preparatory work, and has many elements that must be worked through carefully and systematically with rigorous documentation.</p>
<p>To illustrate, data processing and integration in a systematic manner requires a comprehensive understanding of previous studies (study-specific designs, standard operating procedures, data collection devices, data format and data content, and quality of study-specific data) and requires research content knowledge and analytical skills [<xref ref-type="bibr" rid="ref-5">5</xref>, <xref ref-type="bibr" rid="ref-6">6</xref>]. Even if harmonization procedures (variable selection and pairing rules definition and data processing) are done under the consensus and advice from experts, there is inevitably an element of subjectivity in harmonization procedures. Evaluation of the quality of the harmonized data is required to understand its scientific performance [<xref ref-type="bibr" rid="ref-6">6</xref>]. At least two independent individuals are needed to evaluate inter-coder agreement with regard to their data harmonization procedures or processes (such as Cohen&#x2019;s k statistic) [<xref ref-type="bibr" rid="ref-5">5</xref>]. Furthermore, data harmonization may lead to limited use of information (in terms of the number of variables, variable categories) collected by specific primary studies.</p>
<p>For example, in our research context, maternal ethnicity and household income variables were categorized differently across datasets, broad categories vs. specific categories. Using harmonized data, we had to analyze the data by broad categories. We also had to analyze the anxiety data on the subsample. In contrast, the use of a single study or dataset is less resource-intensive, with more flexibility on using the information collected by primary studies, but has other limitations as described. Additionally, the harmonization strategies used in one context may not necessarily be directly applicable to different contexts due to the variation in heterogeneities, such as large number of studies or data sources and heterogeneous target population and data collections and management systems across studies.</p>
</sec>
<sec>
<title>Conclusion</title>
<p>Data harmonization is an important aspect of conducting research using multiple datasets. It generates comparable data across different data sources and facilitates pooling of relevant data across data sources, leading to unique opportunities for research. Data harmonization and pooling augment the utility and scientific impact of existing data or individual studies, creates a collaborative research environment, minimize the duplication of research, and increase research feasibility. Hence, data harmonization is a very promising avenue to support advancement in population health research that can result in improvements to the health and well-being of populations.</p>
</sec>
<sec>
<title>Ethical statements</title>
<p>Ethics approval for this study was obtained from the Conjoint Health Research Ethics Board at the University of Calgary (REB16-2548). This study used secondary data and all the data were anonymized; therefore, did not require informed consent.</p>
<boxed-text content-type="box">
<caption><title>Box 1: Data harmonization best practices or key lessons</title></caption>
<list list-type="order">
<list-item><p>Appraisal of published and unpublished documents of the primary studies and conversations with the primary study teams, to gain in-depth understanding regarding the primary study methodologies and facilitate the judgement around the homogeneity of study- or data sources-specific factors.</p></list-item>
<list-item><p>Agreement on data access and intellectual property from each study and ethical approval before data harmonization.</p></list-item>
<list-item><p>Preliminary exploration of variables in the datasets before initiating the actual harmonization to understand the variables available in the datasets, the variable coding or data management, the data distributions, or the data quality and comparability.</p></list-item>
<list-item><p>Identification of completely unidentical and completely or partially identical variables across data sources, and variable harmonization, considering multiple features such as construct measured, measurement scale used, cross-sectional or longitudinal measurement, data coding, and overlapping samples.</p></list-item>
<list-item><p>Establishing the consistency of variables across two datasets before data combination.</p></list-item>
<list-item><p>Preserve the integrity of the original data and ensure that the original data are not lost, while seeking to harmonize variables to address own research purposes and exploring unique research opportunities such as overlapping samples and data measured by multiple scales.</p></list-item>
<list-item><p>Documentation of data harmonization procedures and sharing/discussing it with the primary study teams, seek suggestions on data harmonization procedures used and solutions for data errors observed in the original datasets during the data harmonization process.</p></list-item>
</list>
</boxed-text>
</sec>
<sec sec-type="supplementary-material">
<title>Supplementary Files</title>
<supplementary-material id="sup-a">
<label>Supplementary Tables</label> 
<media mimetype="application" mime-subtype="pdf" xlink:href="ijpds-06-1680-s001.pdf"/>
</supplementary-material>
</sec>
</body>
<back>
<ref-list>
<ref id="ref-1"><label>1</label><mixed-citation publication-type="journal"><string-name><surname>Roberts</surname> <given-names>G</given-names></string-name>, <string-name><surname>Binder</surname> <given-names>D</given-names></string-name>. <article-title>Analyses Based on Combining Similar Information from Multiple Surveys</article-title>. <source>Section on Survey Research Methods Joint Statistical Meetings (JSM)</source>; <year>2009</year>. p.<fpage>2138</fpage>&#x2013;<lpage>47</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>2</label><mixed-citation publication-type="journal"><string-name><surname>Rao</surname> <given-names>SR</given-names></string-name>, <string-name><surname>Graubard</surname> <given-names>BI</given-names></string-name>, <string-name><surname>Schmid</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Morton</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Louis</surname> <given-names>TA</given-names></string-name>, <string-name><surname>Zaslavsky</surname> <given-names>AM</given-names></string-name>, <etal>et al</etal>. <article-title>Meta-analysis of survey data: application to health services research</article-title>. <source>Health Services and Outcomes Research Methodology</source>. <year>2008</year>;<volume>8</volume>(<issue>2</issue>):<fpage>98</fpage>&#x2013;<lpage>114</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>3</label><mixed-citation publication-type="journal"><string-name><surname>Fortier</surname> <given-names>I</given-names></string-name>, <string-name><surname>Doiron</surname> <given-names>D</given-names></string-name>, <string-name><surname>Burton</surname> <given-names>P</given-names></string-name>, <string-name><surname>Raina</surname> <given-names>P</given-names></string-name>. <article-title>Invited commentary: consolidating data harmonization&#x2013;how to obtain quality and applicability?</article-title> <source>Am J Epidemiol</source>. <year>2011</year>;<volume>174</volume>(<issue>3</issue>):<fpage>261</fpage>&#x2013;<lpage>4</lpage>; author reply 5-6.</mixed-citation></ref>
<ref id="ref-4"><label>4</label><mixed-citation publication-type="journal"><string-name><surname>Fortier</surname> <given-names>I</given-names></string-name>, <string-name><surname>Doiron</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wolfson</surname> <given-names>C</given-names></string-name>, <string-name><surname>Raina</surname> <given-names>P</given-names></string-name>. <article-title>Harmonizing data for collaborative research on aging: Why should we foster such an agenda?</article-title> <source>Canadian Journal of Aging</source>. <year>2012</year>;<volume>31</volume>:<fpage>95</fpage>&#x2013;<lpage>99</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>5</label><mixed-citation publication-type="journal"><string-name><surname>Fortier</surname> <given-names>I</given-names></string-name>, <string-name><surname>Doiron</surname> <given-names>D</given-names></string-name>, <string-name><surname>Little</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ferretti</surname> <given-names>V</given-names></string-name>, <string-name><surname>L&#x2019;Heureux</surname> <given-names>F</given-names></string-name>, <string-name><surname>Stolk</surname> <given-names>RP</given-names></string-name>, <etal>et al</etal>. <article-title>Is rigorous retrospective harmonization possible? Application of the DataSHaPER approach across 53 large studies.</article-title> <source>Int J Epidemiol</source>. <year>2011</year>; <volume>40</volume>:<fpage>1314</fpage>&#x2013;<lpage>1328</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>6</label><mixed-citation publication-type="journal"><string-name><surname>Fortier</surname> <given-names>I</given-names></string-name>, <string-name><surname>Raina</surname> <given-names>P</given-names></string-name>, Heuvel ERVd, <string-name><surname>Griffith</surname> <given-names>LE</given-names></string-name>, <string-name><surname>Craig</surname> <given-names>C</given-names></string-name>, <string-name><surname>Saliba</surname> <given-names>M</given-names></string-name>, <etal>et al</etal>. <article-title>Maelstrom research guidelines for rigorous retrospective data harmonization</article-title>. <source>Int J Epidemiol</source>. <year>2017</year>;<volume>46</volume> (<issue>1</issue>):<fpage>103</fpage>&#x2013;<lpage>105</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>7</label><mixed-citation publication-type="journal"><string-name><surname>Kaplan</surname> <given-names>BJ</given-names></string-name>, <string-name><surname>Giesbrecht</surname> <given-names>GF</given-names></string-name>, <string-name><surname>Leung</surname> <given-names>BM</given-names></string-name>, <string-name><surname>Field</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Dewey</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bell</surname> <given-names>RC</given-names></string-name>, <etal>et al</etal>. <article-title>The Alberta Pregnancy Outcomes and Nutrition (APrON) cohort study: rationale and methods</article-title>. <source>Matern Child Nutr</source>. <year>2014</year>;<volume>10</volume>(<issue>1</issue>):<fpage>44</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>8</label><mixed-citation publication-type="journal"><string-name><surname>Leung</surname> <given-names>BM</given-names></string-name>, <string-name><surname>McDonald</surname> <given-names>SW</given-names></string-name>, <string-name><surname>Kaplan</surname> <given-names>BJ</given-names></string-name>, <string-name><surname>Giesbrecht</surname> <given-names>GF</given-names></string-name>, <string-name><surname>Tough</surname> <given-names>SC</given-names></string-name>. <article-title>Comparison of sample characteristics in two pregnancy cohorts: community-based versus population-based recruitment methods</article-title>. <source>BMC Med Res Methodol</source>. <year>2013</year>;<volume>13</volume>:<fpage>149</fpage>.</mixed-citation></ref>
<ref id="ref-9"><label>9</label><mixed-citation publication-type="journal"><string-name><surname>McDonald</surname> <given-names>SW</given-names></string-name>, <string-name><surname>Lyon</surname> <given-names>AW</given-names></string-name>, <string-name><surname>Benzies</surname> <given-names>KM</given-names></string-name>, <string-name><surname>McNeil</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Lye</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Dolan</surname> <given-names>SM</given-names></string-name>, <etal>et al</etal>. <article-title>The All Our Babies pregnancy cohort: design, methods, and participant characteristics</article-title>. <source>BMC Pregnancy Childbirth</source>. <year>2013</year>;<volume>13</volume> Suppl <supplement>1</supplement>:<fpage>S2</fpage>.</mixed-citation></ref>
<ref id="ref-10"><label>10</label><mixed-citation publication-type="journal"><string-name><surname>Tough</surname> <given-names>SC</given-names></string-name>, <string-name><surname>McDonald</surname> <given-names>SW</given-names></string-name>, <string-name><surname>Collisson</surname> <given-names>BA</given-names></string-name>, <string-name><surname>Graham</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Kehler</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kingston</surname> <given-names>D</given-names></string-name>, <etal>et al</etal>. <article-title>Cohort Profile: The All Our Babies pregnancy cohort (AOB)</article-title>. <source>Int J Epidemiol</source>. <year>2017</year>;<volume>46</volume>(<issue>5</issue>):<fpage>1389</fpage>&#x2013;<lpage>90k</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>11</label><mixed-citation publication-type="journal"><string-name><surname>Adhikari</surname> <given-names>K</given-names></string-name>, <string-name><surname>Patten</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Williamson</surname> <given-names>T</given-names></string-name>, <string-name><surname>Patel</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Premji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tough</surname> <given-names>S</given-names></string-name>, <string-name><surname>Letourneau</surname> <given-names>N</given-names></string-name>, <string-name><surname>Giesbrecht</surname> <given-names>G</given-names></string-name>, <string-name><surname>Metcalfe</surname> <given-names>A</given-names></string-name>. <article-title>Does Neighbourhood Socioeconomic Status Predict the Risk of Preterm Birth? A Community-based Canadian Cohort Study</article-title>. <source>BMJ Open</source>. <year>2019</year>;<volume>9</volume>:<fpage>e025341</fpage>. <pub-id pub-id-type="doi">10.1136/bmjopen-2018-025341</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12</label><mixed-citation publication-type="journal"><string-name><surname>Adhikari</surname> <given-names>K</given-names></string-name>, <string-name><surname>Patten</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Williamson</surname> <given-names>T</given-names></string-name>, <string-name><surname>Patel</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Premji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tough</surname> <given-names>S</given-names></string-name>, <string-name><surname>Letourneau</surname> <given-names>N</given-names></string-name>, <string-name><surname>Giesbrecht</surname> <given-names>G</given-names></string-name>, <string-name><surname>Metcalfe</surname> <given-names>A</given-names></string-name>. <article-title>Assessment of Anxiety during Pregnancy Using Multiple Anxiety Scales: Do Anxiety Scales Differ in Their Ability to Assess Anxiety During Pregnancy</article-title>? <source>Journal of Psychosomatic Obstetrics &#x0026; Gynecology</source>. <year>2020</year>:<fpage>1</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>13</label><mixed-citation publication-type="journal"><string-name><surname>Adhikari</surname> <given-names>K</given-names></string-name>, <string-name><surname>Patten</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Williamson</surname> <given-names>T</given-names></string-name>, <string-name><surname>Patel</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Premji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tough</surname> <given-names>S</given-names></string-name>, <string-name><surname>Letourneau</surname> <given-names>N</given-names></string-name>, <string-name><surname>Giesbrecht</surname> <given-names>G</given-names></string-name>, <string-name><surname>Metcalfe</surname> <given-names>A</given-names></string-name>. <article-title>Neighbourhood Socioeconomic Status Modifies the Association between Anxiety and Depression during Pregnancy and Preterm Birth: A Community-based Canadian Cohort Study</article-title>. <source>BMJ Open</source>. <year>2020</year>;<volume>10</volume>::<fpage>e031035</fpage>. <pub-id pub-id-type="doi">10.1136/ bmjopen-2019-031035.13.</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14</label><mixed-citation publication-type="journal"><string-name><surname>Pampalon</surname> <given-names>R</given-names></string-name>, <string-name><surname>Raymond</surname> <given-names>G</given-names></string-name>. <article-title>A deprivation index for health and welfare planning in Quebec</article-title>. <source>Chronic Dis Can</source> <year>2000</year>;<volume>21</volume>:<fpage>104</fpage>&#x2013;<lpage>13</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>15</label><mixed-citation publication-type="journal"><collab>Alberta Health Services</collab>. <article-title>How to use the Pampalon Deprivation Index in Alberta: Research and Innovation, Alberta Health Services</article-title>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-16"><label>16</label><mixed-citation publication-type="other"><collab>Statistics Canada</collab>. <article-title>2011 Census Program</article-title>. Retrieved on <month>February</month> <day>01</day>, <year>2021</year>, from: 2011 Census Program: Topics (statcan.gc.ca)</mixed-citation></ref>
<ref id="ref-17"><label>17</label><mixed-citation publication-type="website"><string-name><surname>Kveder</surname> <given-names>A</given-names></string-name>, <string-name><surname>Galico</surname> <given-names>A</given-names></string-name>. <article-title>Guideline for cleaning and harmonization of generations and gender survey data</article-title>. Retrieved on <month>February</month> <day>01</day>, <year>2021</year>, from: <uri>http://www.unece.org/pau/ggp/materials.htm</uri>.</mixed-citation></ref>
<ref id="ref-18"><label>18</label><mixed-citation publication-type="journal"><string-name><surname>Cox</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Holden</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Sagovsky</surname> <given-names>R</given-names></string-name>. <collab>Detection of postnataldepression</collab>. <article-title>Development of the 10-item EdinburghPostnatal Depression Scale</article-title>. <source>Br J Psychiatry</source>. <year>1987</year>;<volume>150</volume>:<fpage>782</fpage>&#x2013;<lpage>786</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>19</label><mixed-citation publication-type="journal"><string-name><surname>Matthey</surname> <given-names>S</given-names></string-name>. <article-title>Using the Edinburgh postnatal depression scale to screen for anxiety disorders</article-title>. <source>Depress Anxiety</source> <year>2008</year>;<volume>25</volume>:<fpage>926</fpage>&#x2013;<lpage>31</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>20</label><mixed-citation publication-type="journal"><string-name><surname>Matthey</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fisher</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rowe</surname> <given-names>H</given-names></string-name>. <article-title>Using the Edinburgh postnatal depression scale to screen for anxiety disorders: conceptual and methodological considerations</article-title>. <source>J Affect Disord</source> <year>2013</year>;<volume>146</volume>:<fpage>224</fpage>&#x2013;<lpage>30</lpage></mixed-citation></ref>
<ref id="ref-21"><label>21</label><mixed-citation publication-type="journal"><string-name><surname>Bergeron</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rachel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Stephanie</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alan</surname> <given-names>B</given-names></string-name>, <string-name><surname>William</surname> <given-names>F</given-names></string-name>, <string-name><surname>Isabel</surname> <given-names>F</given-names></string-name>. <article-title>Cohort Profile: Research advancement through cohort cataloguing and harmonization (ReACH)</article-title>. <source>Int J Epidemiol</source>. <year>2021</year>;<volume>50</volume>(<issue>2</issue>):<fpage>396</fpage>&#x2013;<lpage>397</lpage>.</mixed-citation></ref>
</ref-list>
</back>
</article>