Keyword search

Sort Help
entries

Results

All (88)

All (88) (0 to 10 of 88 results)

  • Journals and periodicals: 12-206-X
    Description: This report summarizes the annual achievements of the Methodology Research and Development Program (MRDP) sponsored by the Modern Statistical Methods and Data Science Branch at Statistics Canada. This program covers research and development activities in statistical methods with potentially broad application in the agency’s statistical programs; these activities would otherwise be less likely to be carried out during the provision of regular methodology services to those programs. The MRDP also includes activities that provide support in the application of past successful developments in order to promote the use of the results of research and development work. Selected prospective research activities are also presented.
    Release date: 2025-10-10

  • Articles and reports: 11-522-X202500100034
    Description: Until now, detailed data on the destination of manufacturing sales have not historically been available to Canadians. Through integration of annual survey data, a destination of sales table by industry and province of origin was developed for the annual and monthly manufacturing surveys at Statistics Canada. Respondents for the annual survey are asked for their distribution of sales as a percentage across 15 destinations. To tackle the difficulty of generating an establishment-level distribution for multi-province respondents, three approaches were compared: using the respondents' total distribution for all their establishments, using optimization, and using the distributions of the single-province respondents. The imputed distribution of destination sales from the annual data was then applied to the monthly survey's sales value. This paper delves into challenges faced with imputing the destination sales (especially for respondents with establishments in multiple provinces), ensuring sales match marginal origin province totals, and allocating a distribution of destinations based on data from the annual program to the monthly estimates.
    Release date: 2025-09-08

  • Articles and reports: 11-522-X202500100035
    Description: Historically, the Canadian census of population Edit and Imputation (E&I) process has operated using a nearest-neighbour donor imputation methodology wherein the distance between a failed unit and a potential donor is obtained through a weighted combination of auxiliary variables. Revision to the model between cycles can be a complicated and time-consuming process given there is no standard approach to variable selection and weighting between topics. This paper will illustrate the potential of the Relief variable selection algorithm to create a machine learning-driven approach to variable selection and weighting that is standardized and comparable between census cycles and among the many topics of the census. An overview on how this process may be applied in practice will be presented, followed by results on several topics that indicate a general improvement over previous methods.
    Release date: 2025-09-08

  • Articles and reports: 12-001-X202500100004
    Description: Survey data collection often is plagued by unit and item nonresponse. To reduce reliance on strong assumptions about the missingness mechanisms, statisticians can use information about population marginal distributions known, for example, from censuses or administrative databases. One approach that does so is the Missing Data with Auxiliary Margins, or MD-AM, framework, which uses multiple imputation for both unit and item nonresponse so that survey-weighted estimates accord with the known marginal distributions. However, this framework relies on specifying and estimating a joint distribution for the survey data and nonresponse indicators, which can be computationally and practically daunting in data with many variables of mixed types. We propose two adaptations to the MD-AM framework to simplify the imputation task. First, rather than specifying a joint model for unit respondents’ data, we use random hot deck imputation while still leveraging the known marginal distributions. Second, instead of sampling from conditional distributions implied by the joint model for the missing data due to item nonresponse, we apply multiple imputation by chained equations for item nonresponse before imputation for unit nonresponse. Using simulation studies with nonignorable missingness mechanisms, we demonstrate that the proposed approach can provide more accurate point and interval estimates than models that do not leverage the auxiliary information. We illustrate the approach using data on voter turnout from the U.S. Current Population Survey.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100013
    Description: This discussion of the paper by Rao and Lohr focuses on the use of machine learning procedures for estimating finite population parameters. While there is growing interest in these methods within national statistical offices, several areas remain largely unexplored and warrant significant attention in the coming years. In this discussion, I highlight potential topics for future research and development in this rapidly evolving field.
    Release date: 2025-06-30

  • Journals and periodicals: 62F0026M
    Description: This series provides detailed documentation on the issues, concepts, methodology, data quality and other relevant research related to household expenditures from the Survey of Household Spending, the Homeowner Repair and Renovation Survey and the Food Expenditure Survey.
    Release date: 2025-05-21

  • Articles and reports: 12-001-X202200100008
    Description:

    The Multiple Imputation of Latent Classes (MILC) method combines multiple imputation and latent class analysis to correct for misclassification in combined datasets. Furthermore, MILC generates a multiply imputed dataset which can be used to estimate different statistics in a straightforward manner, ensuring that uncertainty due to misclassification is incorporated when estimating the total variance. In this paper, it is investigated how the MILC method can be adjusted to be applied for census purposes. More specifically, it is investigated how the MILC method deals with a finite and complete population register, how the MILC method can simultaneously correct misclassification in multiple latent variables and how multiple edit restrictions can be incorporated. A simulation study shows that the MILC method is in general able to reproduce cell frequencies in both low- and high-dimensional tables with low amounts of bias. In addition, variance can also be estimated appropriately, although variance is overestimated when cell frequencies are small.

    Release date: 2022-06-21

  • Articles and reports: 12-001-X202000100006
    Description:

    In surveys, logical boundaries among variables or among waves of surveys make imputation of missing values complicated. We propose a new regression-based multiple imputation method to deal with survey nonresponses with two-sided logical boundaries. This imputation method automatically satisfies the boundary conditions without an additional acceptance/rejection procedure and utilizes the boundary information to derive an imputed value and to determine the suitability of the imputed value. Simulation results show that our new imputation method outperforms the existing imputation methods for both mean and quantile estimations regardless of missing rates, error distributions, and missing-mechanisms. We apply our method to impute the self-reported variable “years of smoking” in successive health screenings of Koreans.

    Release date: 2020-06-30

  • Surveys and statistical programs – Documentation: 12-539-X
    Description:

    This document brings together guidelines and checklists on many issues that need to be considered in the pursuit of quality objectives in the execution of statistical activities. Its focus is on how to assure quality through effective and appropriate design or redesign of a statistical project or program from inception through to data evaluation, dissemination and documentation. These guidelines draw on the collective knowledge and experience of many Statistics Canada employees. It is expected that Quality Guidelines will be useful to staff engaged in the planning and design of surveys and other statistical projects, as well as to those who evaluate and analyze the outputs of these projects.

    Release date: 2019-12-04

  • Articles and reports: 12-001-X201800254957
    Description:

    When a linear imputation method is used to correct non-response based on certain assumptions, total variance can be assigned to non-responding units. Linear imputation is not as limited as it seems, given that the most common methods – ratio, donor, mean and auxiliary value imputation – are all linear imputation methods. We will discuss the inference framework and the unit-level decomposition of variance due to non-response. Simulation results will also be presented. This decomposition can be used to prioritize non-response follow-up or manual corrections, or simply to guide data analysis.

    Release date: 2018-12-20
Data (2)

Data (2) ((2 results))

  • Public use microdata: 82M0010X
    Description:

    The National Population Health Survey (NPHS) program is designed to collect information related to the health of the Canadian population. The first cycle of data collection began in 1994. The institutional component includes long-term residents (expected to stay longer than six months) in health care facilities with four or more beds in Canada with the principal exclusion of the Yukon and the Northwest Teritories. The document has been produced to facilitate the manipulation of the 1996-1997 microdata file containing survey results. The main variables include: demography, health status, chronic conditions, restriction of activity, socio-demographic, and others.

    Release date: 2000-08-02

  • Public use microdata: 12M0010X
    Description:

    Cycle 10 collected data from persons 15 years and older and concentrated on the respondent's family. Topics covered include marital history, common- law unions, biological, adopted and step children, family origins, child leaving and fertility intentions.

    The target population of the GSS (General Social Survey) consisted of all individuals aged 15 and over living in a private household in one of the ten provinces.

    Release date: 1997-02-28
Analysis (62)

Analysis (62) (0 to 10 of 62 results)

  • Journals and periodicals: 12-206-X
    Description: This report summarizes the annual achievements of the Methodology Research and Development Program (MRDP) sponsored by the Modern Statistical Methods and Data Science Branch at Statistics Canada. This program covers research and development activities in statistical methods with potentially broad application in the agency’s statistical programs; these activities would otherwise be less likely to be carried out during the provision of regular methodology services to those programs. The MRDP also includes activities that provide support in the application of past successful developments in order to promote the use of the results of research and development work. Selected prospective research activities are also presented.
    Release date: 2025-10-10

  • Articles and reports: 11-522-X202500100034
    Description: Until now, detailed data on the destination of manufacturing sales have not historically been available to Canadians. Through integration of annual survey data, a destination of sales table by industry and province of origin was developed for the annual and monthly manufacturing surveys at Statistics Canada. Respondents for the annual survey are asked for their distribution of sales as a percentage across 15 destinations. To tackle the difficulty of generating an establishment-level distribution for multi-province respondents, three approaches were compared: using the respondents' total distribution for all their establishments, using optimization, and using the distributions of the single-province respondents. The imputed distribution of destination sales from the annual data was then applied to the monthly survey's sales value. This paper delves into challenges faced with imputing the destination sales (especially for respondents with establishments in multiple provinces), ensuring sales match marginal origin province totals, and allocating a distribution of destinations based on data from the annual program to the monthly estimates.
    Release date: 2025-09-08

  • Articles and reports: 11-522-X202500100035
    Description: Historically, the Canadian census of population Edit and Imputation (E&I) process has operated using a nearest-neighbour donor imputation methodology wherein the distance between a failed unit and a potential donor is obtained through a weighted combination of auxiliary variables. Revision to the model between cycles can be a complicated and time-consuming process given there is no standard approach to variable selection and weighting between topics. This paper will illustrate the potential of the Relief variable selection algorithm to create a machine learning-driven approach to variable selection and weighting that is standardized and comparable between census cycles and among the many topics of the census. An overview on how this process may be applied in practice will be presented, followed by results on several topics that indicate a general improvement over previous methods.
    Release date: 2025-09-08

  • Articles and reports: 12-001-X202500100004
    Description: Survey data collection often is plagued by unit and item nonresponse. To reduce reliance on strong assumptions about the missingness mechanisms, statisticians can use information about population marginal distributions known, for example, from censuses or administrative databases. One approach that does so is the Missing Data with Auxiliary Margins, or MD-AM, framework, which uses multiple imputation for both unit and item nonresponse so that survey-weighted estimates accord with the known marginal distributions. However, this framework relies on specifying and estimating a joint distribution for the survey data and nonresponse indicators, which can be computationally and practically daunting in data with many variables of mixed types. We propose two adaptations to the MD-AM framework to simplify the imputation task. First, rather than specifying a joint model for unit respondents’ data, we use random hot deck imputation while still leveraging the known marginal distributions. Second, instead of sampling from conditional distributions implied by the joint model for the missing data due to item nonresponse, we apply multiple imputation by chained equations for item nonresponse before imputation for unit nonresponse. Using simulation studies with nonignorable missingness mechanisms, we demonstrate that the proposed approach can provide more accurate point and interval estimates than models that do not leverage the auxiliary information. We illustrate the approach using data on voter turnout from the U.S. Current Population Survey.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100013
    Description: This discussion of the paper by Rao and Lohr focuses on the use of machine learning procedures for estimating finite population parameters. While there is growing interest in these methods within national statistical offices, several areas remain largely unexplored and warrant significant attention in the coming years. In this discussion, I highlight potential topics for future research and development in this rapidly evolving field.
    Release date: 2025-06-30

  • Journals and periodicals: 62F0026M
    Description: This series provides detailed documentation on the issues, concepts, methodology, data quality and other relevant research related to household expenditures from the Survey of Household Spending, the Homeowner Repair and Renovation Survey and the Food Expenditure Survey.
    Release date: 2025-05-21

  • Articles and reports: 12-001-X202200100008
    Description:

    The Multiple Imputation of Latent Classes (MILC) method combines multiple imputation and latent class analysis to correct for misclassification in combined datasets. Furthermore, MILC generates a multiply imputed dataset which can be used to estimate different statistics in a straightforward manner, ensuring that uncertainty due to misclassification is incorporated when estimating the total variance. In this paper, it is investigated how the MILC method can be adjusted to be applied for census purposes. More specifically, it is investigated how the MILC method deals with a finite and complete population register, how the MILC method can simultaneously correct misclassification in multiple latent variables and how multiple edit restrictions can be incorporated. A simulation study shows that the MILC method is in general able to reproduce cell frequencies in both low- and high-dimensional tables with low amounts of bias. In addition, variance can also be estimated appropriately, although variance is overestimated when cell frequencies are small.

    Release date: 2022-06-21

  • Articles and reports: 12-001-X202000100006
    Description:

    In surveys, logical boundaries among variables or among waves of surveys make imputation of missing values complicated. We propose a new regression-based multiple imputation method to deal with survey nonresponses with two-sided logical boundaries. This imputation method automatically satisfies the boundary conditions without an additional acceptance/rejection procedure and utilizes the boundary information to derive an imputed value and to determine the suitability of the imputed value. Simulation results show that our new imputation method outperforms the existing imputation methods for both mean and quantile estimations regardless of missing rates, error distributions, and missing-mechanisms. We apply our method to impute the self-reported variable “years of smoking” in successive health screenings of Koreans.

    Release date: 2020-06-30

  • Articles and reports: 12-001-X201800254957
    Description:

    When a linear imputation method is used to correct non-response based on certain assumptions, total variance can be assigned to non-responding units. Linear imputation is not as limited as it seems, given that the most common methods – ratio, donor, mean and auxiliary value imputation – are all linear imputation methods. We will discuss the inference framework and the unit-level decomposition of variance due to non-response. Simulation results will also be presented. This decomposition can be used to prioritize non-response follow-up or manual corrections, or simply to guide data analysis.

    Release date: 2018-12-20

  • Articles and reports: 11-633-X2017006
    Description:

    This paper describes a method of imputing missing postal codes in a longitudinal database. The 1991 Canadian Census Health and Environment Cohort (CanCHEC), which contains information on individuals from the 1991 Census long-form questionnaire linked with T1 tax return files for the 1984-to-2011 period, is used to illustrate and validate the method. The cohort contains up to 28 consecutive fields for postal code of residence, but because of frequent gaps in postal code history, missing postal codes must be imputed. To validate the imputation method, two experiments were devised where 5% and 10% of all postal codes from a subset with full history were randomly removed and imputed.

    Release date: 2017-03-13
Reference (24)

Reference (24) (0 to 10 of 24 results)

  • Surveys and statistical programs – Documentation: 12-539-X
    Description:

    This document brings together guidelines and checklists on many issues that need to be considered in the pursuit of quality objectives in the execution of statistical activities. Its focus is on how to assure quality through effective and appropriate design or redesign of a statistical project or program from inception through to data evaluation, dissemination and documentation. These guidelines draw on the collective knowledge and experience of many Statistics Canada employees. It is expected that Quality Guidelines will be useful to staff engaged in the planning and design of surveys and other statistical projects, as well as to those who evaluate and analyze the outputs of these projects.

    Release date: 2019-12-04

  • Surveys and statistical programs – Documentation: 99-012-X2011006
    Geography: Canada
    Description:

    This reference guide provides information that enables users to effectively use, apply and interpret data from the 2011 National Household Survey (NHS). This guide contains definitions and explanations of concepts, classifications, data quality and comparability to other sources. Additional information is included for specific variables to help general users better understand the concepts and questions used in the NHS.

    Release date: 2013-06-26

  • Surveys and statistical programs – Documentation: 99-012-X2011007
    Description:

    This reference guide provides information that enables users to effectively use, apply and interpret data from the 2011 National Household Survey (NHS). This guide contains definitions and explanations of concepts, classifications, data quality and comparability to other sources. Additional information is included for specific variables to help general users better understand the concepts and questions used in the NHS.

    Release date: 2013-06-26

  • Surveys and statistical programs – Documentation: 99-012-X2011008
    Description:

    This reference guide provides information that enables users to effectively use, apply and interpret data from the 2011 National Household Survey (NHS). This guide contains definitions and explanations of concepts, classifications, data quality and comparability to other sources. Additional information is included for specific variables to help general users better understand the concepts and questions used in the NHS.

    Release date: 2013-06-26

  • Surveys and statistical programs – Documentation: 99-013-X2011006
    Description:

    This reference guide provides information that enables users to effectively use, apply and interpret data from the 2011 National Household Survey (NHS). This guide contains definitions and explanations of concepts, classifications, data quality and comparability to other sources. Additional information is included for specific variables to help general users better understand the concepts and questions used in the NHS.

    Release date: 2013-06-26

  • Surveys and statistical programs – Documentation: 62F0026M2010004
    Description:

    This report describes the quality indicators produced for the 2007 Survey of Household Spending. These quality indicators, such as coefficients of variation, nonresponse rates, slippage rates and imputation rates, help users interpret the survey data.

    Release date: 2010-12-13

  • Surveys and statistical programs – Documentation: 62F0026M2010005
    Description:

    This report describes the quality indicators produced for the 2008 Survey of Household Spending. These quality indicators, such as coefficients of variation, nonresponse rates, slippage rates and imputation rates, help users interpret the survey data.

    Release date: 2010-12-13

  • Surveys and statistical programs – Documentation: 75F0002M2008005
    Description: The Survey of Labour and Income Dynamics (SLID) is a longitudinal survey initiated in 1993. The survey was designed to measure changes in the economic well-being of Canadians as well as the factors affecting these changes. Sample surveys are subject to sampling errors. In order to consider these errors, each estimates presented in the "Income Trends in Canada" series comes with a quality indicator based on the coefficient of variation. However, other factors must also be considered to make sure data are properly used. Statistics Canada puts considerable time and effort to control errors at every stage of the survey and to maximise the fitness for use. Nevertheless, the survey design and the data processing could restrict the fitness for use. It is the policy at Statistics Canada to furnish users with measures of data quality so that the user is able to interpret the data properly. This report summarizes the set of quality measures of SLID data. Among the measures included in the report are sample composition and attrition rates, sampling errors, coverage errors in the form of slippage rates, response rates, tax permission and tax linkage rates, and imputation rates.
    Release date: 2008-08-20

  • Surveys and statistical programs – Documentation: 92-393-X
    Description:

    This report is a brief guide to users of census income data. It provides a general description of the various 2001 Census phases, from data collection, through processing for non-response, to dissemination. Descriptions of, and summary data on, the changes to income data that occurred during the processing stages are given. Comparative data from national accounts and tax data sources at a highly aggregated level are also presented to put the quality of the 2001 Census income data into perspective. For users wishing to compare census income data over time, changes in income content and universe coverage over the years are explained. Finally, a complete description of all census products containing income data is also supplied.

    Release date: 2004-09-16

  • Surveys and statistical programs – Documentation: 92-390-X
    Description:

    This report includes a definition of the 2001 place of work concept and the place of work geography, standard text on data collection and coverage (including data collection methods, special coverage studies, sampling and weighting, edit and follow-up, coverage and content considerations). Both standard and subject-matter specific text pieces are also included for data assimilation (automated as well as interactive coding), edit and imputation and data evaluation. Finally, this technical report includes a section on historical comparability.

    Release date: 2004-08-26