Keyword search

Filter results by

Search Help
Currently selected filters that can be removed

Keyword(s)

Geography

2 facets displayed. 0 facets selected.

Survey or statistical program

64 facets displayed. 0 facets selected.

Content

1 facets displayed. 1 facets selected.
Sort Help
entries

Results

All (332)

All (332) (0 to 10 of 332 results)

  • Stats in brief: 89-20-00062026001
    Description: In this video you will learn about the steps and activities in the data journey. The data journey represents the key stages of the data process. The journey is not necessarily linear. It is intended to represent the different steps and activities that could be undertaken to produce meaningful information from data. Not everyone who uses data will do all of these steps.
    Release date: 2026-07-09

  • Articles and reports: 12-001-X202600100001
    Description: Wayne A. Fuller is a leading figure in statistics whose career at Iowa State University (ISU) began in 1959; he is now Distinguished Professor Emeritus in Statistics and Economics. This article briefly recounts his early life and training in agricultural economics at ISU and highlights influential contributions spanning time series analysis, measurement error models, and survey sampling. It documents his impact through seminal textbooks, methodological advances such as the Dickey-Fuller test and regression estimation, sustained work on major operational surveys (e.g., the National Resources Inventory), and mentorship of many graduate students. The article includes an interview conducted on May 20th, 2025, at Professor Fuller’s home.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100002
    Description: Survey data typically have missing values due to unit and item nonresponse. Sometimes, survey organizations know the marginal distributions of certain categorical variables in the target population. As shown in previous work, survey organizations can leverage these distributions in multiple imputation for nonignorable unit nonresponse, generating imputations that result in plausible completed-data estimates for the variables with known margins. However, this prior work does not use the design weights for unit nonrespondents. We extend this previous work to utilize the design weights for all sampled units. We illustrate the approach using simulation studies.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100003
    Description: Probability-proportional-to-size sampling is widely used by national statistical offices. Here population units are selected with probabilities proportional to an auxiliary variable. Variance formulas in such designs require both first- and second-order inclusion probabilities. The computation of second-order inclusion probabilities is particularly challenging for large populations, and has been the subject of extensive research. This article presents some new exact and approximation formulas for second-order inclusion probabilities in randomized systematic sampling with unequal probabilities and without replacement.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100004
    Description: We test the notion that a quasi-probabilistic method of selecting individuals within households (last birthday, LB) draws in a different sample compared to a non-probabilistic approach that selects respondents according to known parameters on age and gender (frequency matching, FM). With data from an original field experiment, we evaluate fieldwork efficiency (time and completed cases), economy (cost), success in recruiting a representative sample, and differences across a set of attitudinal and behavioral measures. We find that the FM approach performs better on efficiency and cost and achieves a comparable sample; importantly, this comparability extends across measures of personality traits and public opinion. With appropriate caveats, we conclude that researchers’ choice of selection methods should be guided by both theoretical benefits and practical tradeoffs.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100005
    Description: Confidence intervals are very often constructed based on a probability distribution that uses a certain number of degrees of freedom as a parameter. This is the case with the Student and the modified Wilson confidence intervals, discussed in this article, which use quantiles from the Student distribution where the number of degrees of freedom is generally unknown. For the length of a confidence interval to be representative of the reliability of an estimate, the actual coverage rate must match the nominal rate. To that end, the number of degrees of freedom in the probability distribution used in practice to calculate the confidence interval must be estimated as precisely as possible. An approximate rule is often used, although it tends to overestimate the actual number of degrees of freedom. In this article, a more precise version of degrees of freedom, derived from the Satterthwaite approximation, is obtained in the context of the Canadian Census of Population. The sampling design is equivalent to a simple random design without replacement, cluster-stratified, and the variance estimation method is an adaptation of the balanced repeated replication method. An explicit expression of the degrees of freedom is obtained under these conditions, enabling the factors influencing them to be identified. For comparison, the degree of freedom formula is also established for the conventional variance estimator. A simulation study shows that using this version of degrees of freedom corrects the undercoverage problem observed with the approximate rule, showing the importance of accurately assessing this number.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100006
    Description: We introduce a general framework for constructing master samples that preserve desirable design properties across panels. The core procedure is to order an initial probability sample. Since the final sequence must be robust to a uniform random rotation, we define and minimize an objective that aggregates panel-level performance across all possible circular panels. A final random rotation is applied to ensure design validity. The framework is flexible with respect to the choice of design criteria, such as spatial balance or marginal balance, and can be implemented efficiently using simulated annealing to obtain high-quality approximate solutions. By construction, the approach supports both positive and negative sample coordination for spatially balanced, marginally balanced, and doubly balanced samples. The method’s versatility is demonstrated through three applications: constructing a master sample with spatially balanced panels, marginally balanced panels, and doubly balanced panels.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100007
    Description: National statistical institutes operate sample coordination systems to spread the response burden in business surveys. Despite the applied sample coordination and monitoring the response burden, some businesses might still be heavily sampled within a short period. This may lead to a peaking response burden for individual businesses, which could affect response rates and response quality. This paper proposes a new sample coordination method based on Adapted Spatially Correlated Poisson (ASCP) sampling that focuses on businesses with a high response burden. The effects on the response burden will be evaluated in two simulation studies and compared with a stratified approach, a pragmatic method in which sampling fractions are manually adjusted and with the baseline method of ignoring the response burden. For the simulations, real-world scenarios and data from Statistics Netherlands are used. The first simulation study considers a practical situation in which a given sample is adjusted with the aim to avoid the occurrence of businesses with a peaking response burden. The second simulation study analyzes the longer-term effects of the different sample coordination methods and focuses both on the reduction and spread of the response burden. The advantages and disadvantages of the different methods will be explained and discussed in detail, and recommendations for applying these methods at national statistical institutes and other survey agencies will be given.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100008
    Description: This paper introduces an innovative and intuitive finite population sampling method that has been developed using a unique graphical framework. In this approach, first-order inclusion probabilities are represented as bars on a two-dimensional graph. By manipulating the positions of these bars, researchers can create a wide range of different sampling designs. This graphical visualization of sampling designs facilitates the exploration of alternative designs and may simplify certain aspects of the implementation compared to traditional mathematical algorithms. This novel approach holds significant promise for tackling complex challenges in sampling, such as achieving an optimal design. By applying a version of the greedy best-first search algorithm to this graphical approach, the potential for integrating intelligent algorithms into finite population sampling is demonstrated.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100009
    Description: Combining estimates from independent surveys via inverse-variance weights can lead to negative bias when unknown variances are estimated and the target variable is non-negative and positively skewed. In such cases, strong positive correlations typically arise between the estimators and their corresponding variance estimators, causing standard linear combinations with inverse-variance weights to exhibit negative bias. We introduce a strikingly simple method to reduce bias: replace the standard weight with the ratio of the estimator to the variance estimator. Under a linear model linking the two, we show that the new ratio-weighted estimator is approximately unbiased, whereas the conventional inverse-variance combination exhibits downward bias. Through simulations, we demonstrate that the new method brings both the bias and the mean squared error closer to the optimum for a wide range of different target variables. As our method uses only standardly reported summary statistics, it can be immediately adopted to reduce this widespread bias and improve the reliability of scientific findings in various fields.
    Release date: 2026-06-29
Data (4)

Data (4) ((4 results))

  • Profile of a community or region: 46-26-0002
    Description: The National Address Register (NAR) is a list of commercial and residential addresses in Canada that are extracted from Statistics Canada's Building Register and deemed non-confidential.
    Release date: 2026-06-26

  • Table: 89-26-0006
    Description: PASSAGES is an open-source dynamic microsimulation model aimed at supporting policy analysis and research relating to Canadian retirement income system outcomes at the individual and family level. The publicly available version includes a synthetic starting database, a model, and documentation. A confidential starting database is also available.
    Release date: 2026-06-26

  • Public use microdata: 89F0002X
    Description: The SPSD/M is a static microsimulation model designed to analyse financial interactions between governments and individuals in Canada. It can compute taxes paid to and cash transfers received from government. It is comprised of a database, a series of tax/transfer algorithms and models, analytical software and user documentation.
    Release date: 2026-02-12

  • Table: 11-10-0074-01
    Geography: Census tract
    Frequency: Occasional
    Description:

    The divergence index (D-index) describes the degree that families with different income levels are mixing together in neighbourhoods. It compares neighbourhood (census tract, CT) discrete income distributions to a base distribution, which is the income quintiles of the neighbourhood’s census metropolitan area (CMA).

    Release date: 2020-06-22
Analysis (215)

Analysis (215) (0 to 10 of 215 results)

  • Stats in brief: 89-20-00062026001
    Description: In this video you will learn about the steps and activities in the data journey. The data journey represents the key stages of the data process. The journey is not necessarily linear. It is intended to represent the different steps and activities that could be undertaken to produce meaningful information from data. Not everyone who uses data will do all of these steps.
    Release date: 2026-07-09

  • Articles and reports: 12-001-X202600100001
    Description: Wayne A. Fuller is a leading figure in statistics whose career at Iowa State University (ISU) began in 1959; he is now Distinguished Professor Emeritus in Statistics and Economics. This article briefly recounts his early life and training in agricultural economics at ISU and highlights influential contributions spanning time series analysis, measurement error models, and survey sampling. It documents his impact through seminal textbooks, methodological advances such as the Dickey-Fuller test and regression estimation, sustained work on major operational surveys (e.g., the National Resources Inventory), and mentorship of many graduate students. The article includes an interview conducted on May 20th, 2025, at Professor Fuller’s home.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100002
    Description: Survey data typically have missing values due to unit and item nonresponse. Sometimes, survey organizations know the marginal distributions of certain categorical variables in the target population. As shown in previous work, survey organizations can leverage these distributions in multiple imputation for nonignorable unit nonresponse, generating imputations that result in plausible completed-data estimates for the variables with known margins. However, this prior work does not use the design weights for unit nonrespondents. We extend this previous work to utilize the design weights for all sampled units. We illustrate the approach using simulation studies.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100003
    Description: Probability-proportional-to-size sampling is widely used by national statistical offices. Here population units are selected with probabilities proportional to an auxiliary variable. Variance formulas in such designs require both first- and second-order inclusion probabilities. The computation of second-order inclusion probabilities is particularly challenging for large populations, and has been the subject of extensive research. This article presents some new exact and approximation formulas for second-order inclusion probabilities in randomized systematic sampling with unequal probabilities and without replacement.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100004
    Description: We test the notion that a quasi-probabilistic method of selecting individuals within households (last birthday, LB) draws in a different sample compared to a non-probabilistic approach that selects respondents according to known parameters on age and gender (frequency matching, FM). With data from an original field experiment, we evaluate fieldwork efficiency (time and completed cases), economy (cost), success in recruiting a representative sample, and differences across a set of attitudinal and behavioral measures. We find that the FM approach performs better on efficiency and cost and achieves a comparable sample; importantly, this comparability extends across measures of personality traits and public opinion. With appropriate caveats, we conclude that researchers’ choice of selection methods should be guided by both theoretical benefits and practical tradeoffs.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100005
    Description: Confidence intervals are very often constructed based on a probability distribution that uses a certain number of degrees of freedom as a parameter. This is the case with the Student and the modified Wilson confidence intervals, discussed in this article, which use quantiles from the Student distribution where the number of degrees of freedom is generally unknown. For the length of a confidence interval to be representative of the reliability of an estimate, the actual coverage rate must match the nominal rate. To that end, the number of degrees of freedom in the probability distribution used in practice to calculate the confidence interval must be estimated as precisely as possible. An approximate rule is often used, although it tends to overestimate the actual number of degrees of freedom. In this article, a more precise version of degrees of freedom, derived from the Satterthwaite approximation, is obtained in the context of the Canadian Census of Population. The sampling design is equivalent to a simple random design without replacement, cluster-stratified, and the variance estimation method is an adaptation of the balanced repeated replication method. An explicit expression of the degrees of freedom is obtained under these conditions, enabling the factors influencing them to be identified. For comparison, the degree of freedom formula is also established for the conventional variance estimator. A simulation study shows that using this version of degrees of freedom corrects the undercoverage problem observed with the approximate rule, showing the importance of accurately assessing this number.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100006
    Description: We introduce a general framework for constructing master samples that preserve desirable design properties across panels. The core procedure is to order an initial probability sample. Since the final sequence must be robust to a uniform random rotation, we define and minimize an objective that aggregates panel-level performance across all possible circular panels. A final random rotation is applied to ensure design validity. The framework is flexible with respect to the choice of design criteria, such as spatial balance or marginal balance, and can be implemented efficiently using simulated annealing to obtain high-quality approximate solutions. By construction, the approach supports both positive and negative sample coordination for spatially balanced, marginally balanced, and doubly balanced samples. The method’s versatility is demonstrated through three applications: constructing a master sample with spatially balanced panels, marginally balanced panels, and doubly balanced panels.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100007
    Description: National statistical institutes operate sample coordination systems to spread the response burden in business surveys. Despite the applied sample coordination and monitoring the response burden, some businesses might still be heavily sampled within a short period. This may lead to a peaking response burden for individual businesses, which could affect response rates and response quality. This paper proposes a new sample coordination method based on Adapted Spatially Correlated Poisson (ASCP) sampling that focuses on businesses with a high response burden. The effects on the response burden will be evaluated in two simulation studies and compared with a stratified approach, a pragmatic method in which sampling fractions are manually adjusted and with the baseline method of ignoring the response burden. For the simulations, real-world scenarios and data from Statistics Netherlands are used. The first simulation study considers a practical situation in which a given sample is adjusted with the aim to avoid the occurrence of businesses with a peaking response burden. The second simulation study analyzes the longer-term effects of the different sample coordination methods and focuses both on the reduction and spread of the response burden. The advantages and disadvantages of the different methods will be explained and discussed in detail, and recommendations for applying these methods at national statistical institutes and other survey agencies will be given.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100008
    Description: This paper introduces an innovative and intuitive finite population sampling method that has been developed using a unique graphical framework. In this approach, first-order inclusion probabilities are represented as bars on a two-dimensional graph. By manipulating the positions of these bars, researchers can create a wide range of different sampling designs. This graphical visualization of sampling designs facilitates the exploration of alternative designs and may simplify certain aspects of the implementation compared to traditional mathematical algorithms. This novel approach holds significant promise for tackling complex challenges in sampling, such as achieving an optimal design. By applying a version of the greedy best-first search algorithm to this graphical approach, the potential for integrating intelligent algorithms into finite population sampling is demonstrated.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100009
    Description: Combining estimates from independent surveys via inverse-variance weights can lead to negative bias when unknown variances are estimated and the target variable is non-negative and positively skewed. In such cases, strong positive correlations typically arise between the estimators and their corresponding variance estimators, causing standard linear combinations with inverse-variance weights to exhibit negative bias. We introduce a strikingly simple method to reduce bias: replace the standard weight with the ratio of the estimator to the variance estimator. Under a linear model linking the two, we show that the new ratio-weighted estimator is approximately unbiased, whereas the conventional inverse-variance combination exhibits downward bias. Through simulations, we demonstrate that the new method brings both the bias and the mean squared error closer to the optimum for a wide range of different target variables. As our method uses only standardly reported summary statistics, it can be immediately adopted to reduce this widespread bias and improve the reliability of scientific findings in various fields.
    Release date: 2026-06-29
Reference (59)

Reference (59) (10 to 20 of 59 results)

  • Surveys and statistical programs – Documentation: 32-26-0008
    Description: This report describes the main changes, additions or deletions to the Census of Agriculture questionnaire by topic and in the order they appear on the questionnaire.
    Release date: 2025-07-04

  • Surveys and statistical programs – Documentation: 98-20-00052026004
    Description: This report provides detailed insight into the design and methodology of the content test component of the 2024 Census Test. This test evaluated changes to the wording and flow of some questions, as well as the potential addition of new questions, to help determine the content of the 2026 Census of Population.
    Release date: 2025-07-04

  • Surveys and statistical programs – Documentation: 11-633-X2024004
    Description: The Longitudinal Immigration Database (IMDB) is a comprehensive source of data that plays a key role in the understanding of the economic behaviour of immigrants. It is the only annual Canadian dataset that allows users to study the characteristics of immigrants to Canada at the time of admission and their economic outcomes and regional (inter-provincial) mobility over a time span of more than 40 years.
    Release date: 2024-12-09

  • Surveys and statistical programs – Documentation: 11-633-X2024005
    Description: The Analytical Studies and Modelling Branch is the research (ASMB), modelling, training and access hub of Statistics Canada. It focuses on leveraging the agency’s vast data holdings to generate in-depth insights that support evidence-based policy making and to enable others to do so through analytical training and data access. The ASMB, like other program areas in the agency, works to support Statistics Canada’s overall mission of delivering insights through data for a better Canada.
    Release date: 2024-12-06

  • Surveys and statistical programs – Documentation: 98-303-X
    Description: The Coverage Technical Report will present the errors included in census data that result from persons who are either missed (not enumerated) or enumerated more than once. The population coverage error is one of the most important types of errors because it affects the accuracy of not only population counts, but also all the census data results that describe the characteristics of the population universe.
    Release date: 2024-10-23

  • Surveys and statistical programs – Documentation: 89-653-X2024002
    Description: This guide is intended to provide a detailed review of both the 2022 IPS and IPS–NIS with respect to subject matter and methodological approaches. It is designed to help data users by serving as a guide to the concepts and measures of the survey as well as the technical details of the survey’s design, field work and data processing. This guide is meant to provide users with helpful information on how to use and interpret survey results. The discussion on data quality also allows users to review the strengths and limitations of the data for their particular needs.

    Chapter 1 of this guide provides an overview of the 2022 IPS and IPS–NIS by introducing the survey background and objectives. Chapter 2 outlines the survey’s themes and explains the key concepts and definitions used for the survey. Chapters 3 to 6 cover important aspects of the survey methodology, sampling design, data collection and processing. Chapters 7 and 8 review issues of data quality and caution users about comparing 2022 IPS or IPS–NIS data with data from other sources. Chapter 9 outlines the survey products available to the public, including data tables, analytical articles and reference material. The appendices provide a comprehensive list of survey indicators, extra coding categories and standard classifications used on both the IPS and the IPS–NIS. Lastly, a glossary of survey terms and information on confidence intervals is also provided.
    Release date: 2024-08-14

  • Surveys and statistical programs – Documentation: 75-514-G
    Description: The Guide to the Job Vacancy and Wage Survey contains a dictionary of concepts and definitions, and covers topics such as survey methodology, data collection, processing, and data quality. The guide covers both components of the survey: the job vacancy component, which is quarterly, and the wage component, which is annual.
    Release date: 2024-06-18

  • Surveys and statistical programs – Documentation: 32-26-0007
    Description: Census of Agriculture data provide statistical information on farms and farm operators at fine geographic levels and for small subpopulations. Quality evaluation activities are essential to ensure that census data are reliable and that they meet user needs.

    This report provides data quality information pertaining to the Census of Agriculture, such as sources of error, error detection, disclosure control methods, data quality indicators, response rates and collection rates.
    Release date: 2024-02-06

  • Surveys and statistical programs – Documentation: 11-633-X2024001
    Description: The Longitudinal Immigration Database (IMDB) is a comprehensive source of data that plays a key role in the understanding of the economic behaviour of immigrants. It is the only annual Canadian dataset that allows users to study the characteristics of immigrants to Canada at the time of admission and their economic outcomes and regional (inter-provincial) mobility over a time span of more than 35 years.
    Release date: 2024-01-22

  • Surveys and statistical programs – Documentation: 98-306-X
    Description:

    This report describes sampling, weighting and estimation procedures used in the Census of Population. It provides operational and theoretical justifications for them, and presents the results of the evaluations of these procedures.

    Release date: 2023-10-04