Filter results by

Search Help
Currently selected filters that can be removed

Keyword(s)

Year of publication

1 facets displayed. 1 facets selected.

Survey or statistical program

134 facets displayed. 0 facets selected.

Content

1 facets displayed. 0 facets selected.
Sort Help
entries

Results

All (393)

All (393) (0 to 10 of 393 results)

  • Articles and reports: 13-26-0004
    Description: StatCan's accessibility plan aims to ensure that all StatCan and Statistical Survey Operations employees are supported in a barrier-free environment, with their accessibility needs met. Statistics Canada: Road to Accessibility 2023-25 is intended to be evergreen. As we make progress toward achieving an accessible and inclusive StatCan, our actions and commitments will change and evolve, and the Plan will be updated to ensure a continued and relevant focus on the areas needing it most.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200001
    Description: Nested error regression models are commonly used to incorporate unit specific auxiliary variables to improve small area estimates. When the mean structure of the model is misspecified, the design-based mean squared prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP) generally increases. The Observed Best Prediction (OBP) method has been proposed with the intent to improve on the design-based MSPE over EBLUP. In this paper, we conduct a Monte Carlo simulation experiments to understand the effect of misspsecification of mean structures on different small area estimators. Our findings suggest that the OBP using unit-level auxiliary variables does not outperform the EBLUP in terms of design-based MSPE, unless the number of small areas m is extremely large. Conversely, the performance of OBP significantly improves when area-level auxiliary variables are employed. This paper includes both analytical and numerical evidence to demonstrate these observations, providing practical insights for addressing model misspecification in small area estimation (SAE).
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200002
    Description: This study examines interviewer effects on household nonresponse in three waves of the Household Finance and Consumption Survey (HFCS) in Austria using a multilevel model. Addressing nonresponse at its source is crucial for maintaining survey data quality and representativeness. Our findings indicate that the variation in response behavior explained by interviewer effects decreased from about one-third in the first wave to 7% in the third wave. Effective interviewers tend to have a university degree, be married, homeowners, and have a larger workload. Additionally, higher mean wages in the household’s municipality negatively affect survey participation. These insights suggest targeted interviewer selection and training strategies to improve response rates.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200003
    Description: In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200004
    Description: The class of generalized linear models (GLM) is a flexible generalization of ordinary least squares regression that allows the linear model to be related to the response variable via a link function and assumes the magnitude of the variance of each measurement to be a function of its predicted value. Multicollinearity in GLMs can inflate variances of the estimated coefficients and cause poor prediction in certain regions of the regression space. It may also cause a nonsignificant Wald statistic even when the predictors are highly predictive in a model of the family of GLMs. Little previous research has closely investigated the diagnostics of multicollinearity in GLMs, especially when complex survey data are used. In this paper, we develop variance inflation factors (VIFs) that measure the amount that the variance of a parameter estimator is increased due to multicollinearity in GLMs. We also extend VIFs and condition indexes to apply to complex survey data, accounting for design features, e.g. weights, clusters, and strata. Illustrations of these methods are given using data from a household survey of health and nutrition.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200005
    Description: The use of non-probability data sources for statistical purposes and for official statistics has become increasingly popular in recent years. However, statistical inference based on non-probability samples is made more difficult by nature of their biasedness and lack of representativity. In this paper we propose quantile balancing inverse probability weighting estimator (QBIPW) for non-probability samples. We apply the idea of Harms and Duchesne (2006) allowing the use of quantile information in the estimation process to reproduce known totals and the distribution of auxiliary variables. We discuss the estimation of the QBIPW probabilities and its variance. Our simulation study has demonstrated that the proposed estimators are robust against model mis-specification and, as a result, help to reduce bias and mean squared error. Finally, we applied the proposed methods to estimate the share of job vacancies aimed at Ukrainian workers in Poland using an integrated set of administrative and survey data about job vacancies.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200006
    Description: National Statistical Institutes (NSIs) are directing resources into advancing the use of administrative data in official statistics. Administrative data, however, are not developed for the purpose of producing statistics rather as a result of an event or transaction relating to administrative procedures of organizations, public administrations and government agencies. Therefore, it is essential to check the quality of the administrative data with respect to sources of error, particularly representativeness to the target population. In this paper, we utilize the strength of probability-based reference samples or censuses that can be used to detect the lack of representativeness in administrative data and introduce quality indicators based on distance metrics and representativity indicators (R-indicators). We demonstrate their application with a simulation study and discuss a real application applied on a UK Office for National Statistics (ONS) administrative dataset.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200007
    Description: Although probability samples have been regarded as the gold standard to collect information for population-based study, non-probability samples have been used frequently in practice due to low cost, convenience, and the lack of the sampling frame for the survey. Naïve estimates based on non-probability samples without any adjustments may be misleading due to selection bias. Recently, a valid data integration approach that includes mass imputation, propensity score weighting, and calibration has been used to improve the representativeness of non-probability samples. The effectiveness of the mass imputation approach depends on the underlying model assumptions. In this paper, we propose using deep learning for the mass imputation in the combining of probability and non-probability samples and compare it with several modern machine learning-based mass imputation approaches, including generalized additive modeling, regression tree, random forest, and XG-boosting. In the simulation study, deep learning-based approaches have been shown to be more robust and effective than other mass imputation approaches against the failure of underlying model assumptions under non-linearity scenarios.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200008
    Description: Classical design-based survey estimation relies on a properly specified sampling design for valid inference. We consider the properties of regression estimation under a misspecified sample design, in which the nominal and true inclusion probabilities do not necessarily match. This general misspecified sample design setting encompasses many challenges in the modern survey environment. Under this setting, an asymptotic analysis of the regression estimator, an expression of the bias, and an expression of the variance are presented. Further, a consistent variance estimator is derived and an expression which estimates the bias in-part or in-whole is discussed. This later expression may be used as an indicator of the presence of bias due to misspecification by a practitioner. A simulation study is conducted to support the presented theory.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200009
    Description: We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our methodology uses hierarchical global-local prior distributions to make inferences for the proportion of individuals in Florida’s counties who do not have health insurance. Results from an extensive simulation study show that this methodology will provide improved inference by using several data sources. Among the five model variants evaluated the ones using horseshoe priors for all variances have better performance than the ones using lasso priors for the local variances.
    Release date: 2025-12-23
Stats in brief (73)

Stats in brief (73) (0 to 10 of 73 results)

  • Stats in brief: 45-20-00032025009
    Description: Think your Christmas tree knowledge is top-notch? Time to put it to the test! In this holiday special of Eh Sayers, our colleagues face off in a trivia showdown all about Canada’s Christmas tree industry. Discover which nation first sparked the tradition, how much Canadian Christmas tree farmers earned in 2023, the surprising numbers behind exports and imports, and more festive facts. After listening, you’ll be the go-to expert on Christmas trees at your holiday party!
    Release date: 2025-12-19

  • Stats in brief: 11-621-M2025016
    Description: Amid ongoing shifts in tariffs, trade regulations and U.S. policy, the Canadian Survey on Business Conditions has remained an important tool for understanding the impact of tariffs on businesses in Canada. Building on analysis from the second and third quarters of 2025, which captured the introduction of newly imposed tariffs, followed by reactions to countermeasures, this paper presents updated findings for the fourth quarter of 2025 from the survey that was conducted from October 1 to November 5, 2025.
    Release date: 2025-12-18

  • Stats in brief: 89-20-00062025001
    Description: This video is designed to help you critically assess the data presented to you. No data is perfect. By understanding the strengths and limitations of the data, you can avoid being misled—and make smarter, more informed decisions.
    Release date: 2025-12-15

  • Stats in brief: 11-627-M2025059
    Description: This infographic from the Rural Data Lab of Statistics Canada presents a visual overview of key highlights and statistical findings from the 2023 Rural Canada Business Profiles database. The database provides financial profiles for small and medium-sized businesses in Canada with total annual revenues of $30,000 to $5,000,000 and $5,000,001 to $20,000,000, respectively. The infographic provides observations focused on small businesses in rural and small town Canada.
    Release date: 2025-12-12

  • Stats in brief: 11-627-M2025061
    Description: This infographic provides an overview of the lumber industry, showcasing key metrics and trends related to production, exports and price change. It highlights significant data points, illustrating the state of the market and offering insights into the current landscape of lumber in Canada.
    Release date: 2025-12-12

  • Stats in brief: 11-627-M2025055
    Description: Using police-reported data from the Uniform Crime Reporting Survey, this infographic presents data on subsequent contacts with police over a nine-year period for individuals living in rural areas of the Canadian provinces who were accused of a crime in 2014.
    Release date: 2025-12-09

  • Stats in brief: 11-627-M2025051
    Description: Using data from the 2022 Canadian Survey on Disability, this infographic highlights the trends and experiences of racialized persons with disabilities.
    Release date: 2025-12-03

  • Stats in brief: 11-627-M2025052
    Description: Utilizing data from the 2022 Canadian Survey on Disability, this infographic highlights the trends and experiences of persons with dynamic disabilities.
    Release date: 2025-12-03

  • Stats in brief: 11-627-M2025060
    Description: Police-reported homicide victims and rates per 100,000 population, by geographic region (national, provincial, and CMA).
    Release date: 2025-12-02

  • Stats in brief: 11-627-M2025056
    Description: This infographic features government spending data in Canada for the 2024-2025 fiscal year. It gives a breakdown of expenses by the socio-economic purpose for which the funds are used.
    Release date: 2025-11-27
Articles and reports (299)

Articles and reports (299) (0 to 10 of 299 results)

  • Articles and reports: 13-26-0004
    Description: StatCan's accessibility plan aims to ensure that all StatCan and Statistical Survey Operations employees are supported in a barrier-free environment, with their accessibility needs met. Statistics Canada: Road to Accessibility 2023-25 is intended to be evergreen. As we make progress toward achieving an accessible and inclusive StatCan, our actions and commitments will change and evolve, and the Plan will be updated to ensure a continued and relevant focus on the areas needing it most.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200001
    Description: Nested error regression models are commonly used to incorporate unit specific auxiliary variables to improve small area estimates. When the mean structure of the model is misspecified, the design-based mean squared prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP) generally increases. The Observed Best Prediction (OBP) method has been proposed with the intent to improve on the design-based MSPE over EBLUP. In this paper, we conduct a Monte Carlo simulation experiments to understand the effect of misspsecification of mean structures on different small area estimators. Our findings suggest that the OBP using unit-level auxiliary variables does not outperform the EBLUP in terms of design-based MSPE, unless the number of small areas m is extremely large. Conversely, the performance of OBP significantly improves when area-level auxiliary variables are employed. This paper includes both analytical and numerical evidence to demonstrate these observations, providing practical insights for addressing model misspecification in small area estimation (SAE).
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200002
    Description: This study examines interviewer effects on household nonresponse in three waves of the Household Finance and Consumption Survey (HFCS) in Austria using a multilevel model. Addressing nonresponse at its source is crucial for maintaining survey data quality and representativeness. Our findings indicate that the variation in response behavior explained by interviewer effects decreased from about one-third in the first wave to 7% in the third wave. Effective interviewers tend to have a university degree, be married, homeowners, and have a larger workload. Additionally, higher mean wages in the household’s municipality negatively affect survey participation. These insights suggest targeted interviewer selection and training strategies to improve response rates.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200003
    Description: In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200004
    Description: The class of generalized linear models (GLM) is a flexible generalization of ordinary least squares regression that allows the linear model to be related to the response variable via a link function and assumes the magnitude of the variance of each measurement to be a function of its predicted value. Multicollinearity in GLMs can inflate variances of the estimated coefficients and cause poor prediction in certain regions of the regression space. It may also cause a nonsignificant Wald statistic even when the predictors are highly predictive in a model of the family of GLMs. Little previous research has closely investigated the diagnostics of multicollinearity in GLMs, especially when complex survey data are used. In this paper, we develop variance inflation factors (VIFs) that measure the amount that the variance of a parameter estimator is increased due to multicollinearity in GLMs. We also extend VIFs and condition indexes to apply to complex survey data, accounting for design features, e.g. weights, clusters, and strata. Illustrations of these methods are given using data from a household survey of health and nutrition.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200005
    Description: The use of non-probability data sources for statistical purposes and for official statistics has become increasingly popular in recent years. However, statistical inference based on non-probability samples is made more difficult by nature of their biasedness and lack of representativity. In this paper we propose quantile balancing inverse probability weighting estimator (QBIPW) for non-probability samples. We apply the idea of Harms and Duchesne (2006) allowing the use of quantile information in the estimation process to reproduce known totals and the distribution of auxiliary variables. We discuss the estimation of the QBIPW probabilities and its variance. Our simulation study has demonstrated that the proposed estimators are robust against model mis-specification and, as a result, help to reduce bias and mean squared error. Finally, we applied the proposed methods to estimate the share of job vacancies aimed at Ukrainian workers in Poland using an integrated set of administrative and survey data about job vacancies.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200006
    Description: National Statistical Institutes (NSIs) are directing resources into advancing the use of administrative data in official statistics. Administrative data, however, are not developed for the purpose of producing statistics rather as a result of an event or transaction relating to administrative procedures of organizations, public administrations and government agencies. Therefore, it is essential to check the quality of the administrative data with respect to sources of error, particularly representativeness to the target population. In this paper, we utilize the strength of probability-based reference samples or censuses that can be used to detect the lack of representativeness in administrative data and introduce quality indicators based on distance metrics and representativity indicators (R-indicators). We demonstrate their application with a simulation study and discuss a real application applied on a UK Office for National Statistics (ONS) administrative dataset.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200007
    Description: Although probability samples have been regarded as the gold standard to collect information for population-based study, non-probability samples have been used frequently in practice due to low cost, convenience, and the lack of the sampling frame for the survey. Naïve estimates based on non-probability samples without any adjustments may be misleading due to selection bias. Recently, a valid data integration approach that includes mass imputation, propensity score weighting, and calibration has been used to improve the representativeness of non-probability samples. The effectiveness of the mass imputation approach depends on the underlying model assumptions. In this paper, we propose using deep learning for the mass imputation in the combining of probability and non-probability samples and compare it with several modern machine learning-based mass imputation approaches, including generalized additive modeling, regression tree, random forest, and XG-boosting. In the simulation study, deep learning-based approaches have been shown to be more robust and effective than other mass imputation approaches against the failure of underlying model assumptions under non-linearity scenarios.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200008
    Description: Classical design-based survey estimation relies on a properly specified sampling design for valid inference. We consider the properties of regression estimation under a misspecified sample design, in which the nominal and true inclusion probabilities do not necessarily match. This general misspecified sample design setting encompasses many challenges in the modern survey environment. Under this setting, an asymptotic analysis of the regression estimator, an expression of the bias, and an expression of the variance are presented. Further, a consistent variance estimator is derived and an expression which estimates the bias in-part or in-whole is discussed. This later expression may be used as an indicator of the presence of bias due to misspecification by a practitioner. A simulation study is conducted to support the presented theory.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200009
    Description: We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our methodology uses hierarchical global-local prior distributions to make inferences for the proportion of individuals in Florida’s counties who do not have health insurance. Results from an extensive simulation study show that this methodology will provide improved inference by using several data sources. Among the five model variants evaluated the ones using horseshoe priors for all variances have better performance than the ones using lasso priors for the local variances.
    Release date: 2025-12-23
Journals and periodicals (21)

Journals and periodicals (21) (0 to 10 of 21 results)

  • Table: 57-003-X
    Description: This publication presents definitions, explanatory notes, methodology, and energy conversion factors for the Report on Energy Supply and Demand in Canada.
    Release date: 2025-11-13

  • Journals and periodicals: 45-26-0001
    Description: The Departmental Sustainable Development Strategy (DSDS) outlines departmental actions, with measurable performance indicators, that support the implementation strategies of the 2022-2026 Federal Sustainable Development Strategy. The DSDS further outlines Statistics Canada’s sustainable development vision to produce data to help track whether Canada is moving toward a more sustainable future and highlights projects with links to supporting sustainable development goals.
    Release date: 2025-10-31

  • Journals and periodicals: 11-631-X
    Description: Statistics Canada regularly prepares presentations with statistical findings about the country’s economy, society and environment. These presentations may be intended for conferences, meetings with stakeholders, or other events held throughout the year to provide Statistics Canada with an opportunity to promote the role of official statistics and to better understand data users’ needs. This series provides online access to these presentations as well as new presentations created to help communicate research findings on a wide range of subjects to a broad audience.
    Release date: 2025-10-27

  • Journals and periodicals: 81-599-X
    Geography: Canada
    Description: The fact sheets in this series provide an "at-a-glance" overview of particular aspects of education in Canada and summarize key data trends in selected tables published as part of the Pan-Canadian Education Indicators Program (PCEIP).

    The PCEIP mission is to publish a set of statistical measures on education systems in Canada for policy makers, practitioners and the general public to monitor the performance of education systems across jurisdictions and over time. PCEIP is a joint venture of Statistics Canada and the Council of Ministers of Education, Canada (CMEC).

    Release date: 2025-10-24

  • Journals and periodicals: 18-001-X
    Geography: Canada
    Description: Reports on Special Business Projects is an occasional series that focuses primarily on the results of special surveys or special projects conducted by the Centre for Special Business Projects. The reports cover a wide range of topics, which include business performance and trends, custom tabulations of business data, economic impact studies, new measurement frameworks and indicators to support program development, monitoring and performance assessment, territorial economic indicators and other special studies.
    Release date: 2025-10-10

  • Journals and periodicals: 12-206-X
    Description: This report summarizes the annual achievements of the Methodology Research and Development Program (MRDP) sponsored by the Modern Statistical Methods and Data Science Branch at Statistics Canada. This program covers research and development activities in statistical methods with potentially broad application in the agency’s statistical programs; these activities would otherwise be less likely to be carried out during the provision of regular methodology services to those programs. The MRDP also includes activities that provide support in the application of past successful developments in order to promote the use of the results of research and development work. Selected prospective research activities are also presented.
    Release date: 2025-10-10

  • Journals and periodicals: 16-508-X
    Description: Environment fact sheets will include short, focused, single-theme analysis on key issues within the changing environment with regards to all Canadians. Over the course of the series, analysis will include topics on: air and climate, pollution and waste, environmental protection and quality, and natural resources.
    Release date: 2025-10-02

  • Journals and periodicals: 82-622-X
    Geography: Canada
    Description: The Health Research Working Paper Series publishes: analytical work-in-progress; background documentation for specific research projects (e.g methodological papers); lengthy reports intended for specific clients, and; compendiums of data tables. Publication in this series does not preclude publication of specific aspects of the work in a peer-reviewed journal.
    Release date: 2025-10-02

  • Journals and periodicals: 11-522-X
    Description: Since 1984, an annual international symposium on methodological issues has been sponsored by Statistics Canada. Proceedings have been available since 1987.
    Release date: 2025-09-08

  • Journals and periodicals: 14-28-0001
    Description: Statistics Canada's Quality of Employment in Canada publication is intended to provide Canadians and Canadian organizations with a better understanding of quality of employment using an internationally-supported statistical framework. Quality of employment is approached as a multidimensional concept, characterized by different elements, which relate to human needs in various ways. To cover all relevant aspects, the framework identified seven dimensions and twelve sub-dimensions of quality of employment.
    Release date: 2025-08-18