Weighting and estimation
Filter results by
Search HelpKeyword(s)
Type
Survey or statistical program
- Census of Population (3)
- Survey of Household Spending (1)
- Quarterly Demographic Estimates (1)
- Annual Demographic Estimates: Canada, Provinces and Territories (1)
- Estimates of the number of census families for July 1st, Canada, provinces and territories (1)
- Annual Demographic Estimates : Subprovincial Areas (1)
- Labour Force Survey (1)
- Longitudinal Administrative Databank (1)
- Longitudinal and International Study of Adults (1)
- Canadian Income Survey (1)
- Canadian Survey on Business Conditions (1)
Results
All (74)
All (74) (0 to 10 of 74 results)
- Articles and reports: 12-001-X202600100003Description: Probability-proportional-to-size sampling is widely used by national statistical offices. Here population units are selected with probabilities proportional to an auxiliary variable. Variance formulas in such designs require both first- and second-order inclusion probabilities. The computation of second-order inclusion probabilities is particularly challenging for large populations, and has been the subject of extensive research. This article presents some new exact and approximation formulas for second-order inclusion probabilities in randomized systematic sampling with unequal probabilities and without replacement.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100005Description: Confidence intervals are very often constructed based on a probability distribution that uses a certain number of degrees of freedom as a parameter. This is the case with the Student and the modified Wilson confidence intervals, discussed in this article, which use quantiles from the Student distribution where the number of degrees of freedom is generally unknown. For the length of a confidence interval to be representative of the reliability of an estimate, the actual coverage rate must match the nominal rate. To that end, the number of degrees of freedom in the probability distribution used in practice to calculate the confidence interval must be estimated as precisely as possible. An approximate rule is often used, although it tends to overestimate the actual number of degrees of freedom. In this article, a more precise version of degrees of freedom, derived from the Satterthwaite approximation, is obtained in the context of the Canadian Census of Population. The sampling design is equivalent to a simple random design without replacement, cluster-stratified, and the variance estimation method is an adaptation of the balanced repeated replication method. An explicit expression of the degrees of freedom is obtained under these conditions, enabling the factors influencing them to be identified. For comparison, the degree of freedom formula is also established for the conventional variance estimator. A simulation study shows that using this version of degrees of freedom corrects the undercoverage problem observed with the approximate rule, showing the importance of accurately assessing this number.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100009Description: Combining estimates from independent surveys via inverse-variance weights can lead to negative bias when unknown variances are estimated and the target variable is non-negative and positively skewed. In such cases, strong positive correlations typically arise between the estimators and their corresponding variance estimators, causing standard linear combinations with inverse-variance weights to exhibit negative bias. We introduce a strikingly simple method to reduce bias: replace the standard weight with the ratio of the estimator to the variance estimator. Under a linear model linking the two, we show that the new ratio-weighted estimator is approximately unbiased, whereas the conventional inverse-variance combination exhibits downward bias. Through simulations, we demonstrate that the new method brings both the bias and the mean squared error closer to the optimum for a wide range of different target variables. As our method uses only standardly reported summary statistics, it can be immediately adopted to reduce this widespread bias and improve the reliability of scientific findings in various fields.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100010Description: With the exception of two-phase sampling, the standard variance approximation of the generalized regression (GREG) estimator assumes that the population totals in the weighting scheme are observed without error. If the weighting model of the GREG estimator contains population totals that are observed with measurement error sources other than the sampling error of first-phase estimates, then this uncertainty will be ignored by the variance approximation of the GREG estimator. This paper proposes a variance approximation for the GREG estimator that accounts for additional uncertainty arising from measurement error in one or more of the population totals used in the weighting scheme. This approach has been developed for, and is being applied to, the Dutch Labour Force Survey (DLFS). The monthly publications of the DLFS are obtained with a time series model, which corrects for rotation group bias and discontinuities caused by major redesigns and the loss of face-to-face interviews during COVID-19. The GREG estimates for the quarterly figures are benchmarked to the average of the monthly publications to enforce numerical consistency between monthly and quarterly publication tables. The standard variance approximation of the GREG estimator assumes that these population totals are observed without error. This results in an underestimation of the variance of the GREG estimator. The variance approximation proposed in this paper results in more realistic standard errors for the quarterly GREG estimates.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100011Description: We construct a hybrid Bayesian method, which includes a differentially private mechanism, to mask Census county totals for a U.S. state on acreage of a commodity. We use surrogates for data collected at the farm level from a past U.S. Census of Agriculture to illustrate our procedure. We use two Bayesian small area models (parametric and mixture) to accommodate the smaller counties with fewer farms and some counties with large acres. In these models, the Laplace distribution provides a differentially private mechanism. In pre-processing, we also incorporate the Census weights to form the observed total acreage, a scaling factor to the Laplace mechanism for each county, a square-root transformation of the observed total acreage to avoid negative masked estimates especially for small counties, and the p-percent rule and the 3+ rule to partition the counties into suppressed counties, non-sensitive counties and sensitive counties. Because of difficulties in specifying and tuning the privacy budget (an unknown parameter), to balance security and utility, we specify a prior for the privacy budget, where the values are not specified, and the Gibbs sampler is used to fit the hierarchical Bayesian models. In post-processing, we use Bayesian predictive inference to obtain masked county acreages, and this includes a benchmarking so that the masked state total matches the observed state total. As a measure of reliability of the Bayesian procedure, we use the posterior coefficients of variation for the masked posterior means of the counties. As a measure of utility, we use the absolute relative errors for the individual counties, together with other global measures. For the sensitive counties, there are some differences between the two small area models but both are much better than an individual area model; the mixture model being the best compromise for security and utility.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100012Description: We propose small area estimators of general indicators in off-census years, which avoid the use of deprecated census microdata, but are nearly optimal in census years. The procedure is based on replacing the obsolete census file with a larger unit-level survey that adequately covers the areas of interest and contains the values of useful auxiliary variables. However, the minimal data requirement of the proposed method is a single survey with microdata on the target variable and suitable auxiliary variables for the period of interest. We also develop an estimator of the mean squared error (MSE) that accounts for the uncertainty introduced by the large survey used to replace the census of auxiliary information. Our empirical results indicate that the proposed predictors perform clearly better than the alternative predictors when census data are outdated, and are very close to optimal ones when census data are correct. They also illustrate that the proposed total MSE estimator corrects for the bias of purely model-based MSE estimators that do not account for the large survey uncertainty.Release date: 2026-06-29
- Surveys and statistical programs – Documentation: 11-633-X2026002Description: Recent changes in Canada’s immigration levels have heightened interest in understanding how immigration affects housing demand. This article develops a methodological framework for projecting housing use associated with permanent residents (PRs) and non-permanent residents (NPRs) under alternative immigration scenarios. The framework applies observed per capita housing use rates from the Census of Population to estimate incremental housing use by tenure over time.Release date: 2026-04-24
- Articles and reports: 12-001-X202500200001Description: Nested error regression models are commonly used to incorporate unit specific auxiliary variables to improve small area estimates. When the mean structure of the model is misspecified, the design-based mean squared prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP) generally increases. The Observed Best Prediction (OBP) method has been proposed with the intent to improve on the design-based MSPE over EBLUP. In this paper, we conduct a Monte Carlo simulation experiments to understand the effect of misspsecification of mean structures on different small area estimators. Our findings suggest that the OBP using unit-level auxiliary variables does not outperform the EBLUP in terms of design-based MSPE, unless the number of small areas m is extremely large. Conversely, the performance of OBP significantly improves when area-level auxiliary variables are employed. This paper includes both analytical and numerical evidence to demonstrate these observations, providing practical insights for addressing model misspecification in small area estimation (SAE).Release date: 2025-12-23
- Articles and reports: 12-001-X202500200003Description: In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.Release date: 2025-12-23
- Articles and reports: 12-001-X202500200005Description: The use of non-probability data sources for statistical purposes and for official statistics has become increasingly popular in recent years. However, statistical inference based on non-probability samples is made more difficult by nature of their biasedness and lack of representativity. In this paper we propose quantile balancing inverse probability weighting estimator (QBIPW) for non-probability samples. We apply the idea of Harms and Duchesne (2006) allowing the use of quantile information in the estimation process to reproduce known totals and the distribution of auxiliary variables. We discuss the estimation of the QBIPW probabilities and its variance. Our simulation study has demonstrated that the proposed estimators are robust against model mis-specification and, as a result, help to reduce bias and mean squared error. Finally, we applied the proposed methods to estimate the share of job vacancies aimed at Ukrainian workers in Poland using an integrated set of administrative and survey data about job vacancies.Release date: 2025-12-23
- Previous Go to previous page of All results
- 1 (current) Go to page 1 of All results
- 2 Go to page 2 of All results
- 3 Go to page 3 of All results
- 4 Go to page 4 of All results
- 5 Go to page 5 of All results
- 6 Go to page 6 of All results
- 7 Go to page 7 of All results
- 8 Go to page 8 of All results
- Next Go to next page of All results
Data (0)
Data (0) (0 results)
No content available at this time.
Analysis (68)
Analysis (68) (50 to 60 of 68 results)
- Articles and reports: 12-001-X202300100004Description: The Dutch Health Survey (DHS), conducted by Statistics Netherlands, is designed to produce reliable direct estimates at an annual frequency. Data collection is based on a combination of web interviewing and face-to-face interviewing. Due to lockdown measures during the Covid-19 pandemic there was no or less face-to-face interviewing possible, which resulted in a sudden change in measurement and selection effects in the survey outcomes. Furthermore, the production of annual data about the effect of Covid-19 on health-related themes with a delay of about one year compromises the relevance of the survey. The sample size of the DHS does not allow the production of figures for shorter reference periods. Both issues are solved by developing a bivariate structural time series model (STM) to estimate quarterly figures for eight key health indicators. This model combines two series of direct estimates, a series based on complete response and a series based on web response only and provides model-based predictions for the indicators that are corrected for the loss of face-to-face interviews during the lockdown periods. The model is also used as a form of small area estimation and borrows sample information observed in previous reference periods. In this way timely and relevant statistics describing the effects of the corona crisis on the development of Dutch health are published. In this paper the method based on the bivariate STM is compared with two alternative methods. The first one uses a univariate STM where no correction for the lack of face-to-face observation is applied to the estimates. The second one uses a univariate STM that also contains an intervention variable that models the effect of the loss of face-to-face response during the lockdown.Release date: 2023-06-30
- Articles and reports: 12-001-X202300100005Description: Weight smoothing is a useful technique in improving the efficiency of design-based estimators at the risk of bias due to model misspecification. As an extension of the work of Kim and Skinner (2013), we propose using weight smoothing to construct the conditional likelihood for efficient analytic inference under informative sampling. The Beta prime distribution can be used to build a parameter model for weights in the sample. A score test is developed to test for model misspecification in the weight model. A pretest estimator using the score test can be developed naturally. The pretest estimator is nearly unbiased and can be more efficient than the design-based estimator when the weight model is correctly specified, or the original weights are highly variable. A limited simulation study is presented to investigate the performance of the proposed methods.Release date: 2023-06-30
- Articles and reports: 12-001-X202300100011Description: The definition of statistical units is a recurring issue in the domain of sample surveys. Indeed, not all the populations surveyed have a readily available sampling frame. For some populations, the sampled units are distinct from the observation units and producing estimates on the population of interest raises complex questions, which can be addressed by using the weight share method (Deville and Lavallée, 2006). However, the two populations considered in this approach are discrete. In some fields of study, the sampled population is continuous: this is for example the case of forest inventories for which, frequently, the trees surveyed are those located on plots of which the centers are points randomly drawn in a given area. The production of statistical estimates from the sample of trees surveyed poses methodological difficulties, as do the associated variance calculations. The purpose of this paper is to generalize the weight share method to the continuous (sampled population) ? discrete (surveyed population) case, from the extension proposed by Cordy (1993) of the Horvitz-Thompson estimator for drawing points carried out in a continuous universe.Release date: 2023-06-30
- Articles and reports: 12-001-X202200200010Description:
Multilevel time series (MTS) models are applied to estimate trends in time series of antenatal care coverage at several administrative levels in Bangladesh, based on repeated editions of the Bangladesh Demographic and Health Survey (BDHS) within the period 1994-2014. MTS models are expressed in an hierarchical Bayesian framework and fitted using Markov Chain Monte Carlo simulations. The models account for varying time lags of three or four years between the editions of the BDHS and provide predictions for the intervening years as well. It is proposed to apply cross-sectional Fay-Herriot models to the survey years separately at district level, which is the most detailed regional level. Time series of these small domain predictions at the district level and their variance-covariance matrices are used as input series for the MTS models. Spatial correlations among districts, random intercept and slope at the district level, and different trend models at district level and higher regional levels are examined in the MTS models to borrow strength over time and space. Trend estimates at district level are obtained directly from the model outputs, while trend estimates at higher regional and national levels are obtained by aggregation of the district level predictions, resulting in a numerically consistent set of trend estimates.
Release date: 2022-12-15 - Articles and reports: 12-001-X202200200011Description:
Two-phase sampling is a cost effective sampling design employed extensively in surveys. In this paper a method of most efficient linear estimation of totals in two-phase sampling is proposed, which exploits optimally auxiliary survey information. First, a best linear unbiased estimator (BLUE) of any total is formally derived in analytic form, and shown to be also a calibration estimator. Then, a proper reformulation of such a BLUE and estimation of its unknown coefficients leads to the construction of an “optimal” regression estimator, which can also be obtained through a suitable calibration procedure. A distinctive feature of such calibration is the alignment of estimates from the two phases in an one-step procedure involving the combined first-and-second phase samples. Optimal estimation is feasible for certain two-phase designs that are used often in large scale surveys. For general two-phase designs, an alternative calibration procedure gives a generalized regression estimator as an approximate optimal estimator. The proposed general approach to optimal estimation leads to the most effective use of the available auxiliary information in any two-phase survey. The advantages of this approach over existing methods of estimation in two-phase sampling are shown both theoretically and through a simulation study.
Release date: 2022-12-15 - Articles and reports: 12-001-X202200200012Description:
In many applications, the population means of geographically adjacent small areas exhibit a spatial variation. If available auxiliary variables do not adequately account for the spatial pattern, the residual variation will be included in the random effects. As a result, the independent and identical distribution assumption on random effects of the Fay-Herriot model will fail. Furthermore, limited resources often prevent numerous sub-populations from being included in the sample, resulting in non-sampled small areas. The problem can be exacerbated for predicting means of non-sampled small areas using the above Fay-Herriot model as the predictions will be made based solely on the auxiliary variables. To address such inadequacy, we consider Bayesian spatial random-effect models that can accommodate multiple non-sampled areas. Under mild conditions, we establish the propriety of the posterior distributions for various spatial models for a useful class of improper prior densities on model parameters. The effectiveness of these spatial models is assessed based on simulated and real data. Specifically, we examine predictions of statewide four-person family median incomes based on the 1990 Current Population Survey and the 1980 Census for the United States of America.
Release date: 2022-12-15 - Articles and reports: 75F0002M2022006Description:
This technical paper describes how the cost for "other necessities" is estimated in the 2018-base MBM. It provides a brief overview of the theory and application of techniques for estimating costs of "other necessities" in poverty lines and deconstructs the 2018-base MBM other necessities component to provide insights on how it is constructed. The aim of this paper is to provide a more detailed understanding of how the other necessities component of the MBM is estimated.
Release date: 2022-12-08 - Articles and reports: 89-648-X2022001Description:
This report explores the size and nature of the attrition challenges faced by the Longitudinal and International Study of Adults (LISA) survey, as well as the use of a non-response weight adjustment and calibration strategy to mitigate the effects of attrition on the LISA estimates. The study focuses on data from waves 1 (2012) to 4 (2018) and uses practical examples based on selected demographic variables, to illustrate how attrition be assessed and treated.
Release date: 2022-11-14 - Articles and reports: 12-001-X202200100002Description: We consider an intercept only linear random effects model for analysis of data from a two stage cluster sampling design. At the first stage a simple random sample of clusters is drawn, and at the second stage a simple random sample of elementary units is taken within each selected cluster. The response variable is assumed to consist of a cluster-level random effect plus an independent error term with known variance. The objects of inference are the mean of the outcome variable and the random effect variance. With a more complex two stage sampling design, the use of an approach based on an estimated pairwise composite likelihood function has appealing properties. Our purpose is to use our simpler context to compare the results of likelihood inference with inference based on a pairwise composite likelihood function that is treated as an approximate likelihood, in particular treated as the likelihood component in Bayesian inference. In order to provide credible intervals having frequentist coverage close to nominal values, the pairwise composite likelihood function and corresponding posterior density need modification, such as a curvature adjustment. Through simulation studies, we investigate the performance of an adjustment proposed in the literature, and find that it works well for the mean but provides credible intervals for the random effect variance that suffer from under-coverage. We propose possible future directions including extensions to the case of a complex design.Release date: 2022-06-21
- Articles and reports: 12-001-X202200100003Description:
Use of auxiliary data to improve the efficiency of estimators of totals and means through model-assisted survey regression estimation has received considerable attention in recent years. Generalized regression (GREG) estimators, based on a working linear regression model, are currently used in establishment surveys at Statistics Canada and several other statistical agencies. GREG estimators use common survey weights for all study variables and calibrate to known population totals of auxiliary variables. Increasingly, many auxiliary variables are available, some of which may be extraneous. This leads to unstable GREG weights when all the available auxiliary variables, including interactions among categorical variables, are used in the working linear regression model. On the other hand, new machine learning methods, such as regression trees and lasso, automatically select significant auxiliary variables and lead to stable nonnegative weights and possible efficiency gains over GREG. In this paper, a simulation study, based on a real business survey sample data set treated as the target population, is conducted to study the relative performance of GREG, regression trees and lasso in terms of efficiency of the estimators and properties of associated regression weights. Both probability sampling and non-probability sampling scenarios are studied.
Release date: 2022-06-21
- Previous Go to previous page of Analysis results
- 1 Go to page 1 of Analysis results
- 2 Go to page 2 of Analysis results
- 3 Go to page 3 of Analysis results
- 4 Go to page 4 of Analysis results
- 5 Go to page 5 of Analysis results
- 6 (current) Go to page 6 of Analysis results
- 7 Go to page 7 of Analysis results
- Next Go to next page of Analysis results
Reference (6)
Reference (6) ((6 results))
- Surveys and statistical programs – Documentation: 11-633-X2026002Description: Recent changes in Canada’s immigration levels have heightened interest in understanding how immigration affects housing demand. This article develops a methodological framework for projecting housing use associated with permanent residents (PRs) and non-permanent residents (NPRs) under alternative immigration scenarios. The framework applies observed per capita housing use rates from the Census of Population to estimate incremental housing use by tenure over time.Release date: 2026-04-24
- Surveys and statistical programs – Documentation: 91-528-XDescription: The Technical Guide on Demographic Estimates at Statistics Canada provides detailed descriptions of the most current data sources and methods used by the Centre for demography at Statistics Canada to produce demographic estimates as part of the Demographic estimates program. They comprise postcensal and intercensal population estimates; base population; births and deaths; immigrants; emigrants; returning emigrants; non-permanent residents; interprovincial migration; subprovincial estimates of population and intraprovincial migration; population estimates by age and gender; and census family estimates. A glossary of commonly used terms is available at the end of the guide.Release date: 2025-12-17
- Surveys and statistical programs – Documentation: 98-306-XDescription:
This report describes sampling, weighting and estimation procedures used in the Census of Population. It provides operational and theoretical justifications for them, and presents the results of the evaluations of these procedures.
Release date: 2023-10-04 - Notices and consultations: 75F0002M2019006Description:
In 2018, Statistics Canada released two new data tables with estimates of effective tax and transfer rates for individual tax filers and census families. These estimates are derived from the Longitudinal Administrative Databank. This publication provides a detailed description of the methods used to derive the estimates of effective tax and transfer rates.
Release date: 2019-04-16 - Surveys and statistical programs – Documentation: 99-002-XDescription: This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.Release date: 2015-01-28
- Surveys and statistical programs – Documentation: 92-568-XDescription:
This report describes sampling and weighting procedures used in the 2006 Census. It reviews the history of these procedures in Canadian censuses, provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.
Release date: 2009-08-11