Weighting and estimation

Skip to filters. View results.

Sort Help
entries

Results

All (638)

All (638) (0 to 10 of 638 results)

  • Articles and reports: 12-001-X202600100003
    Description: Probability-proportional-to-size sampling is widely used by national statistical offices. Here population units are selected with probabilities proportional to an auxiliary variable. Variance formulas in such designs require both first- and second-order inclusion probabilities. The computation of second-order inclusion probabilities is particularly challenging for large populations, and has been the subject of extensive research. This article presents some new exact and approximation formulas for second-order inclusion probabilities in randomized systematic sampling with unequal probabilities and without replacement.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100005
    Description: Confidence intervals are very often constructed based on a probability distribution that uses a certain number of degrees of freedom as a parameter. This is the case with the Student and the modified Wilson confidence intervals, discussed in this article, which use quantiles from the Student distribution where the number of degrees of freedom is generally unknown. For the length of a confidence interval to be representative of the reliability of an estimate, the actual coverage rate must match the nominal rate. To that end, the number of degrees of freedom in the probability distribution used in practice to calculate the confidence interval must be estimated as precisely as possible. An approximate rule is often used, although it tends to overestimate the actual number of degrees of freedom. In this article, a more precise version of degrees of freedom, derived from the Satterthwaite approximation, is obtained in the context of the Canadian Census of Population. The sampling design is equivalent to a simple random design without replacement, cluster-stratified, and the variance estimation method is an adaptation of the balanced repeated replication method. An explicit expression of the degrees of freedom is obtained under these conditions, enabling the factors influencing them to be identified. For comparison, the degree of freedom formula is also established for the conventional variance estimator. A simulation study shows that using this version of degrees of freedom corrects the undercoverage problem observed with the approximate rule, showing the importance of accurately assessing this number.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100009
    Description: Combining estimates from independent surveys via inverse-variance weights can lead to negative bias when unknown variances are estimated and the target variable is non-negative and positively skewed. In such cases, strong positive correlations typically arise between the estimators and their corresponding variance estimators, causing standard linear combinations with inverse-variance weights to exhibit negative bias. We introduce a strikingly simple method to reduce bias: replace the standard weight with the ratio of the estimator to the variance estimator. Under a linear model linking the two, we show that the new ratio-weighted estimator is approximately unbiased, whereas the conventional inverse-variance combination exhibits downward bias. Through simulations, we demonstrate that the new method brings both the bias and the mean squared error closer to the optimum for a wide range of different target variables. As our method uses only standardly reported summary statistics, it can be immediately adopted to reduce this widespread bias and improve the reliability of scientific findings in various fields.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100010
    Description: With the exception of two-phase sampling, the standard variance approximation of the generalized regression (GREG) estimator assumes that the population totals in the weighting scheme are observed without error. If the weighting model of the GREG estimator contains population totals that are observed with measurement error sources other than the sampling error of first-phase estimates, then this uncertainty will be ignored by the variance approximation of the GREG estimator. This paper proposes a variance approximation for the GREG estimator that accounts for additional uncertainty arising from measurement error in one or more of the population totals used in the weighting scheme. This approach has been developed for, and is being applied to, the Dutch Labour Force Survey (DLFS). The monthly publications of the DLFS are obtained with a time series model, which corrects for rotation group bias and discontinuities caused by major redesigns and the loss of face-to-face interviews during COVID-19. The GREG estimates for the quarterly figures are benchmarked to the average of the monthly publications to enforce numerical consistency between monthly and quarterly publication tables. The standard variance approximation of the GREG estimator assumes that these population totals are observed without error. This results in an underestimation of the variance of the GREG estimator. The variance approximation proposed in this paper results in more realistic standard errors for the quarterly GREG estimates.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100011
    Description: We construct a hybrid Bayesian method, which includes a differentially private mechanism, to mask Census county totals for a U.S. state on acreage of a commodity. We use surrogates for data collected at the farm level from a past U.S. Census of Agriculture to illustrate our procedure. We use two Bayesian small area models (parametric and mixture) to accommodate the smaller counties with fewer farms and some counties with large acres. In these models, the Laplace distribution provides a differentially private mechanism. In pre-processing, we also incorporate the Census weights to form the observed total acreage, a scaling factor to the Laplace mechanism for each county, a square-root transformation of the observed total acreage to avoid negative masked estimates especially for small counties, and the p-percent rule and the 3+ rule to partition the counties into suppressed counties, non-sensitive counties and sensitive counties. Because of difficulties in specifying and tuning the privacy budget (an unknown parameter), to balance security and utility, we specify a prior for the privacy budget, where the values are not specified, and the Gibbs sampler is used to fit the hierarchical Bayesian models. In post-processing, we use Bayesian predictive inference to obtain masked county acreages, and this includes a benchmarking so that the masked state total matches the observed state total. As a measure of reliability of the Bayesian procedure, we use the posterior coefficients of variation for the masked posterior means of the counties. As a measure of utility, we use the absolute relative errors for the individual counties, together with other global measures. For the sensitive counties, there are some differences between the two small area models but both are much better than an individual area model; the mixture model being the best compromise for security and utility.
    Release date: 2026-06-29

  • Articles and reports: 12-001-X202600100012
    Description: We propose small area estimators of general indicators in off-census years, which avoid the use of deprecated census microdata, but are nearly optimal in census years. The procedure is based on replacing the obsolete census file with a larger unit-level survey that adequately covers the areas of interest and contains the values of useful auxiliary variables. However, the minimal data requirement of the proposed method is a single survey with microdata on the target variable and suitable auxiliary variables for the period of interest. We also develop an estimator of the mean squared error (MSE) that accounts for the uncertainty introduced by the large survey used to replace the census of auxiliary information. Our empirical results indicate that the proposed predictors perform clearly better than the alternative predictors when census data are outdated, and are very close to optimal ones when census data are correct. They also illustrate that the proposed total MSE estimator corrects for the bias of purely model-based MSE estimators that do not account for the large survey uncertainty.
    Release date: 2026-06-29

  • Surveys and statistical programs – Documentation: 11-633-X2026002
    Description: Recent changes in Canada’s immigration levels have heightened interest in understanding how immigration affects housing demand. This article develops a methodological framework for projecting housing use associated with permanent residents (PRs) and non-permanent residents (NPRs) under alternative immigration scenarios. The framework applies observed per capita housing use rates from the Census of Population to estimate incremental housing use by tenure over time.
    Release date: 2026-04-24

  • Articles and reports: 12-001-X202500200001
    Description: Nested error regression models are commonly used to incorporate unit specific auxiliary variables to improve small area estimates. When the mean structure of the model is misspecified, the design-based mean squared prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP) generally increases. The Observed Best Prediction (OBP) method has been proposed with the intent to improve on the design-based MSPE over EBLUP. In this paper, we conduct a Monte Carlo simulation experiments to understand the effect of misspsecification of mean structures on different small area estimators. Our findings suggest that the OBP using unit-level auxiliary variables does not outperform the EBLUP in terms of design-based MSPE, unless the number of small areas m is extremely large. Conversely, the performance of OBP significantly improves when area-level auxiliary variables are employed. This paper includes both analytical and numerical evidence to demonstrate these observations, providing practical insights for addressing model misspecification in small area estimation (SAE).
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200003
    Description: In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200005
    Description: The use of non-probability data sources for statistical purposes and for official statistics has become increasingly popular in recent years. However, statistical inference based on non-probability samples is made more difficult by nature of their biasedness and lack of representativity. In this paper we propose quantile balancing inverse probability weighting estimator (QBIPW) for non-probability samples. We apply the idea of Harms and Duchesne (2006) allowing the use of quantile information in the estimation process to reproduce known totals and the distribution of auxiliary variables. We discuss the estimation of the QBIPW probabilities and its variance. Our simulation study has demonstrated that the proposed estimators are robust against model mis-specification and, as a result, help to reduce bias and mean squared error. Finally, we applied the proposed methods to estimate the share of job vacancies aimed at Ukrainian workers in Poland using an integrated set of administrative and survey data about job vacancies.
    Release date: 2025-12-23
Data (0)

Data (0) (0 results)

No content available at this time.

Analysis (610)

Analysis (610) (50 to 60 of 610 results)

  • Articles and reports: 11-522-X202200100015
    Description: We present design-based Horvitz-Thompson and multiplicity estimators of the population size, as well as of the total and mean of a response variable associated with the elements of a hidden population to be used with the link-tracing sampling variant proposed by Félix-Medina and Thompson (2004). Since the computation of the estimators requires to know the inclusion probabilities of the sampled people, but they are unknown, we propose a Bayesian model which allows us to estimate them, and consequently to compute the estimators of the population parameters. The results of a small numeric study indicate that the performance of the proposed estimators is acceptable.
    Release date: 2024-03-25

  • Articles and reports: 11-522-X202200100018
    Description: The Longitudinal Social Data Development Program (LSDDP) is a social data integration approach aimed at providing longitudinal analytical opportunities without imposing additional burden on respondents. The LSDDP uses a multitude of signals from different data sources for the same individual, which helps to better understand their interactions and track changes over time. This article looks at how the ethnicity status of people in Canada can be estimated at the most detailed disaggregated level possible using the results from a variety of business rules applied to linked data and to the LSDDP denominator. It will then show how improvements were obtained using machine learning methods, such as decision trees and random forest techniques.
    Release date: 2024-03-25

  • Articles and reports: 12-001-X202300200002
    Description: Being able to quantify the accuracy (bias, variance) of published output is crucial in official statistics. Output in official statistics is nearly always divided into subpopulations according to some classification variable, such as mean income by categories of educational level. Such output is also referred to as domain statistics. In the current paper, we limit ourselves to binary classification variables. In practice, misclassifications occur and these contribute to the bias and variance of domain statistics. Existing analytical and numerical methods to estimate this effect have two disadvantages. The first disadvantage is that they require that the misclassification probabilities are known beforehand and the second is that the bias and variance estimates are biased themselves. In the current paper we present a new method, a Gaussian mixture model estimated by an Expectation-Maximisation (EM) algorithm combined with a bootstrap, referred to as the EM bootstrap method. This new method does not require that the misclassification probabilities are known beforehand, although it is more efficient when a small audit sample is used that yields a starting value for the misclassification probabilities in the EM algorithm. We compared the performance of the new method with currently available numerical methods: the bootstrap method and the SIMEX method. Previous research has shown that for non-linear parameters the bootstrap outperforms the analytical expressions. For nearly all conditions tested, the bias and variance estimates that are obtained by the EM bootstrap method are closer to their true values than those obtained by the bootstrap and SIMEX methods. We end this paper by discussing the results and possible future extensions of the method.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200003
    Description: We investigate small area prediction of general parameters based on two models for unit-level counts. We construct predictors of parameters, such as quartiles, that may be nonlinear functions of the model response variable. We first develop a procedure to construct empirical best predictors and mean square error estimators of general parameters under a unit-level gamma-Poisson model. We then use a sampling importance resampling algorithm to develop predictors for a generalized linear mixed model (GLMM) with a Poisson response distribution. We compare the two models through simulation and an analysis of data from the Iowa Seat-Belt Use Survey.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200004
    Description: We present a novel methodology to benchmark county-level estimates of crop area totals to a preset state total subject to inequality constraints and random variances in the Fay-Herriot model. For planted area of the National Agricultural Statistics Service (NASS), an agency of the United States Department of Agriculture (USDA), it is necessary to incorporate the constraint that the estimated totals, derived from survey and other auxiliary data, are no smaller than administrative planted area totals prerecorded by other USDA agencies except NASS. These administrative totals are treated as fixed and known, and this additional coherence requirement adds to the complexity of benchmarking the county-level estimates. A fully Bayesian analysis of the Fay-Herriot model offers an appealing way to incorporate the inequality and benchmarking constraints, and to quantify the resulting uncertainties, but sampling from the posterior densities involves difficult integration, and reasonable approximations must be made. First, we describe a single-shrinkage model, shrinking the means while the variances are assumed known. Second, we extend this model to accommodate double shrinkage, borrowing strength across means and variances. This extended model has two sources of extra variation, but because we are shrinking both means and variances, it is expected that this second model should perform better in terms of goodness of fit (reliability) and possibly precision. The computations are challenging for both models, which are applied to simulated data sets with properties resembling the Illinois corn crop.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200012
    Description: In recent decades, many different uses of auxiliary information have enriched survey sampling theory and practice. Jean-Claude Deville contributed significantly to this progress. My comments trace some of the steps on the way to one important theory for the use of auxiliary information: Estimation by calibration.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200013
    Description: Jean-Claude Deville is one of the most prominent researcher in survey sampling theory and practice. His research on balanced sampling, indirect sampling and calibration in particular is internationally recognized and widely used in official statistics. He was also a pioneer in the field of functional data analysis. This discussion gives us the opportunity to recognize the immense work he has accomplished, and to pay tribute to him. In the first part of this article, we recall briefly his contribution to the functional principal analysis. We also detail some recent extension of his work at the intersection of the fields of functional data analysis and survey sampling. In the second part of this paper, we present some extension of Jean-Claude’s work in indirect sampling. These extensions are motivated by concrete applications and illustrate Jean-Claude’s influence on our work as researchers.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200014
    Description: Many things have been written about Jean-Claude Deville in tributes from the statistical community (see Tillé, 2022a; Tillé, 2022b; Christine, 2022; Ardilly, 2022; and Matei, 2022) and from the École nationale de la statistique et de l’administration économique (ENSAE) and the Société française de statistique. Pascal Ardilly, David Haziza, Pierre Lavallée and Yves Tillé provide an in-depth look at Jean-Claude Deville’s contributions to survey theory. To pay tribute to him, I would like to discuss Jean-Claude Deville’s contribution to the more day-to-day application of methodology for all the statisticians at the Institut national de la statistique et des études économiques (INSEE) and at the public statistics service. To do this, I will use my work experience, and particularly the four years (1992 to 1996) I spent working with him in the Statistical Methods Unit and the discussions we had thereafter, especially in the 2000s on the rolling census.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200015
    Description: This article discusses and provides comments on the Ardilly, Haziza, Lavallée and Tillé’s summary presentation of Jean-Claude Deville’s work on survey theory. It sheds light on the context, applications and uses of his findings, and shows how these have become engrained in the role of statisticians, in which Jean-Claude was a trailblazer. It also discusses other aspects of his career and his creative inventions.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X202300200016
    Description: In this discussion, I will present some additional aspects of three major areas of survey theory developed or studied by Jean-Claude Deville: calibration, balanced sampling and the generalized weight-share method.
    Release date: 2024-01-03
Reference (28)

Reference (28) (20 to 30 of 28 results)

  • Surveys and statistical programs – Documentation: 92-371-X
    Description:

    This report deals with sampling and weighting, a process whereby certain characteristics are collected and processed for a random sample of dwellings and persons identified in the complete census enumeration. Data for the whole population are then obtained by scaling up the results for the sample to the full population level. The use of sampling may lead to substantial reductions in costs and respondent burden, or alternatively, can allow the scope of a census to be broadened at the same cost.

    Release date: 1999-12-07

  • Surveys and statistical programs – Documentation: 11-522-X19980015017
    Description:

    Longitudinal studies with repeated observations on individuals permit better characterizations of change and assessment of possible risk factors, but there has been little experience applying sophisticated models for longitudinal data to the complex survey setting. We present results from a comparison of different variance estimation methods for random effects models of change in cognitive function among older adults. The sample design is a stratified sample of people 65 and older, drawn as part of a community-based study designed to examine risk factors for dementia. The model summarizes the population heterogeneity in overall level and rate of change in cognitive function using random effects for intercept and slope. We discuss an unweighted regression including covariates for the stratification variables, a weighted regression, and bootstrapping; we also did preliminary work into using balanced repeated replication and jackknife repeated replication.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015019
    Description:

    The British Labour Force Survey (LFS) is a quarterly household survey with a rotating sample design that can potentially be used to produce longitudinal data, including estimates of labour force gross flows. However, these estimates may be biased due to the effect of non-response. Weighting adjustments are a commonly used method to account for non-response bias. We find that weighting may not fully account for the effect of non-response bias because non-response may depend on the unobserved labour force flows, i.e., the non-response is non-ignorable. To adjust for the effects of non-ignorable non-response, we propose a model for the complex non-response patterns in the LFS which controls for the correlated within-household non-response behaviour found in the survey. The results of modelling suggest that non-response may be non-ignorable in the LFS, causing the weighting estimates to be biased.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015020
    Description:

    At the end of 1993, Eurostat lauched a 'community' panel of households. The first wave, carried out in 1994 in the 12 countries of the European Union, included some 7,300 households in France, and at least 14,000 adults 17 years or over. Each individual was then followed up and interviewed each year, even if they had moved. The individuals leaving the sample present a particular profile. In the first part, we present a sketch of how our sample evolves and an analysis of the main characteristics of the non-respondents. We then propose 2 models to correct for non-response per homogeneous category. We then describe the longitudinal weight distribution obtained from the two models, and the cross-sectional weights using the weight share method. Finally, we compare some indicators calculated using both weighting methods.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015023
    Description:

    The study of social mobility, between labour market statuses or between income levels, for example, is often based on the analysis of mobility matrices. When comparing these transition matrices, with a view to evaluating behavioural changes, one often forgets that the data derive from a sample survey and are therefore affected by sampling variances. Similarly, it is assumed that the responses collected correspond to the ' true value.'

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015026
    Description:

    The purpose of the present study is to utilize panel data from the Current Population Survey (CPS) to examine the effects of unit nonresponse. Because most nonrespondents to the CPS are respondents during at least one month-in-sample, data from other months can be used to compare the characteristics of complete respondents and panel nonrespondents and to evaluate nonresponse adjustment procedures. In the current paper we present analyses utilizing CPS panel data to illustrate the effects of unit nonresponse. After adjusting for nonresponse, additional comparisons are also made to evaluate the effects of nonresponse adjustment. The implications of the findings and suggestions for further research are discussed.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015028
    Description:

    We address the problem of estimation for the income dynamics statistics calculated from complex longitudinal surveys. In addition, we compare two design-based estimators of longitudinal proportions and transition rates in terms of variability under large attrition rates. One estimator is based on the cross-sectional samples for the estimation of the income class boundaries at each time period and on the longitudinal sample for the estimation of the longitudinal counts; the other estimator is entirely based on the longitudinal sample, both for the estimation of the class boundaries and the longitudinal counts. We develop Taylor linearization-type variance estimators for both the longitudinal and the mixed estimator under the assumption of no change in the population, and for the mixed estimator when there is change.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015031
    Description:

    The U.S. Third National Health and Nutrition Examination Survey (NHANES III) was carried out from 1988 to 1994. This survey was intended primarily to provide estimates of cross-sectional parameters believed to be approximately constant over the six-year data collection period. However, for some variable (e.g., serum lead, body mass index and smoking behavior), substantive considerations suggest the possible presence of nontrivial changes in level between 1988 and 1994. For these variables, NHANES III is potentially a valuable source of time-change information, compared to other studies involving more restricted populations and samples. Exploration of possible change over time is complicated by two issues. First, there was of practical concern because some variables displayed substantial regional differences in level. This was of practical concern because some variables displayed substantial regional differences in level. Second, nontrivial changes in level over time can lead to nontrivial biases in some customary NHANES III variance estimators. This paper considers these two problems and discusses some related implications for statistical policy.

    Release date: 1999-10-22