Weighting and estimation

Skip to filters. View results.

Sort Help
entries

Results

All (638)

All (638) (20 to 30 of 638 results)

  • Articles and reports: 12-001-X202500100003
    Description: In recent years, there has been a significant interest in machine learning in national statistical offices. Thanks to their flexibility, these methods may prove useful at the nonresponse treatment stage. In this article, we conduct an empirical investigation in order to compare several machine learning procedures in terms of bias and efficiency. In addition to the classical machine learning procedures, we assess the performance of ensemble approaches that make use of different machine learning procedures to produce a set of weights adjusted for nonresponse.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100005
    Description: In this paper, we derive a second-order unbiased (or nearly unbiased) mean squared prediction error (MSPE) estimator of the empirical best linear unbiased predictor (EBLUP) of a small area mean for a semi-parametric extension to the well-known Fay-Herriot model. Specifically, we derive our MSPE estimator essentially assuming certain moment conditions on both the sampling errors and random effects distributions. The normality-based Prasad-Rao MSPE estimator has a surprising robustness property in that it remains second-order unbiased under the non-normality of random effects when a simple Prasad-Rao method-of-moments estimator is used for the variance component and the sampling error distribution is normal. We show that the normality-based MSPE estimator is no longer second-order unbiased when the sampling error distribution has non-zero kurtosis or when the Fay-Herriot moment method is used to estimate the variance component, even when the sampling error distribution is normal. Interestingly, when the simple method-of moments estimator is used for the variance component, our proposed MSPE estimator does not require the estimation of kurtosis of the random effects. Results of a simulation study on the accuracy of the proposed MSPE estimator, under non-normality of both sampling and random effects distributions, are also presented.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100006
    Description: Survey practitioners have increasingly embraced the benefits of modern machine learning techniques, including classification and regression tree algorithms, in the development of nonresponse adjustments. These methods, which do not require a predefined functional relationship between outcomes and predictors, offer a practical means of conducting variable selection and deriving interpretable structures that link response propensity with explanatory variables. However, when applying these algorithms to survey data, it is common to overlook crucial factors like sampling weights, as well as sample design features such as stratification and clustering. To bridge this shortcoming, we propose an extension of the Chi-square Automatic Interaction Detector (CHAID) approach, and we describe the design-based asymptotic properties of the resulting “survey CHAID” (sCHAID) method. To facilitate the practical use of sCHAID, we incorporate a Rao-Scott correction into the splitting criterion, accounting for the survey design. Using data from the U.S. American Community Survey, we illustrate the use of the method and evaluate its performance through comparisons with existing weighted and unweighted algorithms.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100007
    Description: We introduce a novel approach to model-assisted calibration estimation in survey sampling using generalized entropy. The method builds upon recent work by Kwon, Kim and Qiu (2024) and extends it to a model-assisted framework. Unlike traditional calibration techniques, this approach employs a generalized entropy function as the objective for optimization and incorporates a debiasing calibration constraint to ensure design consistency. The proposed estimator is shown to be asymptotically equivalent to an augmented generalized regression (GREG) estimator. It allows for unequal model variance, potentially improving efficiency when the sampling design is informative. The paper presents both design-based and model-based justifications for the method, along with asymptotic properties and variance estimation techniques. Computational aspects are discussed, including an unconstrained optimization approach that facilitates implementation, especially for high-dimensional auxiliary variables. The method’s performance is evaluated through a simulation study, demonstrating its effectiveness in improving estimation efficiency, particularly when the sampling design is informative.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100008
    Description: Tightened budgets, continuing decrease of response rates in traditional probability surveys and increasing pressure by users for more timely data, has stimulated research on the use of nonprobability sample data, such as administrative records, web scraping, mobile phone data and voluntary internet surveys, for inference on finite population parameters like means and totals. These data are often easier, faster and cheaper to collect than traditional probability samples. However, a major concern with the use of this kind of data for official statistics is their nonrepresentativeness due to possible selection bias, which if not accounted for properly, could bias the inference. In this article, we review and discuss methods considered in the literature to deal with this problem and propose new methods, distinguishing between methods based on integration of the nonprobability sample with an appropriate probability sample, and methods that base the inference solely on the nonprobability sample. Empirical illustrations, based on simulated data are provided.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100010
    Description: The discussants highlight promising research topics for improving the quality and granularity of estimates from surveys. We agree that continued research is needed to evaluate models used for inference, and suggest development of measures of model dependence.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100011
    Description: This discussion examines some advancements in survey design and estimation, inspired by the comprehensive appraisal of Professors Jon Rao and Sharon Lohr on current trends in the field. It delves into three specific areas: balanced sampling, calibration, and small area estimation. Probabilistic balanced sampling methods, such as the cube method and penalized balanced sampling, are explored, with an emphasis on addressing emerging challenges, including extensions to linear mixed models, nonparametric regression models, and spatially balanced designs. Calibration is discussed using a modular framework that incorporates modern regression techniques, and highlights innovative uses of model calibration for data editing and causal inference. Small area estimation is considered in the context of latent variable modeling and data integration, emphasizing its role when the variable(s) of interest cannot be measured either directly or without error. Applications in integrating probability and non-probability data and conducting causal analysis at local level are also discussed.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100012
    Description: In this discussion, we complement the excellent overview by Profs. Lohr and Rao with some additional topics. The first topic is a call for more recognition of the central role of modeling in survey estimation. The second is a brief discussion of the use of partial frame information in survey design. Finally, we draw the attention to recent increases of synthetic methods, in particular, multilevel regression and poststratification (MRP) in small area estimation applications.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202500100014
    Description: Rao (1999) summarized trends in sample survey theory and methods at the turn of the millenium. We provide an updated discussion of some current trends in survey design and estimation methods for the 50th anniversary of Survey Methodology. Recent innovations in survey design include research on anticipating nonsampling errors at the design stage and development of balanced and adaptive sampling designs to take advantage of detailed sampling frame information or data gathered during the survey process. Nonparametric and machine learning methods are increasingly used for data editing as well as for model-assisted estimation and nonresponse adjustments. Small area models have been expanded to incorporate spatial and time series information, increase the flexibility and robustness of the linking and variance models, benchmark to large-area direct estimators, and (for unit level models) account for informative sampling designs. The increasing availability of large administrative datasets, sensor and satellite data, and convenience samples has spurred research on how to use these sources - on their own and when integrated with probability samples. We conclude by discussing some frontiers for survey research.
    Release date: 2025-06-30

  • Articles and reports: 12-001-X202400200004
    Description: While we avoid specifying the parametric relationship between the study variable and covariates, we illustrate the advantage of including a spatial component to better account for the covariates in our models to make Bayesian predictive inference. We treat each unique covariate combination as an individual stratum, then we use small area estimation techniques to make inference about the finite population mean of the continuous response variable. The two spatial models used are the conditional autoregressive and simple conditional autoregressive models. We include the spatial effects by creating the adjacency matrix via the Mahalanobis distance between covariates. We also show how to incorporate survey weights into the spatial models when dealing with probability survey data. We compare the results of two non-spatial models including the Scott-Smith model and the Battese, Harter, and Fuller model to the spatial models. We illustrate the comparison between the aforementioned models with an application using BMI data from eight counties in California. Our goal is to have neighboring strata yield similar predictions, and to increase the difference between strata that are not neighbors. Ultimately, using the spatial models shows less global pooling compared to the non-spatial models, which was the desired outcome.
    Release date: 2024-12-20
Data (0)

Data (0) (0 results)

No content available at this time.

Analysis (610)

Analysis (610) (590 to 600 of 610 results)

  • Articles and reports: 12-001-X197800254833
    Description: Owners of small businesses complain about the quantity of forms they are required to collectors of statistics. Administrative data are an alternative source but do not usually include all the information required by the survey takers.

    The “Tax Data Imputation System” makes use of tax data collected from a large number of businesses by Revenue Canada and data obtained by sample survey for a small subset of these businesses. Survey data is imputed (estimated) for all the businesses not actually surveyed using a “hot-deck” technique, with adjustments made to ensure certain edit rules are satisfied. The results of a simulation study suggest that this procedure has reasonable statistical properties. Estimators (of means or totals) are unbiased with variances of comparable size to the corresponding ratio estimators.
    Release date: 1978-12-15

  • Articles and reports: 12-001-X197800254835
    Description: Some estimators alternative to the usual PPS estimator are suggested in this paper for situations where the size measure used for PPS sampling is not correlated with the study variable and where data are available on another supplementary variable (size measure). Properties of these estimators are studied under super-population models and also empirically.
    Release date: 1978-12-15

  • Articles and reports: 12-001-X197800154831
    Description: The impact on linear statistics of the sample design used in obtaining survey data is the subject of much of sampling literature. Recently, more attention has been paid to the design’s impact on non-linear statistics; the major factor inhibiting these investigations has been the problem of estimating at least the first two moments of such statistics. The present article examines the problem of estimating the variances of non-linear statistics from complex samples, in the light of existing literature. The behaviour of the chi-square statistic computed from a complex sample to test hypotheses of goodness of fit or independence is studied. Alternative tests are developed and their properties studied in simulation experiments.
    Release date: 1978-06-15

  • Articles and reports: 12-001-X197800154833
    Description: The total variance of a survey estimate incorporates sampling variance, simple response variance and correlated response variance. The last component reflects the part of the total variance due to a common influence on a group of respondents. In the Canadian census, self-enumeration was adopted as the standard method of enumeration in the 1971 Census. One factor in favor of introducing this method was evidence, from the 1961 Census, that correlated response variance made an important contribution to the total variance of census estimates. Based on a study conducted using interpenetration of interviewers, this article compares correlated response variances from the 1961, 1971 and 1976 Censuses. The empirical results demonstrate that although the self-enumeration adopted in the 1971 Census did not completely remove the correlated response variance, this approach has considerably reduced the magnitude of this component of variance for almost all the characteristics examined.
    Release date: 1978-06-15

  • Articles and reports: 12-001-X197800154835
    Description: Raking ratio estimators give estimates of the population values of characteristics examined on a sample basis utilizing the row and column totals of a contingency table of characteristics examined on a 100% basis. In this paper, the asymptotic variance of the maximum likelihood estimator of a sample characteristic subject to the marginal constraints of the above contingency table is derived. From this, we are able to compute the loss in efficiency of the raking ratio estimators relative to the maximum likelihood estimator in an empirical study.
    Release date: 1978-06-15

  • Articles and reports: 12-001-X197700254829
    Description: Results of an earlier paper on the use of raking ratio estimators are extended to the case of cluster sampling. An empirical study is discussed.
    Release date: 1977-12-12

  • Articles and reports: 12-001-X197700254832
    Description: In periodic household surveys, area samples are usually selected in geographic strata with probability of selection of areal units proportional to population size in these units. The design-based estimates for areas composed of domains within strata can have poor precision due to cluster sampling with a few primary sampling units per stratum. In this paper, synthetic estimates are investigated as an alternative to these estimates. An empirical evaluation based on the design of the Canadian Labour Force, Survey is given.
    Release date: 1977-12-12

  • Articles and reports: 12-001-X197700100002
    Description: In multi-stage sampling when selection is without replacement at the first stage, estimation of the variance of the estimate of the population total is often done assuming sampling with replacement. This estimate is biased and the degree of bias is not negligible. In this paper, a procedure which gives unbiased estimates of the variance making use of only estimated primary sampling unit totals is suggested for the case when sampling at the second and subsequent stages is simple random without replacement. This procedure is based on sub-samples drawn from the selected second and subsequent stage units.
    Release date: 1977-06-20

  • Articles and reports: 12-001-X197700100004
    Description: The 1971 and 1976 Censuses of Population and Housing have utilized the raking ratio estimation procedure to obtain estimates for variables collected only on a sample basis. This paper derives large sample approximations for the bias and variance of such estimates and examines their performance in an empirical study.
    Release date: 1977-06-20

  • Articles and reports: 12-001-X197700100005
    Description: Objective yield surveys have been conducted annually in the Niagara Peninsula since 1964. The aim of each of these annual surveys is to provide a forecast of the marketable production change in the region from the previous year. These estimates are determined far enough in advance of the harvest to enable them to serve as important factors in price negotiations between growers and processors, as well as indicators of particular crop situations which could necessitate immediate changes in strategy by the marketing agencies. In 1973 an extensive redesign project was initiated. This report provides a summary of the sample design, data collection procedures and estimation procedures which were incorporated in the redesign of the sour cherry, peach and grape objective yield surveys.
    Release date: 1977-06-20
Reference (28)

Reference (28) (0 to 10 of 28 results)

  • Surveys and statistical programs – Documentation: 11-633-X2026002
    Description: Recent changes in Canada’s immigration levels have heightened interest in understanding how immigration affects housing demand. This article develops a methodological framework for projecting housing use associated with permanent residents (PRs) and non-permanent residents (NPRs) under alternative immigration scenarios. The framework applies observed per capita housing use rates from the Census of Population to estimate incremental housing use by tenure over time.
    Release date: 2026-04-24

  • Surveys and statistical programs – Documentation: 91-528-X
    Description: The Technical Guide on Demographic Estimates at Statistics Canada provides detailed descriptions of the most current data sources and methods used by the Centre for demography at Statistics Canada to produce demographic estimates as part of the Demographic estimates program. They comprise postcensal and intercensal population estimates; base population; births and deaths; immigrants; emigrants; returning emigrants; non-permanent residents; interprovincial migration; subprovincial estimates of population and intraprovincial migration; population estimates by age and gender; and census family estimates. A glossary of commonly used terms is available at the end of the guide.
    Release date: 2025-12-17

  • Surveys and statistical programs – Documentation: 98-306-X
    Description:

    This report describes sampling, weighting and estimation procedures used in the Census of Population. It provides operational and theoretical justifications for them, and presents the results of the evaluations of these procedures.

    Release date: 2023-10-04

  • Notices and consultations: 75F0002M2019006
    Description:

    In 2018, Statistics Canada released two new data tables with estimates of effective tax and transfer rates for individual tax filers and census families. These estimates are derived from the Longitudinal Administrative Databank. This publication provides a detailed description of the methods used to derive the estimates of effective tax and transfer rates.

    Release date: 2019-04-16

  • Surveys and statistical programs – Documentation: 75F0002M2015003
    Description:

    This note discusses revised income estimates from the Survey of Labour and Income Dynamics (SLID). These revisions to the SLID estimates make it possible to compare results from the Canadian Income Survey (CIS) to earlier years. The revisions address the issue of methodology differences between SLID and CIS.

    Release date: 2015-12-17

  • Surveys and statistical programs – Documentation: 13-605-X201500414166
    Description:

    Estimates of the underground economy by province and territory for the period 2007 to 2012 are now available for the first time. The objective of this technical note is to explain how the methodology employed to derive upper-bound estimates of the underground economy for the provinces and territories differs from that used to derive national estimates.

    Release date: 2015-04-29

  • Surveys and statistical programs – Documentation: 99-002-X2011001
    Description:

    This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.

    Release date: 2015-01-28

  • Surveys and statistical programs – Documentation: 99-002-X
    Description: This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.
    Release date: 2015-01-28

  • Surveys and statistical programs – Documentation: 92-568-X
    Description:

    This report describes sampling and weighting procedures used in the 2006 Census. It reviews the history of these procedures in Canadian censuses, provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.

    Release date: 2009-08-11

  • Surveys and statistical programs – Documentation: 71F0031X2006003
    Description:

    This paper introduces and explains modifications made to the Labour Force Survey estimates in January 2006. Some of these modifications include changes to the population estimates, improvements to the public and private sector estimates and historical updates to several small Census Agglomerations (CA).

    Release date: 2006-01-25