Weighting and estimation

Skip to filters. View results.

Sort Help
entries

Results

All (638)

All (638) (10 to 20 of 638 results)

  • Articles and reports: 12-001-X202500200009
    Description: We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our methodology uses hierarchical global-local prior distributions to make inferences for the proportion of individuals in Florida’s counties who do not have health insurance. Results from an extensive simulation study show that this methodology will provide improved inference by using several data sources. Among the five model variants evaluated the ones using horseshoe priors for all variances have better performance than the ones using lasso priors for the local variances.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200010
    Description: In this paper, we study the performance of hierarchical Bayes (HB) small area estimators using noninformative and informative priors. We apply the Bayesian models of You and Chapman (2006) and You (2021) to the Canadian Labor Force Survey (LFS) data and evaluate the impact of the priors on the HB estimators. A Bayesian model comparison and simulation study are also conducted. Our results indicate that a correct informative prior can lead to very good results, and noninformative priors can also perform very well. Incorrect informative priors can lead to poor results in terms of large bias and large coefficient of variation (CV). Noninformative priors are recommended in practice for HB small area estimation unless correctly specified informative priors are available. Informative priors are particularly useful when the number of small areas is relatively small.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200011
    Description: We propose an approximate hierarchical Bayes approach that uses the Natural Exponential Family with Quadratic Variance Function (NEF-QVF) in combining information from multiple sources to improve traditional survey estimates of finite population means for small areas. Unlike other Bayesian approaches in finite population sampling, we do not assume a model for all units of the finite population and do not require linking sampled units to the finite population frame. We assume a model only for the finite population units in which the outcome variable is observed; because, for these units, the assumed model can be checked using existing statistical tools. We do not posit an elaborate model on the true means for unobserved units. Instead, we assume that population means of cells with the same combination of factor levels are identical across small areas, and that the population mean for a cell is identical to the mean of the observed units in that cell. We apply our proposed methodology to a real-life survey, linking information from multiple disparate data sources. We also provide practical ways of model selection that can be applied to a wider class of models under similar setting but for a diverse range of scientific problems.
    Release date: 2025-12-23

  • Articles and reports: 12-001-X202500200012
    Description: The observed best prediction (OBP) under a nested-error regression (NER) model was previously proposed using a design-based mean squared prediction error (MSPE) as a tool to derive the best predictive estimator (BPE). A recent study showed the OBP under the NER model may suffer from numerical instability when computing the BPE. We propose several modifications of the OBP under the NER model, including ones using a model-based MSPE to derive the BPE, to improve the numerical stability and predictive performance. We compare the performance of the modified OBP strategies with the existing methods in a simulation study. A real-data example is discussed.
    Release date: 2025-12-23

  • Surveys and statistical programs – Documentation: 91-528-X
    Description: The Technical Guide on Demographic Estimates at Statistics Canada provides detailed descriptions of the most current data sources and methods used by the Centre for demography at Statistics Canada to produce demographic estimates as part of the Demographic estimates program. They comprise postcensal and intercensal population estimates; base population; births and deaths; immigrants; emigrants; returning emigrants; non-permanent residents; interprovincial migration; subprovincial estimates of population and intraprovincial migration; population estimates by age and gender; and census family estimates. A glossary of commonly used terms is available at the end of the guide.
    Release date: 2025-12-17

  • Articles and reports: 11-522-X202500100006
    Description: Small area estimation is frequently used to produce estimates at a disaggregated level where direct survey estimation does not have sufficient sample to produce precise estimates. Often this is done using the area-level Fay-Herriot model, by assuming the direct estimates are independent under the design and have a known variance, and applying a smoothing process to the variance estimates of the direct estimates to better meet that last assumption. It is not rare that small area estimates are benchmarked/raked to aggregated level direct estimates. This article shows that wrongly assuming independence can have a big impact on the MSE of the raked estimates. Values of the covariances between direct estimates are thus required for good point and MSE estimates. Getting good estimates of those covariances is difficult given the small sample sizes in some areas. An original way of deriving values for those covariances, by reverse-engineering a hypothetical raking process, is presented.
    Release date: 2025-09-08

  • Articles and reports: 11-522-X202500100007
    Description: This paper employs the Pseudo Maximum Likelihood (PML) estimator to the non-probability two-phase sampling when relevant auxiliary information is available from both probability survey sample and non-probability survey sample. To accommodate various weight adjustments and estimates variance beyond totals and means such as medians and quantiles, a simplified pseudo-population bootstrap procedure is proposed to approximately estimate the second-phase variance. Specifically, the simplification ignores the second phase sampling variability (i.e., treated as fixed, while in fact it is random), if the first-phase sampling fraction of the non-probability sample is negligible. Using the Bank of Canada 2020 Cash Alternative Survey Wave 2, the performance of the proposed method is compared to alternative methods, which either do not explicitly model the selection probability (i.e., raking) or ignore the valuable information from Phase 1 (i.e., Phase-2-Only). The results show that the PML-based approach performs better than raking and Phase-2-Only estimates in terms of reducing the selection bias for both phases' payment-related variables, especially for the low-response youth group. Estimated variances of the PML-based estimates are stable.

    Release date: 2025-09-08

  • Articles and reports: 11-522-X202500100009
    Description: Three series of web panels were implemented at Statistics Canada from 2020 to 2024. Participants for these web panel series were recruited from respondents of large probabilistic social surveys (recruitment surveys), and subsequently were invited to complete a series of short online surveys. Estimates of recruitment survey variables were calculated using both recruitment survey weights and web panel weights, and these were compared; differences signal the possibility of residual bias that was not corrected by the web panel weighting process. This investigation found more significant differences than would be expected if the web panel estimator fully corrected for the bias resulting from the web panel response process. Questions related to certain topics such as politics and voting, sense of belonging, and media consumption were found to have the most significant differences between web panel estimates and recruitment survey estimates.
    Release date: 2025-09-08

  • Articles and reports: 11-522-X202500100011
    Description: The use of modern "data"-driven imputation methods to treat non-response in the context of surveys processed in the Integrated Business Statistics Program at Statistics Canada has previously been explored. It was observed that these methods can lead to high quality imputation and further have the potential to result in broad efficiencies when setting up a particular survey's edit and imputation strategy. However, estimation of the associated total variance, more specifically the component due to imputation, remains a challenge. In this article, two methods for estimation of total variance are proposed and show preliminary results that have motivated us to pursue further research in this area.
    Release date: 2025-09-08

  • Articles and reports: 11-522-X202500100028
    Description: The United Nations Sustainable Development Goals require detailed, disaggregated data, typically obtained through household surveys. However, surveys alone cannot meet these needs for granular statistics. To address this, National Statistical Institutes adopt small area methods, but these face challenges as auxiliary variables, often derived from surveys, introduce measurement errors into the models. The aim is the application of measurement error correction in classic Fay-Herriot area-level model. The results demonstrate the robustness of the standard approach and ignoring measurement error but show there are specific scenarios where correction for measurement errors is beneficial. The approach is applied to a case study utilizing Indonesian household survey data.
    Release date: 2025-09-08
Data (0)

Data (0) (0 results)

No content available at this time.

Analysis (610)

Analysis (610) (590 to 600 of 610 results)

  • Articles and reports: 12-001-X197800254833
    Description: Owners of small businesses complain about the quantity of forms they are required to collectors of statistics. Administrative data are an alternative source but do not usually include all the information required by the survey takers.

    The “Tax Data Imputation System” makes use of tax data collected from a large number of businesses by Revenue Canada and data obtained by sample survey for a small subset of these businesses. Survey data is imputed (estimated) for all the businesses not actually surveyed using a “hot-deck” technique, with adjustments made to ensure certain edit rules are satisfied. The results of a simulation study suggest that this procedure has reasonable statistical properties. Estimators (of means or totals) are unbiased with variances of comparable size to the corresponding ratio estimators.
    Release date: 1978-12-15

  • Articles and reports: 12-001-X197800254835
    Description: Some estimators alternative to the usual PPS estimator are suggested in this paper for situations where the size measure used for PPS sampling is not correlated with the study variable and where data are available on another supplementary variable (size measure). Properties of these estimators are studied under super-population models and also empirically.
    Release date: 1978-12-15

  • Articles and reports: 12-001-X197800154831
    Description: The impact on linear statistics of the sample design used in obtaining survey data is the subject of much of sampling literature. Recently, more attention has been paid to the design’s impact on non-linear statistics; the major factor inhibiting these investigations has been the problem of estimating at least the first two moments of such statistics. The present article examines the problem of estimating the variances of non-linear statistics from complex samples, in the light of existing literature. The behaviour of the chi-square statistic computed from a complex sample to test hypotheses of goodness of fit or independence is studied. Alternative tests are developed and their properties studied in simulation experiments.
    Release date: 1978-06-15

  • Articles and reports: 12-001-X197800154833
    Description: The total variance of a survey estimate incorporates sampling variance, simple response variance and correlated response variance. The last component reflects the part of the total variance due to a common influence on a group of respondents. In the Canadian census, self-enumeration was adopted as the standard method of enumeration in the 1971 Census. One factor in favor of introducing this method was evidence, from the 1961 Census, that correlated response variance made an important contribution to the total variance of census estimates. Based on a study conducted using interpenetration of interviewers, this article compares correlated response variances from the 1961, 1971 and 1976 Censuses. The empirical results demonstrate that although the self-enumeration adopted in the 1971 Census did not completely remove the correlated response variance, this approach has considerably reduced the magnitude of this component of variance for almost all the characteristics examined.
    Release date: 1978-06-15

  • Articles and reports: 12-001-X197800154835
    Description: Raking ratio estimators give estimates of the population values of characteristics examined on a sample basis utilizing the row and column totals of a contingency table of characteristics examined on a 100% basis. In this paper, the asymptotic variance of the maximum likelihood estimator of a sample characteristic subject to the marginal constraints of the above contingency table is derived. From this, we are able to compute the loss in efficiency of the raking ratio estimators relative to the maximum likelihood estimator in an empirical study.
    Release date: 1978-06-15

  • Articles and reports: 12-001-X197700254829
    Description: Results of an earlier paper on the use of raking ratio estimators are extended to the case of cluster sampling. An empirical study is discussed.
    Release date: 1977-12-12

  • Articles and reports: 12-001-X197700254832
    Description: In periodic household surveys, area samples are usually selected in geographic strata with probability of selection of areal units proportional to population size in these units. The design-based estimates for areas composed of domains within strata can have poor precision due to cluster sampling with a few primary sampling units per stratum. In this paper, synthetic estimates are investigated as an alternative to these estimates. An empirical evaluation based on the design of the Canadian Labour Force, Survey is given.
    Release date: 1977-12-12

  • Articles and reports: 12-001-X197700100002
    Description: In multi-stage sampling when selection is without replacement at the first stage, estimation of the variance of the estimate of the population total is often done assuming sampling with replacement. This estimate is biased and the degree of bias is not negligible. In this paper, a procedure which gives unbiased estimates of the variance making use of only estimated primary sampling unit totals is suggested for the case when sampling at the second and subsequent stages is simple random without replacement. This procedure is based on sub-samples drawn from the selected second and subsequent stage units.
    Release date: 1977-06-20

  • Articles and reports: 12-001-X197700100004
    Description: The 1971 and 1976 Censuses of Population and Housing have utilized the raking ratio estimation procedure to obtain estimates for variables collected only on a sample basis. This paper derives large sample approximations for the bias and variance of such estimates and examines their performance in an empirical study.
    Release date: 1977-06-20

  • Articles and reports: 12-001-X197700100005
    Description: Objective yield surveys have been conducted annually in the Niagara Peninsula since 1964. The aim of each of these annual surveys is to provide a forecast of the marketable production change in the region from the previous year. These estimates are determined far enough in advance of the harvest to enable them to serve as important factors in price negotiations between growers and processors, as well as indicators of particular crop situations which could necessitate immediate changes in strategy by the marketing agencies. In 1973 an extensive redesign project was initiated. This report provides a summary of the sample design, data collection procedures and estimation procedures which were incorporated in the redesign of the sour cherry, peach and grape objective yield surveys.
    Release date: 1977-06-20
Reference (28)

Reference (28) (0 to 10 of 28 results)

  • Surveys and statistical programs – Documentation: 11-633-X2026002
    Description: Recent changes in Canada’s immigration levels have heightened interest in understanding how immigration affects housing demand. This article develops a methodological framework for projecting housing use associated with permanent residents (PRs) and non-permanent residents (NPRs) under alternative immigration scenarios. The framework applies observed per capita housing use rates from the Census of Population to estimate incremental housing use by tenure over time.
    Release date: 2026-04-24

  • Surveys and statistical programs – Documentation: 91-528-X
    Description: The Technical Guide on Demographic Estimates at Statistics Canada provides detailed descriptions of the most current data sources and methods used by the Centre for demography at Statistics Canada to produce demographic estimates as part of the Demographic estimates program. They comprise postcensal and intercensal population estimates; base population; births and deaths; immigrants; emigrants; returning emigrants; non-permanent residents; interprovincial migration; subprovincial estimates of population and intraprovincial migration; population estimates by age and gender; and census family estimates. A glossary of commonly used terms is available at the end of the guide.
    Release date: 2025-12-17

  • Surveys and statistical programs – Documentation: 98-306-X
    Description:

    This report describes sampling, weighting and estimation procedures used in the Census of Population. It provides operational and theoretical justifications for them, and presents the results of the evaluations of these procedures.

    Release date: 2023-10-04

  • Notices and consultations: 75F0002M2019006
    Description:

    In 2018, Statistics Canada released two new data tables with estimates of effective tax and transfer rates for individual tax filers and census families. These estimates are derived from the Longitudinal Administrative Databank. This publication provides a detailed description of the methods used to derive the estimates of effective tax and transfer rates.

    Release date: 2019-04-16

  • Surveys and statistical programs – Documentation: 75F0002M2015003
    Description:

    This note discusses revised income estimates from the Survey of Labour and Income Dynamics (SLID). These revisions to the SLID estimates make it possible to compare results from the Canadian Income Survey (CIS) to earlier years. The revisions address the issue of methodology differences between SLID and CIS.

    Release date: 2015-12-17

  • Surveys and statistical programs – Documentation: 13-605-X201500414166
    Description:

    Estimates of the underground economy by province and territory for the period 2007 to 2012 are now available for the first time. The objective of this technical note is to explain how the methodology employed to derive upper-bound estimates of the underground economy for the provinces and territories differs from that used to derive national estimates.

    Release date: 2015-04-29

  • Surveys and statistical programs – Documentation: 99-002-X2011001
    Description:

    This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.

    Release date: 2015-01-28

  • Surveys and statistical programs – Documentation: 99-002-X
    Description: This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.
    Release date: 2015-01-28

  • Surveys and statistical programs – Documentation: 92-568-X
    Description:

    This report describes sampling and weighting procedures used in the 2006 Census. It reviews the history of these procedures in Canadian censuses, provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.

    Release date: 2009-08-11

  • Surveys and statistical programs – Documentation: 71F0031X2006003
    Description:

    This paper introduces and explains modifications made to the Labour Force Survey estimates in January 2006. Some of these modifications include changes to the population estimates, improvements to the public and private sector estimates and historical updates to several small Census Agglomerations (CA).

    Release date: 2006-01-25