Weighting and estimation
Filter results by
Search HelpKeyword(s)
Type
Survey or statistical program
Results
All (638)
All (638) (600 to 610 of 638 results)
- Articles and reports: 12-001-X198400214357Description:
A finite population of size N is supposed to contain M (unknown) units of a specified category A (say) constituting a domain with mean \mu. A procedure which involves drawing units using simple random sampling without replacement till a preassigned number of members of the domain is reached is proposed. An unbiased estimator of \mu is also derived. This is seen to be superior to the corresponding possibly biased estimator based on a comparable SRSWOR scheme with a fixed number of draws. The proposed scheme is also shown to admit unbiased estimators of M and the domain total T.
Release date: 1984-12-14 - Articles and reports: 12-001-X198300214342Description:
This study considers the suitability of composite estimation techniques for the Canadian Labour Force Survey. The performance of a class of AK composite estimators introduced initially by Gurney and Daly is investigated for several characteristics. While the ordinary composite estimate has a large bias, the AK composite estimate is capable of reducing the bias. Composite estimates having minimum variance and minimum mean square error are compared.
Release date: 1983-12-15 - Articles and reports: 12-001-X198300214344Description:
In order to improve the timeliness, accuracy and consistency of population estimates for different geographic areas, Statistics Canada has developed new methods of estimation for sub-provincial areas (census divisions and census metropolitan areas). Beginning with 1982, two sets of population estimates (regression and component based) will be published yearly, appearing 3-4 months and 12-15 months, respectively, from the reference date.
The regression technique uses family allowance recipients as the main symptomatic indicator and where available, additional indicators - reference population from provincial health insurance files and hydro accounts - to derive population change for the current year. The first set is obtained by adding this change to the second set for the previous year produced by the component method, with births and deaths from vital registers, and estimated migration from Revenue Canada taxation files. The two sets were found to be statistically similar with respect to accuracy, though the first set is more timely, and the second provides more details on the components of population change.
Release date: 1983-12-15 - 604. The methodology of the Canadian Air Scheduled International Passenger Origin and Destination estimation system ArchivedArticles and reports: 12-001-X198300114333Description:
The Air Scheduled International Passenger Origin and Destination (ASIPOD) estimation system uses the data from two air traffic surveys to produce origin-destination estimates of international passengers. The “assignment technique” is the solution to the problem caused by the non-coverage of non-interlining traffic. The assumptions of the technique are sufficiently questionable to warrant an evaluation of the bias of the estimates. However, major improvements will be made in the new system which will decrease the bias in the estimates. Also, estimates of reliability will be produced. And as a result, knowledge of the strength of the inferences made with respect to air traffic markets from these estimates will be improved in international bilateral air negotiations.
Release date: 1983-06-15 - Articles and reports: 12-001-X198300114335Description:
The Canadian Labour Force Survey is a household survey conducted each month for the purpose of producing point-in-time estimates of the number of persons employed, unemployed and not in the labor force. The survey has a rotating panel design in which all individuals in a sampled household location are interviewed each month, for six consecutive months. In the past, little use has been made of this longitudinal structure, although considerable interest has been expressed in the month-to-month gross flows (transitions) amongst the labour force status categories. In this paper we discuss methods being considered by Statistics Canada for the production of gross flow estimates, but from a model-based perspective.
Release date: 1983-06-15 - Articles and reports: 12-001-X198300114336Description:
The peach, sour cherry and the grape objective yield surveys have been carried out annually in the Niagara Peninsula since 1964 in order to forecast the magnitude of change in marketable fruit production from the previous year. Timeliness of the estimates is essential in order to enable the Ontario Tender Fruit Growers Marketing Board (OTFGMB) and the Ontario Grape Growers Marketing Board (OGGMB) to establish the marketing strategies well ahead of the harvest. This paper summarizes the major changes due to the second redesign initiated in 1982. In particular, the sample design, data collection operation and modifications of the estimation procedures are elaborated upon.
Release date: 1983-06-15 - 607. A timely and accurate potato acreage estimate from Landsat: Results of a demonstration ArchivedArticles and reports: 12-001-X198300114337Description:
This paper describes the procedures used and results of a joint Canada Centre for Remote Sensing (CCRS) and Statistics Canada project to provide a timely potato acreage estimate for New Brunswick, a major potato producing province in Canada. The project has demonstrated that satellite imagery combined with more traditional potato area estimation procedures can lower respondent burden, produce timely crop distribution maps and produce reliable estimates for subregions.
Release date: 1983-06-15 - 608. Sampling on two occasions with probabilities proportional to size without replacement (PPSWOR) ArchivedArticles and reports: 12-001-X198300114340Description:
A theory of sampling on two occasions with unequal probabilities and without replacement is presented. Fellegi’s (1963) method, which yields the same selection probabilities for a given unit on each occasion, is used to select the units for the rotation sample. The variances of composite estimators of the population total on the second occasion are developed. Numerical results are presented for small sample sizes and efficiency comparisons are made with a competing strategy.
Release date: 1983-06-15 - Articles and reports: 12-001-X198200114328Description:
Estimates from sample surveys are sometimes required for domains whose boundaries do not coincide with those of design strata. Taking the Canadian Labour Force Survey as an example of a survey utilizing a clustered sample design, some alternative small area estimation techniques available in the literature are evaluated empirically including synthetic, domain (simple and post-stratified) and composite estimators which are linear combinations of synthetic and post-stratified domain estimators. A sample dependent estimator which attaches weight to the post-stratified domain estimate depending on the amount of sample in the domain is proposed and its performance is also evaluated.
Release date: 1982-06-15 - 610. Computerization of complex survey estimates ArchivedArticles and reports: 12-001-X198200114331Description:
Survey data collected by statistical agencies is most likely to be processed through to the tabulation stage by these agencies. The computer programs associated with this processing are also most likely tailored to the particular design and variables used. The statistics computed from such surveys typically range from simple descriptive totals and means to these required for analytic studies such as comparison of domains, regression analysis and contingency tables analysis. This paper describes a computer program which computes these statistics and their associated sampling errors for commonly used sampling designs.
Release date: 1982-06-15
- Previous Go to previous page of All results
- 1 Go to page 1 of All results
- ...
- 58 Go to page 58 of All results
- 59 Go to page 59 of All results
- 60 Go to page 60 of All results
- 61 (current) Go to page 61 of All results
- 62 Go to page 62 of All results
- 63 Go to page 63 of All results
- 64 Go to page 64 of All results
- Next Go to next page of All results
Data (0)
Data (0) (0 results)
No content available at this time.
Analysis (610)
Analysis (610) (0 to 10 of 610 results)
- Articles and reports: 12-001-X202600100003Description: Probability-proportional-to-size sampling is widely used by national statistical offices. Here population units are selected with probabilities proportional to an auxiliary variable. Variance formulas in such designs require both first- and second-order inclusion probabilities. The computation of second-order inclusion probabilities is particularly challenging for large populations, and has been the subject of extensive research. This article presents some new exact and approximation formulas for second-order inclusion probabilities in randomized systematic sampling with unequal probabilities and without replacement.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100005Description: Confidence intervals are very often constructed based on a probability distribution that uses a certain number of degrees of freedom as a parameter. This is the case with the Student and the modified Wilson confidence intervals, discussed in this article, which use quantiles from the Student distribution where the number of degrees of freedom is generally unknown. For the length of a confidence interval to be representative of the reliability of an estimate, the actual coverage rate must match the nominal rate. To that end, the number of degrees of freedom in the probability distribution used in practice to calculate the confidence interval must be estimated as precisely as possible. An approximate rule is often used, although it tends to overestimate the actual number of degrees of freedom. In this article, a more precise version of degrees of freedom, derived from the Satterthwaite approximation, is obtained in the context of the Canadian Census of Population. The sampling design is equivalent to a simple random design without replacement, cluster-stratified, and the variance estimation method is an adaptation of the balanced repeated replication method. An explicit expression of the degrees of freedom is obtained under these conditions, enabling the factors influencing them to be identified. For comparison, the degree of freedom formula is also established for the conventional variance estimator. A simulation study shows that using this version of degrees of freedom corrects the undercoverage problem observed with the approximate rule, showing the importance of accurately assessing this number.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100009Description: Combining estimates from independent surveys via inverse-variance weights can lead to negative bias when unknown variances are estimated and the target variable is non-negative and positively skewed. In such cases, strong positive correlations typically arise between the estimators and their corresponding variance estimators, causing standard linear combinations with inverse-variance weights to exhibit negative bias. We introduce a strikingly simple method to reduce bias: replace the standard weight with the ratio of the estimator to the variance estimator. Under a linear model linking the two, we show that the new ratio-weighted estimator is approximately unbiased, whereas the conventional inverse-variance combination exhibits downward bias. Through simulations, we demonstrate that the new method brings both the bias and the mean squared error closer to the optimum for a wide range of different target variables. As our method uses only standardly reported summary statistics, it can be immediately adopted to reduce this widespread bias and improve the reliability of scientific findings in various fields.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100010Description: With the exception of two-phase sampling, the standard variance approximation of the generalized regression (GREG) estimator assumes that the population totals in the weighting scheme are observed without error. If the weighting model of the GREG estimator contains population totals that are observed with measurement error sources other than the sampling error of first-phase estimates, then this uncertainty will be ignored by the variance approximation of the GREG estimator. This paper proposes a variance approximation for the GREG estimator that accounts for additional uncertainty arising from measurement error in one or more of the population totals used in the weighting scheme. This approach has been developed for, and is being applied to, the Dutch Labour Force Survey (DLFS). The monthly publications of the DLFS are obtained with a time series model, which corrects for rotation group bias and discontinuities caused by major redesigns and the loss of face-to-face interviews during COVID-19. The GREG estimates for the quarterly figures are benchmarked to the average of the monthly publications to enforce numerical consistency between monthly and quarterly publication tables. The standard variance approximation of the GREG estimator assumes that these population totals are observed without error. This results in an underestimation of the variance of the GREG estimator. The variance approximation proposed in this paper results in more realistic standard errors for the quarterly GREG estimates.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100011Description: We construct a hybrid Bayesian method, which includes a differentially private mechanism, to mask Census county totals for a U.S. state on acreage of a commodity. We use surrogates for data collected at the farm level from a past U.S. Census of Agriculture to illustrate our procedure. We use two Bayesian small area models (parametric and mixture) to accommodate the smaller counties with fewer farms and some counties with large acres. In these models, the Laplace distribution provides a differentially private mechanism. In pre-processing, we also incorporate the Census weights to form the observed total acreage, a scaling factor to the Laplace mechanism for each county, a square-root transformation of the observed total acreage to avoid negative masked estimates especially for small counties, and the p-percent rule and the 3+ rule to partition the counties into suppressed counties, non-sensitive counties and sensitive counties. Because of difficulties in specifying and tuning the privacy budget (an unknown parameter), to balance security and utility, we specify a prior for the privacy budget, where the values are not specified, and the Gibbs sampler is used to fit the hierarchical Bayesian models. In post-processing, we use Bayesian predictive inference to obtain masked county acreages, and this includes a benchmarking so that the masked state total matches the observed state total. As a measure of reliability of the Bayesian procedure, we use the posterior coefficients of variation for the masked posterior means of the counties. As a measure of utility, we use the absolute relative errors for the individual counties, together with other global measures. For the sensitive counties, there are some differences between the two small area models but both are much better than an individual area model; the mixture model being the best compromise for security and utility.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100012Description: We propose small area estimators of general indicators in off-census years, which avoid the use of deprecated census microdata, but are nearly optimal in census years. The procedure is based on replacing the obsolete census file with a larger unit-level survey that adequately covers the areas of interest and contains the values of useful auxiliary variables. However, the minimal data requirement of the proposed method is a single survey with microdata on the target variable and suitable auxiliary variables for the period of interest. We also develop an estimator of the mean squared error (MSE) that accounts for the uncertainty introduced by the large survey used to replace the census of auxiliary information. Our empirical results indicate that the proposed predictors perform clearly better than the alternative predictors when census data are outdated, and are very close to optimal ones when census data are correct. They also illustrate that the proposed total MSE estimator corrects for the bias of purely model-based MSE estimators that do not account for the large survey uncertainty.Release date: 2026-06-29
- Articles and reports: 12-001-X202500200001Description: Nested error regression models are commonly used to incorporate unit specific auxiliary variables to improve small area estimates. When the mean structure of the model is misspecified, the design-based mean squared prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP) generally increases. The Observed Best Prediction (OBP) method has been proposed with the intent to improve on the design-based MSPE over EBLUP. In this paper, we conduct a Monte Carlo simulation experiments to understand the effect of misspsecification of mean structures on different small area estimators. Our findings suggest that the OBP using unit-level auxiliary variables does not outperform the EBLUP in terms of design-based MSPE, unless the number of small areas m is extremely large. Conversely, the performance of OBP significantly improves when area-level auxiliary variables are employed. This paper includes both analytical and numerical evidence to demonstrate these observations, providing practical insights for addressing model misspecification in small area estimation (SAE).Release date: 2025-12-23
- Articles and reports: 12-001-X202500200003Description: In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.Release date: 2025-12-23
- Articles and reports: 12-001-X202500200005Description: The use of non-probability data sources for statistical purposes and for official statistics has become increasingly popular in recent years. However, statistical inference based on non-probability samples is made more difficult by nature of their biasedness and lack of representativity. In this paper we propose quantile balancing inverse probability weighting estimator (QBIPW) for non-probability samples. We apply the idea of Harms and Duchesne (2006) allowing the use of quantile information in the estimation process to reproduce known totals and the distribution of auxiliary variables. We discuss the estimation of the QBIPW probabilities and its variance. Our simulation study has demonstrated that the proposed estimators are robust against model mis-specification and, as a result, help to reduce bias and mean squared error. Finally, we applied the proposed methods to estimate the share of job vacancies aimed at Ukrainian workers in Poland using an integrated set of administrative and survey data about job vacancies.Release date: 2025-12-23
- Articles and reports: 12-001-X202500200009Description: We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our methodology uses hierarchical global-local prior distributions to make inferences for the proportion of individuals in Florida’s counties who do not have health insurance. Results from an extensive simulation study show that this methodology will provide improved inference by using several data sources. Among the five model variants evaluated the ones using horseshoe priors for all variances have better performance than the ones using lasso priors for the local variances.Release date: 2025-12-23
- Previous Go to previous page of Analysis results
- 1 (current) Go to page 1 of Analysis results
- 2 Go to page 2 of Analysis results
- 3 Go to page 3 of Analysis results
- 4 Go to page 4 of Analysis results
- 5 Go to page 5 of Analysis results
- 6 Go to page 6 of Analysis results
- 7 Go to page 7 of Analysis results
- ...
- 61 Go to page 61 of Analysis results
- Next Go to next page of Analysis results
Reference (28)
Reference (28) (0 to 10 of 28 results)
- Surveys and statistical programs – Documentation: 11-633-X2026002Description: Recent changes in Canada’s immigration levels have heightened interest in understanding how immigration affects housing demand. This article develops a methodological framework for projecting housing use associated with permanent residents (PRs) and non-permanent residents (NPRs) under alternative immigration scenarios. The framework applies observed per capita housing use rates from the Census of Population to estimate incremental housing use by tenure over time.Release date: 2026-04-24
- Surveys and statistical programs – Documentation: 91-528-XDescription: The Technical Guide on Demographic Estimates at Statistics Canada provides detailed descriptions of the most current data sources and methods used by the Centre for demography at Statistics Canada to produce demographic estimates as part of the Demographic estimates program. They comprise postcensal and intercensal population estimates; base population; births and deaths; immigrants; emigrants; returning emigrants; non-permanent residents; interprovincial migration; subprovincial estimates of population and intraprovincial migration; population estimates by age and gender; and census family estimates. A glossary of commonly used terms is available at the end of the guide.Release date: 2025-12-17
- Surveys and statistical programs – Documentation: 98-306-XDescription:
This report describes sampling, weighting and estimation procedures used in the Census of Population. It provides operational and theoretical justifications for them, and presents the results of the evaluations of these procedures.
Release date: 2023-10-04 - Notices and consultations: 75F0002M2019006Description:
In 2018, Statistics Canada released two new data tables with estimates of effective tax and transfer rates for individual tax filers and census families. These estimates are derived from the Longitudinal Administrative Databank. This publication provides a detailed description of the methods used to derive the estimates of effective tax and transfer rates.
Release date: 2019-04-16 - 5. Revisions to 2006 to 2011 income data ArchivedSurveys and statistical programs – Documentation: 75F0002M2015003Description:
This note discusses revised income estimates from the Survey of Labour and Income Dynamics (SLID). These revisions to the SLID estimates make it possible to compare results from the Canadian Income Survey (CIS) to earlier years. The revisions address the issue of methodology differences between SLID and CIS.
Release date: 2015-12-17 - Surveys and statistical programs – Documentation: 13-605-X201500414166Description:
Estimates of the underground economy by province and territory for the period 2007 to 2012 are now available for the first time. The objective of this technical note is to explain how the methodology employed to derive upper-bound estimates of the underground economy for the provinces and territories differs from that used to derive national estimates.
Release date: 2015-04-29 - Surveys and statistical programs – Documentation: 99-002-X2011001Description:
This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.
Release date: 2015-01-28 - Surveys and statistical programs – Documentation: 99-002-XDescription: This report describes sampling and weighting procedures used in the 2011 National Household Survey. It provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.Release date: 2015-01-28
- Surveys and statistical programs – Documentation: 92-568-XDescription:
This report describes sampling and weighting procedures used in the 2006 Census. It reviews the history of these procedures in Canadian censuses, provides operational and theoretical justifications for them, and presents the results of the evaluation studies of these procedures.
Release date: 2009-08-11 - Surveys and statistical programs – Documentation: 71F0031X2006003Description:
This paper introduces and explains modifications made to the Labour Force Survey estimates in January 2006. Some of these modifications include changes to the population estimates, improvements to the public and private sector estimates and historical updates to several small Census Agglomerations (CA).
Release date: 2006-01-25