Keyword search
Filter results by
Search HelpKeyword(s)
Subject
Type
Year of publication
Survey or statistical program
Results
All (39)
All (39) (10 to 20 of 39 results)
- 11. Theoretical and empirical properties of model assisted decision-based regression estimators ArchivedArticles and reports: 12-001-X201400114004Description:
In 2009, two major surveys in the Governments Division of the U.S. Census Bureau were redesigned to reduce sample size, save resources, and improve the precision of the estimates (Cheng, Corcoran, Barth and Hogue 2009). The new design divides each of the traditional state by government-type strata with sufficiently many units into two sub-strata according to each governmental unit’s total payroll, in order to sample less from the sub-stratum with small size units. The model-assisted approach is adopted in estimating population totals. Regression estimators using auxiliary variables are obtained either within each created sub-stratum or within the original stratum by collapsing two sub-strata. A decision-based method was proposed in Cheng, Slud and Hogue (2010), applying a hypothesis test to decide which regression estimator is used within each original stratum. Consistency and asymptotic normality of these model-assisted estimators are established here, under a design-based or model-assisted asymptotic framework. Our asymptotic results also suggest two types of consistent variance estimators, one obtained by substituting unknown quantities in the asymptotic variances and the other by applying the bootstrap. The performance of all the estimators of totals and of their variance estimators are examined in some empirical studies. The U.S. Annual Survey of Public Employment and Payroll (ASPEP) is used to motivate and illustrate our study.
Release date: 2014-06-27 - Articles and reports: 12-001-X201300211888Description:
When the study variables are functional and storage capacities are limited or transmission costs are high, using survey techniques to select a portion of the observations of the population is an interesting alternative to using signal compression techniques. In this context of functional data, our focus in this study is on estimating the mean electricity consumption curve over a one-week period. We compare different estimation strategies that take account of a piece of auxiliary information such as the mean consumption for the previous period. The first strategy consists in using a simple random sampling design without replacement, then incorporating the auxiliary information into the estimator by introducing a functional linear model. The second approach consists in incorporating the auxiliary information into the sampling designs by considering unequal probability designs, such as stratified and pi designs. We then address the issue of constructing confidence bands for these estimators of the mean. When effective estimators of the covariance function are available and the mean estimator satisfies a functional central limit theorem, it is possible to use a fast technique for constructing confidence bands, based on the simulation of Gaussian processes. This approach is compared with bootstrap techniques that have been adapted to take account of the functional nature of the data.
Release date: 2014-01-15 - Articles and reports: 12-001-X200900211044Description:
In large scaled sample surveys it is common practice to employ stratified multistage designs where units are selected using simple random sampling without replacement at each stage. Variance estimation for these types of designs can be quite cumbersome to implement, particularly for non-linear estimators. Various bootstrap methods for variance estimation have been proposed, but most of these are restricted to single-stage designs or two-stage cluster designs. An extension of the rescaled bootstrap method (Rao and Wu 1988) to stratified multistage designs is proposed which can easily be extended to any number of stages. The proposed method is suitable for a wide range of reweighting techniques, including the general class of calibration estimators. A Monte Carlo simulation study was conducted to examine the performance of the proposed multistage rescaled bootstrap variance estimator.
Release date: 2009-12-23 - Articles and reports: 11-536-X200900110806Description:
Recent work using a pseudo empirical likelihood (EL) method for finite population inferences with complex survey data focused primarily on a single survey sample, non-stratified or stratified, with considerable effort devoted to computational procedures. In this talk we present a pseudo empirical likelihood approach to inference from multiple surveys and multiple-frame surveys, two commonly encountered problems in survey practice. We show that inferences about the common parameter of interest and the effective use of various types of auxiliary information can be conveniently carried out through the constrained maximization of joint pseudo EL function. We obtain asymptotic results which are used for constructing the pseudo EL ratio confidence intervals, either using a chi-square approximation or a bootstrap calibration. All related computational problems can be handled using existing algorithms on stratified sampling after suitable re-formulation.
Release date: 2009-08-11 - Articles and reports: 12-001-X200900110882Description:
The bootstrap technique is becoming more and more popular in sample surveys conducted by national statistical agencies. In most of its implementations, several sets of bootstrap weights accompany the survey microdata file given to analysts. So far, the use of the technique in practice seems to have been mostly limited to variance estimation problems. In this paper, we propose a bootstrap methodology for testing hypotheses about a vector of unknown model parameters when the sample has been drawn from a finite population. The probability sampling design used to select the sample may be informative or not. Our method uses model-based test statistics that incorporate the survey weights. Such statistics are usually easily obtained using classical software packages. We approximate the distribution under the null hypothesis of these weighted model-based statistics by using bootstrap weights. An advantage of our bootstrap method over existing methods of hypothesis testing with survey data is that, once sets of bootstrap weights are provided to analysts, it is very easy to apply even when no specialized software dealing with complex surveys is available. Also, our simulation results suggest that, overall, it performs similarly to the Rao-Scott procedure and better than the Wald and Bonferroni procedures when testing hypotheses about a vector of linear regression model parameters.
Release date: 2009-06-22 - 16. In this issue (Vol. 35, no. 1) ArchivedArticles and reports: 12-001-X200900110892Description:
In this Issue is a column where the Editor biefly presents each paper of the current issue of Survey Methodology. As well, it sometimes contain informations on structure or management changes in the journal.
Release date: 2009-06-22 - Articles and reports: 12-001-X200800210756Description:
In longitudinal surveys nonresponse often occurs in a pattern that is not monotone. We consider estimation of time-dependent means under the assumption that the nonresponse mechanism is last-value-dependent. Since the last value itself may be missing when nonresponse is nonmonotone, the nonresponse mechanism under consideration is nonignorable. We propose an imputation method by first deriving some regression imputation models according to the nonresponse mechanism and then applying nonparametric regression imputation. We assume that the longitudinal data follow a Markov chain with finite second-order moments. No other assumption is imposed on the joint distribution of longitudinal data and their nonresponse indicators. A bootstrap method is applied for variance estimation. Some simulation results and an example concerning the Current Employment Survey are presented.
Release date: 2008-12-23 - 18. Methodological challenges in analyzing nutrition data from the Canadian Community Health Survey - Nutrition ArchivedArticles and reports: 11-522-X200600110394Description:
Statistics Canada conducted the Canadian Community Health Survey - Nutrition in 2004. The survey's main objective was to estimate the distributions of Canadians' usual dietary intake at the provincial level for 15 age-sex groups. Such distributions are generally estimated with the SIDE application, but with the choices that were made concerning sample design and method of estimating sampling variability, obtaining those estimates is not a simple matter. This article describes the methodological challenges in estimating usual intake distributions from the survey data using SIDE.
Release date: 2008-03-17 - Articles and reports: 11-522-X200600110416Description:
Application of standard methods to survey data without accounting for the design features and weight adjustments can lead to erroneous inferences. Bootstrap methods offer an attractive option to the analyst for taking account of the design features and weight adjustments. The data file consists of the full-sample final weights and associated bootstrap final weights for a large number of bootstrap replicates as well as the observed data on the sample elements. We show how such data files can be used to analyze survey data in a straightforward manner using weighted estimating equations. A one-step estimating function bootstrap method that avoids some difficulties with the bootstrap is also discussed.
Release date: 2008-03-17 - 20. Disclosure risk and variance estimation ArchivedArticles and reports: 11-522-X200600110434Description:
Protecting respondents from disclosure of their identity in publicly released survey data is of practical concern to many government agencies. Methods for doing so include suppression of cluster and stratum identifiers and altering or swapping record values between respondents. Unfortunately, stratum and cluster identifiers are usually needed for variance estimation using linearization and for replication methods as resampling is typically done on first-stage sampling units within strata. One might feel that releasing a set of replicate weights that also have stratum and cluster identifiers suppressed might circumvent this problem to some extent, especially using some random resampling such as the bootstrap. In this article, we first demonstrate that by viewing the replicate weights as observations in a high dimensional space one can easily use clustering algorithms to reconstruct the cluster identifiers irrespective of the resampling method even if the resampling weights are randomly altered. We then propose a fast algorithm for swapping cluster and strata identifiers of ultimate units before creating replicate weights without significantly impacting resulting variance estimates of characteristics of interest. The methods are illustrated by application to publicly released data from the National Health and Nutrition Examination Surveys, where such disclosure issues are extremely important..
Release date: 2008-03-17
Data (0)
Data (0) (0 results)
No content available at this time.
Analysis (38)
Analysis (38) (20 to 30 of 38 results)
- 21. Efficient bootstrap for business surveys ArchivedArticles and reports: 12-001-X200700210494Description:
The Australian Bureau of Statistics has recently developed a generalized estimation system for processing its large scale annual and sub-annual business surveys. Designs for these surveys have a large number of strata, use Simple Random Sampling within Strata, have non-negligible sampling fractions, are overlapping in consecutive periods, and are subject to frame changes. A significant challenge was to choose a variance estimation method that would best meet the following requirements: valid for a wide range of estimators (e.g., ratio and generalized regression), requires limited computation time, can be easily adapted to different designs and estimators, and has good theoretical properties measured in terms of bias and variance. This paper describes the Without Replacement Scaled Bootstrap (WOSB) that was implemented at the ABS and shows that it is appreciably more efficient than the Rao and Wu (1988)'s With Replacement Scaled Bootstrap (WSB). The main advantages of the Bootstrap over alternative replicate variance estimators are its efficiency (i.e., accuracy per unit of storage space) and the relative simplicity with which it can be specified in a system. This paper describes the WOSB variance estimator for point-in-time and movement estimates that can be expressed as a function of finite population means. Simulation results obtained as part of the evaluation process show that the WOSB was more efficient than the WSB, especially when the stratum sample sizes are sometimes as small as 5.
Release date: 2008-01-03 - 22. Mean - Adjusted bootstrap for two - Phase sampling ArchivedArticles and reports: 12-001-X20070019853Description:
Two-phase sampling is a useful design when the auxiliary variables are unavailable in advance. Variance estimation under this design, however, is complicated particularly when sampling fractions are high. This article addresses a simple bootstrap method for two-phase simple random sampling without replacement at each phase with high sampling fractions. It works for the estimation of distribution functions and quantiles since no rescaling is performed. The method can be extended to stratified two-phase sampling by independently repeating the proposed procedure in different strata. Variance estimation of some conventional estimators, such as the ratio and regression estimators, is studied for illustration. A simulation study is conducted to compare the proposed method with existing variance estimators for estimating distribution functions and quantiles.
Release date: 2007-06-28 - Articles and reports: 12-001-X20060029549Description:
In this article, we propose a Bernoulli-type bootstrap method that can easily handle multi-stage stratified designs where sampling fractions are large, provided simple random sampling without replacement is used at each stage. The method provides a set of replicate weights which yield consistent variance estimates for both smooth and non-smooth estimators. The method's strength is in its simplicity. It can easily be extended to any number of stages without much complication. The main idea is to either keep or replace a sampling unit at each stage with preassigned probabilities, to construct the bootstrap sample. A limited simulation study is presented to evaluate performance and, as an illustration, we apply the method to the 1997 Japanese National Survey of Prices.
Release date: 2006-12-21 - 24. Link-tracing sampling with an initial sample of sites sequentially selected: Estimation of the population size ArchivedArticles and reports: 11-522-X20040018750Description:
This paper modifies the link-tracing sampling with a sequential sample of sites and proposes a maximum likelihood estimator or another one derived under the Bayesian approach. It proposes that confidence intervals be constructed by Bootstrap methods.
Release date: 2005-10-27 - Articles and reports: 12-002-X20050018030Description:
People often wish to use survey micro-data to study whether the rate of occurrence of a particular condition in a subpopulation is the same as the rate of occurrence in the full population. This paper describes some alternatives for making inferences about such a rate difference and shows whether and how these alternatives may be implemented in three different survey software packages. The software packages illustrated - SUDAAN, WesVar and Bootvar - all can make use of bootstrap weights provided by the analyst to carry out variance estimation.
Release date: 2005-06-23 - Articles and reports: 12-002-X20050018031Description:
This article presents revisions to a Stata "bswreg" ado file that calculates variance estimates using bootstrap weights. This revision adds new output and analytic features. The main feature added to the program enables researchers to apply mean bootstrap weights while accounting for the number of weights used to generate the average bootstrap weight. The Workplace and Employee Survey dataset will be used to illustrate the usefulness of this program. This revised version of the "bswreg" command is still an easy to use flexible tool, which is compatible with a wide variety of regression analytical techniques and datasets. The bswreg command and design-based bootstrap weights should only be used for inference when it is theoretically valid.
Release date: 2005-06-23 - Articles and reports: 11-522-X20030017720Description:
This paper examines a jackknife method proposed in Rao (2003) for estimating mean squared errors (MSEs) when generalized linear models or other non-linear models are used for the response of interest. It demonstrate the method's performance in a simulation study.
Release date: 2005-01-26 - 28. Using bootstrap weights with Wes Var and SUDAAN ArchivedArticles and reports: 12-002-X20040027032Description:
This article examines why many Statistics Canada surveys supply bootstrap weights with their microdata for the purpose of design-based variance estimation. Bootstrap weights are not supported by commercially available software such as SUDAAN and WesVar, but there are ways to use these applications to produce boostrap variance estimates.
The paper concludes with a brief discussion of other design-based approaches to variance estimation as well as software, programs and procedures where these methods have been employed.
Release date: 2004-10-05 - Articles and reports: 11-522-X20020016713Description:
This paper explores the relationship between low income and prevalence of asthma. The genetic and environmental determinants are incompletely understood. It has been observed in a previous study that Canadians with low incomes are at increased risk of asthma. Based on data from 17,605 subjects 12 years of age or older who participated in the first cycle of the National Population Health Survey (NPHS) from 1994 to 1995, males and females with low incomes had 1.44- and 1.33-fold increases, respectively, in the prevalence of asthma compared with their counterparts with high incomes. However, there was no significant difference observed between middle and high income categories. Therefore, it is not clear if there is a more systematic relationship between income adequacy and asthma occurrence. A much larger sample size of the second cycle of the NPHS allowed us to further explore if the prevalence of asthma increases with decreasing income adequacy among Canadians.
Release date: 2004-09-13 - Articles and reports: 12-001-X20040016996Description:
This article studies the use of the sample distribution for the prediction of finite population totals under single-stage sampling. The proposed predictors employ the sample values of the target study variable, the sampling weights of the sample units and possibly known population values of auxiliary variables. The prediction problem is solved by estimating the expectation of the study values for units outside the sample as a function of the corresponding expectation under the sample distribution and the sampling weights. The prediction mean square error is estimated by a combination of an inverse sampling procedure and a re-sampling method. An interesting outcome of the present analysis is that several familiar estimators in common use are shown to be special cases of the proposed approach, thus providing them a new interpretation. The performance of the new and some old predictors in common use is evaluated and compared by a Monte Carlo simulation study using a real data set.
Release date: 2004-07-14
Reference (1)
Reference (1) ((1 result))
- Surveys and statistical programs – Documentation: 11-522-X19980015017Description:
Longitudinal studies with repeated observations on individuals permit better characterizations of change and assessment of possible risk factors, but there has been little experience applying sophisticated models for longitudinal data to the complex survey setting. We present results from a comparison of different variance estimation methods for random effects models of change in cognitive function among older adults. The sample design is a stratified sample of people 65 and older, drawn as part of a community-based study designed to examine risk factors for dementia. The model summarizes the population heterogeneity in overall level and rate of change in cognitive function using random effects for intercept and slope. We discuss an unweighted regression including covariates for the stratification variables, a weighted regression, and bootstrapping; we also did preliminary work into using balanced repeated replication and jackknife repeated replication.
Release date: 1999-10-22