Keyword search
Results
All (7)
All (7) ((7 results))
- Articles and reports: 12-001-X20030016605Description:
In this paper, we examine the effects of model choice on different types of estimators for totals of domains (including small domains or small areas) for a sampled finite population. The paper asks how different estimator types compare for a common underlying model statement. We argue that estimator type - synthetic, generalized regression (GREG), composite, empirical best linear unbiased predicition (EBLUP), hierarchical Bayes, and so on - is one important aspect of domain estimation, and that the choice of the model, including its parameters and effects, is a second aspect, conceptually different from the first. Earlier work has not always made this distinction clear. For a given estimator type, one can derive different estimators, depending on the choice of model. In recent literature, a number of estimator types have been proposed, but there is relatively little impartial comparisons made among them. In this paper, we discuss three types: synthetic, GREG, and, to a limited extent, composite. We show that model improvement - the transition from a weaker to a stronger model - has very different effects on the different estimator types. We also show that the difference in accuracy between the different estimator types depends on the choice of model. For a well-specified model, the difference in accuracy between synthetic and GREG is negligible, but it can be substantial if the model is mis-specified. The synthetic type then tends to be highly inaccurate. We rely partly on theoretical results (for simple random sampling only) and partly on empirical results. The empirical results are based on simulations with repeated samples drawn from two finite populations, one artificially constructed, the other constructed from the real data of the Finnish Labour Force Survey.
Release date: 2003-07-31 - Articles and reports: 12-001-X20030016613Description:
The Illinois Department of Employment Security is using small domain estimation techniques to estimate employment at the county or industry divisional level. The estimator is a standard synthetic estimator, based on the ability to match Current Employment Statistics sample data to ES202 administrative records and an assumed model relationship between the two data sources. This paper is a case study that reviews the steps taken to evaluate the appropriateness of the model and the difficulties encountered in linking the two data sources.
Release date: 2003-07-31 - 3. Alternative Work Practices and Quit Rates: Methodological Issues and Empirical Evidence for Canada ArchivedArticles and reports: 11F0019M2003199Geography: CanadaDescription:
Using a nationally representative sample of establishments, we have examined whether selected alternative work practices (AWPs) tend to reduce quit rates. Overall, our analysis provides strong evidence of a negative association between these AWPs and quit rates among establishments of more than 10 employees operating in high-skill services. We also found some evidence of a negative association in low-skill services. However, the magnitude of this negative association was reduced substantially when we added an indicator of whether the workplace has a formal policy of information sharing. There was very little evidence of a negative association in manufacturing. While establishments with self-directed workgroups have lower quit rates than others, none of the bundles of work practices considered yielded a negative and statistically significant effect. We surmise that key AWPs might be more successful in reducing labour turnover in technologically complex environments than in low-skill ones.
Release date: 2003-03-17 - 4. A hierarchical Bayesian nonignorable nonresponse model for multinomial data from small areas ArchivedArticles and reports: 12-001-X20020026428Description:
The analysis of survey data from different geographical areas where the data from each area are polychotomous can be easily performed using hierarchical Bayesian models, even if there are small cell counts in some of these areas. However, there are difficulties when the survey data have missing information in the form of non-response, especially when the characteristics of the respondents differ from the non-respondents. We use the selection approach for estimation when there are non-respondents because it permits inference for all the parameters. Specifically, we describe a hierarchical Bayesian model to analyse multinomial non-ignorable non-response data from different geographical areas; some of them can be small. For the model, we use a Dirichlet prior density for the multinomial probabilities and a beta prior density for the response probabilities. This permits a 'borrowing of strength' of the data from larger areas to improve the reliability in the estimates of the model parameters corresponding to the smaller areas. Because the joint posterior density of all the parameters is complex, inference is sampling-based and Markov chain Monte Carlo methods are used. We apply our method to provide an analysis of body mass index (BMI) data from the third National Health and Nutrition Examination Survey (NHANES III). For simplicity, the BMI is categorized into 3 natural levels, and this is done for each of 8 age-race-sex domains and 34 counties. We assess the performance of our model using the NHANES III data and simulated examples, which show our model works reasonably well.
Release date: 2003-01-29 - 5. A generalization of the Lavallée and Hidiroglou algorithm for stratification in business Surveys ArchivedArticles and reports: 12-001-X20020026432Description:
This paper suggests stratification algorithms that account for a discrepancy between the stratification variable and the study variable when planning a stratified survey design. Two models are proposed for the change between these two variables. One is a log-linear regression model; the other postulates that the study variable and the stratification variable coincide for most units, and that large discrepancies occur for some units. Then, the Lavallée and Hidiroglou (1988) stratification algorithm is modified to incorporate these models in the determination of the optimal sample sizes and of the optimal stratum boundaries for a stratified sampling design. An example illustrates the performance of the new stratification algorithm. A discussion of the numerical implementation of this algorithm is also presented.
Release date: 2003-01-29 - Articles and reports: 12-001-X20020026434Description:
In theory, it is customary to define general regression estimators in terms of full-rank weighting models (i.e., the design matrix that corresponds to the weighting model is of full rank). For such weighting models, it is well known that the general regression weights reproduce the (known) population totals of the auxiliary variables involved. In practice, however, the weighting model often is not of full rank, especially when the weighting model is for incomplete post-stratification. By means of the theory of generalized inverse matrices, it is shown under which circumstances this consistency property remains valid. In this paper,, we discuss the non-trivial example of consistent weighting between persons and households as proposed by Lemaître and Dufour (1987). We then show how the theory is implemented in Bascula.
Release date: 2003-01-29 - Articles and reports: 12-001-X20020029058Description:
Linearization (or Taylor series) methods are widely used to estimate standard errors for the co-efficients of linear regression models fit to multi-stage samples. When the number of primary sampling units (PSUs) is large, linearization can produce accurate standard errors under quite general conditions. However, when the number of PSUs is small or a co-efficient depends primarily on data from a small number of PSUs, linearization estimators can have large negative bias.
In this paper, we characterize features of the design matrix that produce large bias in linearization standard errors for linear regression co-efficients. We then propose a new method, bias reduced linearization (BRL), based on residuals adjusted to better approximate the covariance of the true errors. When the errors are independent and identically distributed (i.i.d.), the BRL estimator is unbiased for the variance. Furthermore, a simulation study shows that BRL can greatly reduce the bias, even if the errors are not i.i.d. We also propose using a Satterthwaite approximation to determine the degrees of freedom of the reference distribution for tests and confidence intervals about linear combinations of co-efficients based on the BRL estimator. We demonstrate that the jackknife estimator also tends to be biased in situations where linearization is biased. However, the jackknife's bias tends to be positive. Our bias-reduced linearization estimator can be viewed as a compromise between the traditional linearization and jackknife estimators.
Release date: 2003-01-29
Data (0)
Data (0) (0 results)
No content available at this time.
Analysis (7)
Analysis (7) ((7 results))
- Articles and reports: 12-001-X20030016605Description:
In this paper, we examine the effects of model choice on different types of estimators for totals of domains (including small domains or small areas) for a sampled finite population. The paper asks how different estimator types compare for a common underlying model statement. We argue that estimator type - synthetic, generalized regression (GREG), composite, empirical best linear unbiased predicition (EBLUP), hierarchical Bayes, and so on - is one important aspect of domain estimation, and that the choice of the model, including its parameters and effects, is a second aspect, conceptually different from the first. Earlier work has not always made this distinction clear. For a given estimator type, one can derive different estimators, depending on the choice of model. In recent literature, a number of estimator types have been proposed, but there is relatively little impartial comparisons made among them. In this paper, we discuss three types: synthetic, GREG, and, to a limited extent, composite. We show that model improvement - the transition from a weaker to a stronger model - has very different effects on the different estimator types. We also show that the difference in accuracy between the different estimator types depends on the choice of model. For a well-specified model, the difference in accuracy between synthetic and GREG is negligible, but it can be substantial if the model is mis-specified. The synthetic type then tends to be highly inaccurate. We rely partly on theoretical results (for simple random sampling only) and partly on empirical results. The empirical results are based on simulations with repeated samples drawn from two finite populations, one artificially constructed, the other constructed from the real data of the Finnish Labour Force Survey.
Release date: 2003-07-31 - Articles and reports: 12-001-X20030016613Description:
The Illinois Department of Employment Security is using small domain estimation techniques to estimate employment at the county or industry divisional level. The estimator is a standard synthetic estimator, based on the ability to match Current Employment Statistics sample data to ES202 administrative records and an assumed model relationship between the two data sources. This paper is a case study that reviews the steps taken to evaluate the appropriateness of the model and the difficulties encountered in linking the two data sources.
Release date: 2003-07-31 - 3. Alternative Work Practices and Quit Rates: Methodological Issues and Empirical Evidence for Canada ArchivedArticles and reports: 11F0019M2003199Geography: CanadaDescription:
Using a nationally representative sample of establishments, we have examined whether selected alternative work practices (AWPs) tend to reduce quit rates. Overall, our analysis provides strong evidence of a negative association between these AWPs and quit rates among establishments of more than 10 employees operating in high-skill services. We also found some evidence of a negative association in low-skill services. However, the magnitude of this negative association was reduced substantially when we added an indicator of whether the workplace has a formal policy of information sharing. There was very little evidence of a negative association in manufacturing. While establishments with self-directed workgroups have lower quit rates than others, none of the bundles of work practices considered yielded a negative and statistically significant effect. We surmise that key AWPs might be more successful in reducing labour turnover in technologically complex environments than in low-skill ones.
Release date: 2003-03-17 - 4. A hierarchical Bayesian nonignorable nonresponse model for multinomial data from small areas ArchivedArticles and reports: 12-001-X20020026428Description:
The analysis of survey data from different geographical areas where the data from each area are polychotomous can be easily performed using hierarchical Bayesian models, even if there are small cell counts in some of these areas. However, there are difficulties when the survey data have missing information in the form of non-response, especially when the characteristics of the respondents differ from the non-respondents. We use the selection approach for estimation when there are non-respondents because it permits inference for all the parameters. Specifically, we describe a hierarchical Bayesian model to analyse multinomial non-ignorable non-response data from different geographical areas; some of them can be small. For the model, we use a Dirichlet prior density for the multinomial probabilities and a beta prior density for the response probabilities. This permits a 'borrowing of strength' of the data from larger areas to improve the reliability in the estimates of the model parameters corresponding to the smaller areas. Because the joint posterior density of all the parameters is complex, inference is sampling-based and Markov chain Monte Carlo methods are used. We apply our method to provide an analysis of body mass index (BMI) data from the third National Health and Nutrition Examination Survey (NHANES III). For simplicity, the BMI is categorized into 3 natural levels, and this is done for each of 8 age-race-sex domains and 34 counties. We assess the performance of our model using the NHANES III data and simulated examples, which show our model works reasonably well.
Release date: 2003-01-29 - 5. A generalization of the Lavallée and Hidiroglou algorithm for stratification in business Surveys ArchivedArticles and reports: 12-001-X20020026432Description:
This paper suggests stratification algorithms that account for a discrepancy between the stratification variable and the study variable when planning a stratified survey design. Two models are proposed for the change between these two variables. One is a log-linear regression model; the other postulates that the study variable and the stratification variable coincide for most units, and that large discrepancies occur for some units. Then, the Lavallée and Hidiroglou (1988) stratification algorithm is modified to incorporate these models in the determination of the optimal sample sizes and of the optimal stratum boundaries for a stratified sampling design. An example illustrates the performance of the new stratification algorithm. A discussion of the numerical implementation of this algorithm is also presented.
Release date: 2003-01-29 - Articles and reports: 12-001-X20020026434Description:
In theory, it is customary to define general regression estimators in terms of full-rank weighting models (i.e., the design matrix that corresponds to the weighting model is of full rank). For such weighting models, it is well known that the general regression weights reproduce the (known) population totals of the auxiliary variables involved. In practice, however, the weighting model often is not of full rank, especially when the weighting model is for incomplete post-stratification. By means of the theory of generalized inverse matrices, it is shown under which circumstances this consistency property remains valid. In this paper,, we discuss the non-trivial example of consistent weighting between persons and households as proposed by Lemaître and Dufour (1987). We then show how the theory is implemented in Bascula.
Release date: 2003-01-29 - Articles and reports: 12-001-X20020029058Description:
Linearization (or Taylor series) methods are widely used to estimate standard errors for the co-efficients of linear regression models fit to multi-stage samples. When the number of primary sampling units (PSUs) is large, linearization can produce accurate standard errors under quite general conditions. However, when the number of PSUs is small or a co-efficient depends primarily on data from a small number of PSUs, linearization estimators can have large negative bias.
In this paper, we characterize features of the design matrix that produce large bias in linearization standard errors for linear regression co-efficients. We then propose a new method, bias reduced linearization (BRL), based on residuals adjusted to better approximate the covariance of the true errors. When the errors are independent and identically distributed (i.i.d.), the BRL estimator is unbiased for the variance. Furthermore, a simulation study shows that BRL can greatly reduce the bias, even if the errors are not i.i.d. We also propose using a Satterthwaite approximation to determine the degrees of freedom of the reference distribution for tests and confidence intervals about linear combinations of co-efficients based on the BRL estimator. We demonstrate that the jackknife estimator also tends to be biased in situations where linearization is biased. However, the jackknife's bias tends to be positive. Our bias-reduced linearization estimator can be viewed as a compromise between the traditional linearization and jackknife estimators.
Release date: 2003-01-29
Reference (0)
Reference (0) (0 results)
No content available at this time.