7. Application
Qi Dong, Michael R. Elliott and Trivellore E. Raghunathan
Previous | Next
In this
section, we use data from the 2006 National Health Interview Survey (NHIS) and
the 2006 Medical Expenditure Panel Survey (MEPS) to evaluate the performance of
the nonparametric method in a stratified clustering sampling design. The
National Health Interview Survey (NHIS) is a nationwide, face-to-face health
survey based on a stratified multistage design, with oversamples of black,
Hispanic, and elderly populations. For confidentiality purposes, the true
stratification and primary sampling unit (PSU) variables are not publicly-released;
instead pseudo-strata and PSUs (two per stratum) are released. The MEPS is a
subsample of the previous year's NHIS sample, and retains the same stratified
multistage design.
Both NHIS and
MEPS ask respondents whether they are covered by any health insurance and, if
so, what type health insurance they are using (private versus government-sponsored such as Medicare or Medicaid). We
estimate overall health insurance coverage rates as well as coverage rates in
subpopulations defined by demographic variables such as gender, race, income
level, or combinations thereof: specifically, we estimate health insurance
coverage for males, non-Hispanic whites, and non-Hispanic whites with household
income between $25,000 and $35,000 per year. We delete the cases with
item-missing values and focus on our simulation on the complete cases. This
results in 20,147 and 20,893 cases in the NHIS and MEPS data respectively.
7.1 Estimation of health insurance coverage from
the NHIS and MEPS
In this simulation study, we will use the nonparametric
method to adjust for the stratified clustering sampling used by the 2006 NHIS
and MEPS and generate synthetic populations that can be analyzed as simple
random samples. We also consider a model-based approach for generating
synthetic populations using a log-linear model for the health insurance status
by six independent demographic variables: gender, race, census region,
education level, age (categorical), and income level (categorical). Then we
evaluate the method by comparing the estimates of the health insurance coverage
rate for the whole population and selected subdomains obtained from both the
non-parametric and log-linear model synthetic populations to those obtained
from the actual data.
7.1.1 Generating nonparametric
synthetic populations
Using the nonparametric method developed in Section 3,
we generate 200 synthetic populations for each survey. Specifically, we
generate 200 BB samples and for each BB sample, we
generate 10 FPBB of size Thus, each synthetic population is 50 times as
big as the actual sample (1,007,350 for NHIS, 1,044,650 for MEPS). Each
synthetic population is analyzed as a simple random sample and the estimates
are combined as described in Section 5.
7.1.2 Generating synthetic
populations via log-linear models
In the common situation that the survey data of interest
are in the form of a multidimensional contingency table, a log-linear model
might be considered as a parametric approach to generate draws from a posterior
predictive distribution. For simplicity of exposition, assume is the variable of our interest with levels, and is a design variable with levels (e.g.,
gender, race, etc.) whose marginal
distribution is known for the population. Assume represents the cell proportion of the cell, A fully saturated log-linear model is given by
(Agresti 2002):
where is the log of the probability that
one observation falls in cell of the contingency table, is the main effect for is the main effect for and is the interaction effect for and This model includes all possible
one-way and two-way effects and thus is saturated as it has the same number of
effects as cells in the contingency table. To avoid over-fitting the data in
the example, we can consider non-saturated models that exclude some or all of
the interaction terms, choosing the model based on likelihood ratio tests or
AIC or BIC criteria.
The synthetic populations can be generated from the
posterior predictive distribution from the model. However, when the data is
collected under a complex sampling design, we are not aware of standard
statistical software that can produce both the point estimate and covariance
estimate of the regression coefficients. Instead, we have to use a jackknife
replication method to adjust for stratification, clustering and weighting.
Specifically, the parametric synthetic populations can be generated from the
following steps:
1. Estimate
coefficients and covariance matrix:
Under the selected model (assume the two-dimensional saturated model here just for illustration),
estimate the coefficients
and the covariance matrix of the estimates after taking into account the complex design
features using jackknife repeated replication (JRR):
- For
each replication, withdraw one cluster, and inflate the weights for the
respondents in the other clusters within the same stratum by (replication weights), where denotes the number of clusters within stratum Assume we have clusters in total, then we have replications. For each replication, we fit the
log-linear model and obtain the maximum likelihood estimates (MLE) of the
coefficients,
- For
each replication, use the replication weights to fit the log-linear model.
Specifically, use the replication weights to calculate the size of each cell of
the contingency table, which is used to fit the log-linear model. We denote the
MLE for the replication by a column vector, for stratum Notice that is a by 1 column vector. We denote Similarly, are also by 1 column vectors denoted by
The MLE of the coefficients can be obtained by For the by covariance matrix, the jackknife replication
estimate of the element is the covariance between the and coefficients, which is given by:
where and This gives us the correct
variance estimate of
2.
Approximate the posterior distribution of the coefficients:
Let denote the Cholesky decomposition such that Generate a vector of random normal deviates and define
3. Impute the
unobserved values of the population:
Suppose draws, are made from the approximate posterior
distribution of For each
we can
generate one synthetic table using the assumed model:
Once the cell
proportions are determined, we can generate the synthetic table of any size.
The results below are based on a seven-dimension
contingency table (see Table 7.1 for the specific covariate categories). BIC
measures indicated that a model with all 2-way but no 3-way interactions
provided the most parsimonious fit.
Table 7.1
Variables and response categories for the 2006 NHIS and MEPS used in log-linear model.
| Variables of Interest |
Response Categories |
| Age |
1: [18; 24]; 2: [25; 34]; 3: [35; 44]; 4: [45; 54]; 5: [55; 64]; 6: >= 65 |
| Census Region |
1: Northeast; 2: Midwest; 3: South; 4: West |
| Education |
1: Less than high school; 2: High school; 3: Some college; 4: College |
| Gender |
1: Male; 2: Female |
| Health Insurance Coverage |
1: Any Private Insurance; 2: Public Insurance; 3: Uninsured |
| Income |
1: (0; 10,000); 2: [10,000; 15,000); 3: [15,000; 20,000); 4: [20,000; 25,000); 5: [25,000; 35,000); 6: [35,000; 75,000); 7: >= 75,000 |
| Race |
1: Hispanic; 2: Non-Hispanic White; 3: Non-Hispanic Black; 4: Non-Hispanic All other race groups |
7.2 Results
The results are summarized in Table 7.2. For the total
population and the larger subpopulations, we can see that the point estimates
(posterior mean) of health insurance rates are the same for both the
nonparametric and log-linear approach, and are almost identical to those
obtained from the actual data after complex sampling design features are
accounted for. Both methods yield synthetic populations with slightly higher
(posterior) variances than the actual data, reflecting the information loss in
the synthesis. In the NHIS, the loss for the non-parametric estimator averaged
a little over 20% and was slightly greater than for the log-linear model, which
averaged around 10%. Both had losses of about 10% over the actual data in MEPS.
However, for the smaller subpopulation (non-Hispanic whites earning
$25,000-$35,000 per year), the log-linear model produced biased results, due to
the fact that the log-linear model did not include all possible interactions. The
nonparametric method yields estimates almost identical to those obtained from
the actual data after complex sampling design features are accounted for. The
log-linear model also substantially underestimated the variance of insurance
coverage by 30-40% in these cells, versus
an overestimation in the nonparametric approach of 10-40%.
Table 7.2
Estimates from actual data and from the synthetic populations (Nonparametric and log-linear model) for the 2006 NHIS and MEPS.
Table summary
This table displays the estimates from actual data and from the synthetic populations (Nonparametric and log-linear model) for the 2006 NHIS and MEPS. The information is grouped by domain (appearing as row headers), Actual Data (Complex Design), Synthetic Populations (appearing as column headers).
|
Domain
|
Actual Data (Complex Design)
|
Synthetic Populations
|
|
Nonparametric
|
Log-linear Model
|
|
Types
|
NHIS
|
MEPS
|
NHIS
|
MEPS
|
NHIS
|
MEPS
|
|
Whole Population
|
Proportion
|
|
Private
|
0.746
|
0.735
|
0.746
|
0.736
|
0.746
|
0.734
|
|
Public
|
0.075
|
0.133
|
0.075
|
0.132
|
0.076
|
0.133
|
|
Uninsured
|
0.179
|
0.132
|
0.179
|
0.132
|
0.178
|
0.132
|
|
Variance
|
|
Private
|
2.46E-05
|
2.78E-05
|
3.15E-05
|
3.31E-05
|
2.66E-05
|
2.86E-05
|
|
Public
|
6.29E-06
|
1.44E-05
|
8.06E-06
|
1.59E-05
|
7.99E-06
|
1.77E-05
|
|
Uninsured
|
1.84E-05
|
1.41E-05
|
2.29E-05
|
1.71E-05
|
1.81E-05
|
1.56E-05
|
|
Male
|
Proportion
|
|
Private
|
0.74
|
0.735
|
0.74
|
0.736
|
0.74
|
0.735
|
|
Public
|
0.06
|
0.101
|
0.06
|
0.1
|
0.06
|
0.102
|
|
Uninsured
|
0.2
|
0.164
|
0.2
|
0.164
|
0.2
|
0.164
|
|
Variance
|
|
Private
|
3.32E-05
|
3.87E-05
|
3.93E-05
|
4.31E-05
|
3.70E-05
|
3.52E-05
|
|
Public
|
6.82E-06
|
1.53E-05
|
8.81E-06
|
1.63E-05
|
7.91E-06
|
1.91E-05
|
|
Uninsured
|
2.94E-05
|
2.64E-05
|
3.29E-05
|
2.79E-05
|
3.19E-05
|
2.56E-05
|
|
Non-Hispanic White
|
Proportion
|
|
Private
|
0.805
|
0.788
|
0.804
|
0.788
|
0.804
|
0.788
|
|
Public
|
0.062
|
0.116
|
0.062
|
0.116
|
0.062
|
0.117
|
|
Uninsured
|
0.134
|
0.096
|
0.134
|
0.096
|
0.134
|
0.096
|
|
Variance
|
|
Private
|
2.99E-05
|
3.35E-05
|
3.79E-05
|
4.12E-05
|
3.07E-05
|
3.98E-05
|
|
Public
|
8.20E-06
|
1.81E-05
|
1.04E-05
|
2.00E-05
|
1.10E-05
|
2.45E-05
|
|
Uninsured
|
2.02E-05
|
1.51E-05
|
2.35E-05
|
1.80E-05
|
1.82E-05
|
1.82E-05
|
Non-Hispanic White & Income [25,000; 35,000)
|
Proportion
|
|
Private
|
0.827
|
0.813
|
0.827
|
0.814
|
0.84
|
0.838
|
|
Public
|
0.039
|
0.079
|
0.039
|
0.079
|
0.037
|
0.067
|
|
Uninsured
|
0.134
|
0.108
|
0.134
|
0.107
|
0.122
|
0.096
|
|
Variance
|
|
Private
|
1.00E-04
|
1.39E-04
|
1.48E-04
|
1.63E-04
|
6.80E-05
|
8.59E-05
|
|
Public
|
2.82E-05
|
6.31E-05
|
3.86E-05
|
7.28E-05
|
1.79E-05
|
4.25E-05
|
|
Uninsured
|
7.24E-05
|
8.92E-05
|
9.55E-05
|
1.11E-04
|
4.38E-05
|
5.79E-05
|
Previous | Next