Weighting and estimation
Filter results by
Search HelpKeyword(s)
Type
Survey or statistical program
- Survey of Labour and Income Dynamics (5)
- Census of Population (5)
- Survey of Household Spending (2)
- Longitudinal and International Study of Adults (2)
- Survey of Employment, Payrolls and Hours (1)
- Canadian Cancer Registry (1)
- Canadian Community Health Survey - Annual Component (1)
- Uniform Crime Reporting Survey (1)
- Quarterly Demographic Estimates (1)
- Annual Demographic Estimates: Canada, Provinces and Territories (1)
- Estimates of the number of census families for July 1st, Canada, provinces and territories (1)
- Annual Demographic Estimates : Subprovincial Areas (1)
- Labour Force Survey (1)
- Longitudinal Administrative Databank (1)
- General Social Survey - Social Identity (1)
- Canadian Community Health Survey - Nutrition (1)
- Canadian Income Survey (1)
- Residential Property Values (1)
- Canadian Survey on Business Conditions (1)
Results
All (638)
All (638) (20 to 30 of 638 results)
- Articles and reports: 12-001-X202500100003Description: In recent years, there has been a significant interest in machine learning in national statistical offices. Thanks to their flexibility, these methods may prove useful at the nonresponse treatment stage. In this article, we conduct an empirical investigation in order to compare several machine learning procedures in terms of bias and efficiency. In addition to the classical machine learning procedures, we assess the performance of ensemble approaches that make use of different machine learning procedures to produce a set of weights adjusted for nonresponse.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100005Description: In this paper, we derive a second-order unbiased (or nearly unbiased) mean squared prediction error (MSPE) estimator of the empirical best linear unbiased predictor (EBLUP) of a small area mean for a semi-parametric extension to the well-known Fay-Herriot model. Specifically, we derive our MSPE estimator essentially assuming certain moment conditions on both the sampling errors and random effects distributions. The normality-based Prasad-Rao MSPE estimator has a surprising robustness property in that it remains second-order unbiased under the non-normality of random effects when a simple Prasad-Rao method-of-moments estimator is used for the variance component and the sampling error distribution is normal. We show that the normality-based MSPE estimator is no longer second-order unbiased when the sampling error distribution has non-zero kurtosis or when the Fay-Herriot moment method is used to estimate the variance component, even when the sampling error distribution is normal. Interestingly, when the simple method-of moments estimator is used for the variance component, our proposed MSPE estimator does not require the estimation of kurtosis of the random effects. Results of a simulation study on the accuracy of the proposed MSPE estimator, under non-normality of both sampling and random effects distributions, are also presented.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100006Description: Survey practitioners have increasingly embraced the benefits of modern machine learning techniques, including classification and regression tree algorithms, in the development of nonresponse adjustments. These methods, which do not require a predefined functional relationship between outcomes and predictors, offer a practical means of conducting variable selection and deriving interpretable structures that link response propensity with explanatory variables. However, when applying these algorithms to survey data, it is common to overlook crucial factors like sampling weights, as well as sample design features such as stratification and clustering. To bridge this shortcoming, we propose an extension of the Chi-square Automatic Interaction Detector (CHAID) approach, and we describe the design-based asymptotic properties of the resulting “survey CHAID” (sCHAID) method. To facilitate the practical use of sCHAID, we incorporate a Rao-Scott correction into the splitting criterion, accounting for the survey design. Using data from the U.S. American Community Survey, we illustrate the use of the method and evaluate its performance through comparisons with existing weighted and unweighted algorithms.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100007Description: We introduce a novel approach to model-assisted calibration estimation in survey sampling using generalized entropy. The method builds upon recent work by Kwon, Kim and Qiu (2024) and extends it to a model-assisted framework. Unlike traditional calibration techniques, this approach employs a generalized entropy function as the objective for optimization and incorporates a debiasing calibration constraint to ensure design consistency. The proposed estimator is shown to be asymptotically equivalent to an augmented generalized regression (GREG) estimator. It allows for unequal model variance, potentially improving efficiency when the sampling design is informative. The paper presents both design-based and model-based justifications for the method, along with asymptotic properties and variance estimation techniques. Computational aspects are discussed, including an unconstrained optimization approach that facilitates implementation, especially for high-dimensional auxiliary variables. The method’s performance is evaluated through a simulation study, demonstrating its effectiveness in improving estimation efficiency, particularly when the sampling design is informative.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100008Description: Tightened budgets, continuing decrease of response rates in traditional probability surveys and increasing pressure by users for more timely data, has stimulated research on the use of nonprobability sample data, such as administrative records, web scraping, mobile phone data and voluntary internet surveys, for inference on finite population parameters like means and totals. These data are often easier, faster and cheaper to collect than traditional probability samples. However, a major concern with the use of this kind of data for official statistics is their nonrepresentativeness due to possible selection bias, which if not accounted for properly, could bias the inference. In this article, we review and discuss methods considered in the literature to deal with this problem and propose new methods, distinguishing between methods based on integration of the nonprobability sample with an appropriate probability sample, and methods that base the inference solely on the nonprobability sample. Empirical illustrations, based on simulated data are provided.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100010Description: The discussants highlight promising research topics for improving the quality and granularity of estimates from surveys. We agree that continued research is needed to evaluate models used for inference, and suggest development of measures of model dependence.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100011Description: This discussion examines some advancements in survey design and estimation, inspired by the comprehensive appraisal of Professors Jon Rao and Sharon Lohr on current trends in the field. It delves into three specific areas: balanced sampling, calibration, and small area estimation. Probabilistic balanced sampling methods, such as the cube method and penalized balanced sampling, are explored, with an emphasis on addressing emerging challenges, including extensions to linear mixed models, nonparametric regression models, and spatially balanced designs. Calibration is discussed using a modular framework that incorporates modern regression techniques, and highlights innovative uses of model calibration for data editing and causal inference. Small area estimation is considered in the context of latent variable modeling and data integration, emphasizing its role when the variable(s) of interest cannot be measured either directly or without error. Applications in integrating probability and non-probability data and conducting causal analysis at local level are also discussed.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100012Description: In this discussion, we complement the excellent overview by Profs. Lohr and Rao with some additional topics. The first topic is a call for more recognition of the central role of modeling in survey estimation. The second is a brief discussion of the use of partial frame information in survey design. Finally, we draw the attention to recent increases of synthetic methods, in particular, multilevel regression and poststratification (MRP) in small area estimation applications.Release date: 2025-06-30
- Articles and reports: 12-001-X202500100014Description: Rao (1999) summarized trends in sample survey theory and methods at the turn of the millenium. We provide an updated discussion of some current trends in survey design and estimation methods for the 50th anniversary of Survey Methodology. Recent innovations in survey design include research on anticipating nonsampling errors at the design stage and development of balanced and adaptive sampling designs to take advantage of detailed sampling frame information or data gathered during the survey process. Nonparametric and machine learning methods are increasingly used for data editing as well as for model-assisted estimation and nonresponse adjustments. Small area models have been expanded to incorporate spatial and time series information, increase the flexibility and robustness of the linking and variance models, benchmark to large-area direct estimators, and (for unit level models) account for informative sampling designs. The increasing availability of large administrative datasets, sensor and satellite data, and convenience samples has spurred research on how to use these sources - on their own and when integrated with probability samples. We conclude by discussing some frontiers for survey research.Release date: 2025-06-30
- Articles and reports: 12-001-X202400200004Description: While we avoid specifying the parametric relationship between the study variable and covariates, we illustrate the advantage of including a spatial component to better account for the covariates in our models to make Bayesian predictive inference. We treat each unique covariate combination as an individual stratum, then we use small area estimation techniques to make inference about the finite population mean of the continuous response variable. The two spatial models used are the conditional autoregressive and simple conditional autoregressive models. We include the spatial effects by creating the adjacency matrix via the Mahalanobis distance between covariates. We also show how to incorporate survey weights into the spatial models when dealing with probability survey data. We compare the results of two non-spatial models including the Scott-Smith model and the Battese, Harter, and Fuller model to the spatial models. We illustrate the comparison between the aforementioned models with an application using BMI data from eight counties in California. Our goal is to have neighboring strata yield similar predictions, and to increase the difference between strata that are not neighbors. Ultimately, using the spatial models shows less global pooling compared to the non-spatial models, which was the desired outcome.Release date: 2024-12-20
- Previous Go to previous page of All results
- 1 Go to page 1 of All results
- 2 Go to page 2 of All results
- 3 (current) Go to page 3 of All results
- 4 Go to page 4 of All results
- 5 Go to page 5 of All results
- 6 Go to page 6 of All results
- 7 Go to page 7 of All results
- ...
- 64 Go to page 64 of All results
- Next Go to next page of All results
Data (0)
Data (0) (0 results)
No content available at this time.
Analysis (610)
Analysis (610) (0 to 10 of 610 results)
- Articles and reports: 12-001-X202600100003Description: Probability-proportional-to-size sampling is widely used by national statistical offices. Here population units are selected with probabilities proportional to an auxiliary variable. Variance formulas in such designs require both first- and second-order inclusion probabilities. The computation of second-order inclusion probabilities is particularly challenging for large populations, and has been the subject of extensive research. This article presents some new exact and approximation formulas for second-order inclusion probabilities in randomized systematic sampling with unequal probabilities and without replacement.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100005Description: Confidence intervals are very often constructed based on a probability distribution that uses a certain number of degrees of freedom as a parameter. This is the case with the Student and the modified Wilson confidence intervals, discussed in this article, which use quantiles from the Student distribution where the number of degrees of freedom is generally unknown. For the length of a confidence interval to be representative of the reliability of an estimate, the actual coverage rate must match the nominal rate. To that end, the number of degrees of freedom in the probability distribution used in practice to calculate the confidence interval must be estimated as precisely as possible. An approximate rule is often used, although it tends to overestimate the actual number of degrees of freedom. In this article, a more precise version of degrees of freedom, derived from the Satterthwaite approximation, is obtained in the context of the Canadian Census of Population. The sampling design is equivalent to a simple random design without replacement, cluster-stratified, and the variance estimation method is an adaptation of the balanced repeated replication method. An explicit expression of the degrees of freedom is obtained under these conditions, enabling the factors influencing them to be identified. For comparison, the degree of freedom formula is also established for the conventional variance estimator. A simulation study shows that using this version of degrees of freedom corrects the undercoverage problem observed with the approximate rule, showing the importance of accurately assessing this number.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100009Description: Combining estimates from independent surveys via inverse-variance weights can lead to negative bias when unknown variances are estimated and the target variable is non-negative and positively skewed. In such cases, strong positive correlations typically arise between the estimators and their corresponding variance estimators, causing standard linear combinations with inverse-variance weights to exhibit negative bias. We introduce a strikingly simple method to reduce bias: replace the standard weight with the ratio of the estimator to the variance estimator. Under a linear model linking the two, we show that the new ratio-weighted estimator is approximately unbiased, whereas the conventional inverse-variance combination exhibits downward bias. Through simulations, we demonstrate that the new method brings both the bias and the mean squared error closer to the optimum for a wide range of different target variables. As our method uses only standardly reported summary statistics, it can be immediately adopted to reduce this widespread bias and improve the reliability of scientific findings in various fields.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100010Description: With the exception of two-phase sampling, the standard variance approximation of the generalized regression (GREG) estimator assumes that the population totals in the weighting scheme are observed without error. If the weighting model of the GREG estimator contains population totals that are observed with measurement error sources other than the sampling error of first-phase estimates, then this uncertainty will be ignored by the variance approximation of the GREG estimator. This paper proposes a variance approximation for the GREG estimator that accounts for additional uncertainty arising from measurement error in one or more of the population totals used in the weighting scheme. This approach has been developed for, and is being applied to, the Dutch Labour Force Survey (DLFS). The monthly publications of the DLFS are obtained with a time series model, which corrects for rotation group bias and discontinuities caused by major redesigns and the loss of face-to-face interviews during COVID-19. The GREG estimates for the quarterly figures are benchmarked to the average of the monthly publications to enforce numerical consistency between monthly and quarterly publication tables. The standard variance approximation of the GREG estimator assumes that these population totals are observed without error. This results in an underestimation of the variance of the GREG estimator. The variance approximation proposed in this paper results in more realistic standard errors for the quarterly GREG estimates.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100011Description: We construct a hybrid Bayesian method, which includes a differentially private mechanism, to mask Census county totals for a U.S. state on acreage of a commodity. We use surrogates for data collected at the farm level from a past U.S. Census of Agriculture to illustrate our procedure. We use two Bayesian small area models (parametric and mixture) to accommodate the smaller counties with fewer farms and some counties with large acres. In these models, the Laplace distribution provides a differentially private mechanism. In pre-processing, we also incorporate the Census weights to form the observed total acreage, a scaling factor to the Laplace mechanism for each county, a square-root transformation of the observed total acreage to avoid negative masked estimates especially for small counties, and the p-percent rule and the 3+ rule to partition the counties into suppressed counties, non-sensitive counties and sensitive counties. Because of difficulties in specifying and tuning the privacy budget (an unknown parameter), to balance security and utility, we specify a prior for the privacy budget, where the values are not specified, and the Gibbs sampler is used to fit the hierarchical Bayesian models. In post-processing, we use Bayesian predictive inference to obtain masked county acreages, and this includes a benchmarking so that the masked state total matches the observed state total. As a measure of reliability of the Bayesian procedure, we use the posterior coefficients of variation for the masked posterior means of the counties. As a measure of utility, we use the absolute relative errors for the individual counties, together with other global measures. For the sensitive counties, there are some differences between the two small area models but both are much better than an individual area model; the mixture model being the best compromise for security and utility.Release date: 2026-06-29
- Articles and reports: 12-001-X202600100012Description: We propose small area estimators of general indicators in off-census years, which avoid the use of deprecated census microdata, but are nearly optimal in census years. The procedure is based on replacing the obsolete census file with a larger unit-level survey that adequately covers the areas of interest and contains the values of useful auxiliary variables. However, the minimal data requirement of the proposed method is a single survey with microdata on the target variable and suitable auxiliary variables for the period of interest. We also develop an estimator of the mean squared error (MSE) that accounts for the uncertainty introduced by the large survey used to replace the census of auxiliary information. Our empirical results indicate that the proposed predictors perform clearly better than the alternative predictors when census data are outdated, and are very close to optimal ones when census data are correct. They also illustrate that the proposed total MSE estimator corrects for the bias of purely model-based MSE estimators that do not account for the large survey uncertainty.Release date: 2026-06-29
- Articles and reports: 12-001-X202500200001Description: Nested error regression models are commonly used to incorporate unit specific auxiliary variables to improve small area estimates. When the mean structure of the model is misspecified, the design-based mean squared prediction error (MSPE) of Empirical Best Linear Unbiased Predictors (EBLUP) generally increases. The Observed Best Prediction (OBP) method has been proposed with the intent to improve on the design-based MSPE over EBLUP. In this paper, we conduct a Monte Carlo simulation experiments to understand the effect of misspsecification of mean structures on different small area estimators. Our findings suggest that the OBP using unit-level auxiliary variables does not outperform the EBLUP in terms of design-based MSPE, unless the number of small areas m is extremely large. Conversely, the performance of OBP significantly improves when area-level auxiliary variables are employed. This paper includes both analytical and numerical evidence to demonstrate these observations, providing practical insights for addressing model misspecification in small area estimation (SAE).Release date: 2025-12-23
- Articles and reports: 12-001-X202500200003Description: In this paper a model-based inference procedure based on a multivariate structural time series model is developed for the production of monthly figures about consumer confidence. The input for the model are five series of direct estimates for the indices that measure consumer confidence, which are derived from the Dutch Consumer Survey. The model improves the accuracy of the direct estimates, since it provides a better separation of measurement errors and sampling errors from estimated target parameters. The standard errors for the month-to-month changes are clearly smaller under the time series model. A second problem addressed in this paper is related to the transition to a new survey process in 2017. Structural time series models in combination with a parallel run are applied to estimate discontinuities induced by the redesign. An algorithm designed for the consumer confidence variables is developed to construct uninterrupted input series for the aforementioned structural time series model. This inference method facilitated a smooth transition to a new survey design and resulted in uninterrupted series about consumer confidence that date back to 1986. The method is implemented for the production of official monthly figures on consumer confidence in the Netherlands.Release date: 2025-12-23
- Articles and reports: 12-001-X202500200005Description: The use of non-probability data sources for statistical purposes and for official statistics has become increasingly popular in recent years. However, statistical inference based on non-probability samples is made more difficult by nature of their biasedness and lack of representativity. In this paper we propose quantile balancing inverse probability weighting estimator (QBIPW) for non-probability samples. We apply the idea of Harms and Duchesne (2006) allowing the use of quantile information in the estimation process to reproduce known totals and the distribution of auxiliary variables. We discuss the estimation of the QBIPW probabilities and its variance. Our simulation study has demonstrated that the proposed estimators are robust against model mis-specification and, as a result, help to reduce bias and mean squared error. Finally, we applied the proposed methods to estimate the share of job vacancies aimed at Ukrainian workers in Poland using an integrated set of administrative and survey data about job vacancies.Release date: 2025-12-23
- Articles and reports: 12-001-X202500200009Description: We present and apply methodology to improve inference for small area parameters by using data from several sources. This work extends Cahoy and Sedransk (2023) who showed how to integrate summary statistics from several sources. Our methodology uses hierarchical global-local prior distributions to make inferences for the proportion of individuals in Florida’s counties who do not have health insurance. Results from an extensive simulation study show that this methodology will provide improved inference by using several data sources. Among the five model variants evaluated the ones using horseshoe priors for all variances have better performance than the ones using lasso priors for the local variances.Release date: 2025-12-23
- Previous Go to previous page of Analysis results
- 1 (current) Go to page 1 of Analysis results
- 2 Go to page 2 of Analysis results
- 3 Go to page 3 of Analysis results
- 4 Go to page 4 of Analysis results
- 5 Go to page 5 of Analysis results
- 6 Go to page 6 of Analysis results
- 7 Go to page 7 of Analysis results
- ...
- 61 Go to page 61 of Analysis results
- Next Go to next page of Analysis results
Reference (28)
Reference (28) (10 to 20 of 28 results)
- 11. The Effects of the Revised Estimation Methodology on Estimates from Household Expenditure Surveys ArchivedSurveys and statistical programs – Documentation: 62F0026M2005002Description:
This document will provide an overview of the differences between the old and the new weighting methodologies and the effect of the new weighting system on estimations.
Release date: 2005-06-30 - 12. Variance estimation with plausible value achievement data: Two STATA programs for use with the YITS/PISA data ArchivedSurveys and statistical programs – Documentation: 12-002-X20040016891Description:
These two programs are designed to estimate variability due to measurement error beyond the sampling variance introduced by the survey design in the Youth in Transition Survey / Programme of International Student Assessment (YITS/PISA). Program code is included in an appendix.
Release date: 2004-04-15 - 13. Chain Fisher Volume Index Methodology ArchivedSurveys and statistical programs – Documentation: 13-604-M2003042Description:
On May 31, 2001, the quarterly income and expenditure accounts adopted the Chain Fisher Index formula, chained quarterly, as the official measure of real gross domestic product (GDP) in terms of expenditures. This formula was also adopted for the Provincial Accounts on October 31, 2002.
There were two reasons for adopting this formula: to provide users with a more accurate measure of real GDP growth between two consecutive periods and to make the Canadian measure comparable with the Income and Product Accounts of the United States, which has used the Chain Fisher Index formula since 1996 to measure real GDP.
Release date: 2003-11-06 - Surveys and statistical programs – Documentation: 71F0031X2000001Description:
This paper introduces and explains modifications made to the Labour Force Survey estimates in January 2000. Some of these modifications include the adjustment of all LFS estimates to reflect population counts based on the 1996 Census plus the implementation of a new estimation methodology called composite estimation. This new method results in more efficient estimates of month to month change, while improving the quality of monthly level estimates.
Release date: 2001-06-29 - 15. A donor imputation system to create a census database fully adjusted for underenumeration ArchivedSurveys and statistical programs – Documentation: 11-522-X19990015668Description:
Following the problems with estimating underenumeration in the 1991 Census of England and Wales the aim for the 2001 Census is to create a database that is fully adjusted to net underenumeration. To achieve this, the paper investigates weighted donor imputation methodology that utilises information from both the census and census coverage survey (CCS). The US Census Bureau has considered a similar approach for their 2000 Census (see Isaki et al 1998). The proposed procedure distinguishes between individuals who are not counted by the census because their household is missed and those who are missed in counted households. Census data is linked to data from the CCS. Multinomial logistic regression is used to estimate the probabilities that households are missed by the census and the probabilities that individuals are missed in counted households. Household and individual coverage weights are constructed from the estimated probabilities and these feed into the donor imputation procedure.
Release date: 2000-03-02 - Surveys and statistical programs – Documentation: 11-522-X19990015672Description:
Data fusion as discussed here means to create a set of data on not jointly observed variables from two different sources. Suppose for instance that observations are available for (X,Z) on a set of individuals and for (Y,Z) on a different set of individuals. Each of X, Y and Z may be a vector variable. The main purpose is to gain insight into the joint distribution of (X,Y) using Z as a so-called matching variable. At first however, it is attempted to recover as much information as possible on the joint distribution of (X,Y,Z) from the distinct sets of data. Such fusions can only be done at the cost of implementing some distributional properties for the fused data. These are conditional independencies given the matching variables. Fused data are typically discussed from the point of view of how appropriate this underlying assumption is. Here we give a different perspective. We formulate the problem as follows: how can distributions be estimated in situations when only observations from certain marginal distributions are available. It can be solved by applying the maximum entropy criterium. We show in particular that data created by fusing different sources can be interpreted as a special case of this situation. Thus, we derive the needed assumption of conditional independence as a consequence of the type of data available.
Release date: 2000-03-02 - Surveys and statistical programs – Documentation: 11-522-X19990015674Description:
The effect of the environment on health is of increasing concern, in particular the effects of the release of industrial pollutants into the air, the ground and into water. An assessment of the risks to public health of any particular pollution source is often made using the routine health, demographic and environmental data collected by government agencies. These datasets have important differences in sampling geography and in sampling epochs which affect the epidemiological analyses which draw them together. In the UK, health events are recorded for individuals, giving cause codes, a data of diagnosis or death, and using the unit postcode as a geographical reference. In contrast, small area demographic data are recorded only at the decennial census, and released as area level data in areas distinct from postcode geography. Environmental exposure data may be available at yet another resolution, depending on the type of exposure and the source of the measurements.
Release date: 2000-03-02 - Surveys and statistical programs – Documentation: 11-522-X19990015680Description:
To augment the amount of available information, data from different sources are increasingly being combined. These databases are often combined using record linkage methods. When there is no unique identifier, a probabilistic linkage is used. In that case, a record on a first file is associated with a probability that is linked to a record on a second file, and then a decision is taken on whether a possible link is a true link or not. This usually requires a non-negligible amount of manual resolution. It might then be legitimate to evaluate if manual resolution can be reduced or even eliminated. This issue is addressed in this paper where one tries to produce an estimate of a total (or a mean) of one population, when using a sample selected from another population linked somehow to the first population. In other words, having two populations linked through probabilistic record linkage, we try to avoid any decision concerning the validity of links and still be able to produce an unbiased estimate for a total of the one of two populations. To achieve this goal, we suggest the use of the Generalised Weight Share Method (GWSM) described by Lavallée (1995).
Release date: 2000-03-02 - 19. Simultaneous calibration of several surveys ArchivedSurveys and statistical programs – Documentation: 11-522-X19990015684Description:
Often, the same information is gathered almost simultaneously for several different surveys. In France, this practice is institutionalized for household surveys that have a common set of demographic variables, i.e., employment, residence and income. These variables are important co-factors for the variables of interest in each survey, and if used carefully, can reinforce the estimates derived from each survey. Techniques for calibrating uncertain data can apply naturally in this context. This involves finding the best unbiased estimator in common variables and calibrating each survey based on that estimator. The estimator thus obtained in each survey is always a linear estimator, the weightings of which can be easily explained and the variance can be obtained with no new problems, as can the variance estimate. To supplement the list of regression estimators, this technique can also be seen as a ridge-regression estimator, or as a Bayesian-regression estimator.
Release date: 2000-03-02 - Surveys and statistical programs – Documentation: 11-522-X19990015690Description:
The artificial sample was generated in two steps. The first step, based on a master panel, was a Multiple Correspondence Analysis (MCA) carried out on basic variables. Then, "dummy" individuals were generated randomly using the distribution of each "significant" factor in the analysis. Finally, for each individual, a value was generated for each basic variable most closely linked to one of the previous factors. This method ensured that sets of variables were drawn independently. The second step consisted in grafting some other data bases, based on certain property requirements. A variable was generated to be added on the basis of its estimated distribution, using a generalized linear model for common variables and those already added. The same procedure was then used to graft the other samples. This method was applied to the generation of an artificial sample taken from two surveys. The artificial sample that was generated was validated using sample comparison testing. The results were positive, demonstrating the feasibility of this method.
Release date: 2000-03-02