How to decompose the non-response variance: A total survey error approach
Section 1. Introduction
Total survey error is described by Biemer (2010) as the “accumulation of all errors that may arise in the design, collection, processing and analysis of survey data”. He classified survey error components into sampling error and nonsampling errors, such as, non-response, coverage, measurement and data processing errors. These errors may affect variance, bias, or both. The total survey error paradigm aims at maximizing survey quality by minimizing total survey error within prespecified resource constraints like budget, people, or time.
At Statistics Canada, the Corporate Business Architecture initiated the Integrated Business Statistics Program (IBSP) as the standardized platform for more than 140 economic surveys with the objective of achieving efficiency, enhancing quality and improving responsiveness. In particular, reducing collection costs while managing non-response error was identified as one of the program’s pillars. Consequently, an adaptive design where different units may receive different treatments became a keystone for this program. For more details on IBSP, see Statistics Canada (2015). Groves and Heeringa (2006) showed how paradata could be used to increase the response rate. Schouten, Calinescu and Luiten (2013) gave a general framework for an adaptive design and explained how the R-indicator could be used in this context.
A new survey process model called Rolling Estimates has been developed as an attempt to address the IBSP’s pillar mentioned above. The Rolling Estimates model is based on iterative processing and estimation cycles throughout the collection period. Basically, the idea of this model is to compute key estimates with their associated quality indicators at several specific times during the collection period. At the beginning, all units are assigned to the self-response survey treatment which means that the respondents are asked to complete the online questionnaire. Collection efforts like computer-assisted telephone interview non-response follow-ups are then performed on units contributing the most to the estimates where the quality is low based on the preliminary results of the Rolling Estimates. This can be viewed as an adaptive design since the treatments on the units depend on the quality of the estimates produced during the collection period. Most of the work regarding the development of the IBSP’s adaptive design has been done since 2010. Godbout, Beaucage and Turmelle (2011), Turmelle, Godbout and Bosa (2012), Mills, Godbout, Bosa and Turmelle (2013) and Bosa and Godbout (2014) made use of this idea in the context of the IBSP adaptive design to minimize the number of follow-ups in order to reach a targeted quality in terms of coefficient of variation.
This paper revisits the work done so far for IBSP and presents an approach to decompose non-response variance into an item-level score for a given variable of interest within a domain. This item score is basically an attempt to estimate the contribution to the variance borrowed by a single unit. Units with a large score will contribute the most to reduce the variance and the coefficient of variation which is often used as a quality indicator in surveys. However, there are generally many important variables and domains in a survey. The proposed approach first computes, for a given unit, item-level scores for important variables and domains. Then, item scores can be combined into a single unit-level score in order to rank units. For example, the unit score can be a weighted sum or the maximum of its item scores. The most attractive use of the resulting unit-level score is to prioritize units, the ones with the highest scores, for the most expensive collection operations such as telephone follow-up, computer-assisted telephone interview or computer-assisted personal interview. This paper assumes total and partial non-response are both treated in the adaptive design, but treatments may be different depending on the type of non-response. For instance, telephone follow-ups could be made in the case of total non-response whereas questionnaires with partial non-response could be reviewed by analysts. This type of adaptive design generates strong interactions between collection operations, observed data and measured quality. Bosa and Godbout (2014) showed how this methodology was implemented in IBSP under the Rolling Estimates model.
Emphasis will be placed on the derivation of the item-level score throughout this paper. Therefore, the special case of only one variable of interest within a domain is studied. Also, only one imputation method is used to impute the variable of interest in the case of non-response so as to simplify the notation and to ease comprehension for the reader.
Section 2 describes the inference framework. In Section 3, the decomposition of the variance at the unit-level is expressed. In other words, the contribution of each nonresponding unit to the variance is computed. A simulation study was conducted to evaluate the proposed score. It is described in Section 4. Finally, Section 5 expresses some thoughts and conclusions.
- Date modified: