Sort Help
entries

Results

All (16)

All (16) (0 to 10 of 16 results)

  • Articles and reports: 12-001-X202300200017
    Description: Jean-Claude Deville, who passed away in October 2021, was one of the most influential researchers in the field of survey statistics over the past 40 years. This article traces some of his contributions that have had a profound impact on both survey theory and practice. This article will cover the topics of balanced sampling using the cube method, calibration, the weight-sharing method, the development of variance expressions of complex estimators using influence function and quota sampling.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X201900300004
    Description:

    Social or economic studies often need to have a global view of society. For example, in agricultural studies, the characteristics of farms can be linked to the social activities of individuals. Hence, studies of a given phenomenon should be done by considering variables of interest referring to different target populations that are related to each other. In order to get an insight into an underlying phenomenon, the observations must be carried out in an integrated way, in which the units of a given population have to be observed jointly with related units of the other population. In the agricultural example, this means that a sample of rural households should be selected that have some relationship with the farm sample to be used for the study. There are several ways to select integrated samples. This paper studies the problem of defining an optimal sampling strategy for this situation: the solution proposed minimizes the sampling cost, ensuring a predefined estimation precision for the variables of interest (of either one or both populations) describing the phenomenon. Indirect sampling provides a natural framework for this setting since the units belonging to a population can become carriers of information on another population that is the object of a given survey. The problem is studied for different contexts which characterize the information concerning the links available in the sampling design phase, ranging from situations in which the links among the different units are known in the design phase to a situation in which the available information on links is very poor. An empirical study of agricultural data for a developing country is presented. It shows how controlling the inclusion probabilities at the design phase using the available information (namely the links) is effective, can significantly reduce the errors of the estimates for the indirectly observed population. The need for good models for predicting the unknown variables or the links is also demonstrated.

    Release date: 2019-12-17

  • Articles and reports: 12-001-X201300111829
    Description:

    Indirect Sampling is used when the sampling frame is not the same as the target population, but related to the latter. The estimation process for Indirect Sampling is carried out using the Generalised Weight Share Method (GWSM), which is an unbiased procedure (see Lavallée 2002, 2007). For business surveys, Indirect Sampling is applied as follows: the sampling frame is one of establishments, while the target population is one of enterprises. Enterprises are selected through their establishments. This allows stratifying according to the establishment characteristics, rather than those associated with enterprises. Because the variables of interest of establishments are generally highly skewed (a small portion of the establishments covers the major portion of the economy), the GWSM results in unbiased estimates, but their variance can be large. The purpose of this paper is to suggest some adjustments to the weights to reduce the variance of the estimates in the context of skewed populations, while keeping the method unbiased. After a brief overview of Indirect Sampling and the GWSM, we describe the required adjustments to the GWSM. The estimates produced with these adjustments are compared to those from the original GWSM, via a small numerical example, and using real data originating from the Statistics Canada's Business Register.

    Release date: 2013-06-28

  • Articles and reports: 12-001-X200900211038
    Description:

    We examine overcoming the overestimation in using generalized weight share method (GWSM) caused by link nonresponse in indirect sampling. A few adjustment methods incorporating link nonresponse in using GWSM have been constructed for situations both with and without the availability of auxiliary variables. A simulation study on a longitudinal survey is presented using some of the adjustment methods we recommend. The simulation results show that these adjusted GWSMs perform well in reducing both estimation bias and variance. The advancement in bias reduction is significant.

    Release date: 2009-12-23

  • Articles and reports: 12-001-X200700210490
    Description:

    The European Union's Statistics on Income and Living Conditions (SILC) survey was introduced in 2004 as a replacement for the European Panel. It produces annual statistics on income distribution, poverty and social exclusion. First conducted in France in May 2004, it is a longitudinal survey of all individuals over the age of 15 in 16,000 dwellings selected from the master sample and the new-housing sample frame. All respondents are tracked over time, even when they move to a different dwelling. The survey also has to produce cross-sectional estimates of good quality.

    To limit the response burden, the sample design recommended by Eurostat is a rotation scheme consisting of four panels that remain in the sample for four years, with one panel replaced each year. France, however, decided to increase the panel duration to nine years. The rotating sample design meets the survey's longitudinal and cross-sectional requirements, but it presents some weighting challenges.

    Following a review of the inference context of a longitudinal survey, the paper discusses the longitudinal and cross-sectional weighting, which are designed to produce approximately unbiased estimators.

    Release date: 2008-01-03

  • Articles and reports: 12-001-X20070019851
    Description:

    To model economic depreciation, a database is used that contains information on assets discarded by companies. The acquisition and resale prices are known along with the length of use of these assets. However, the assets for which prices are known are only those that were involved in a transaction. While an asset depreciates on a continuous basis during its service life, the value of the asset is only known when there has been a transaction. This article proposes an ex post weighting to offset the effect of source of error in building econometric models.

    Release date: 2007-06-28

  • Articles and reports: 11-522-X20050019494
    Description:

    Traditionally, data quality indicators reported by surveys have been the sampling variance, coverage error, non-response rate and imputation rate. To obtain an imputation rate when combining survey data and administrative data, one of the problems is to compute the imputation rate itself. The presentation will discuss how to solve this problem. First, we will discuss the desired properties when developing a rate in a general context. Second, we will develop some concepts and definitions that will help us to develop combine rates. Third, we will propose different combined rates for the case of imputation. We will then present three different combined rates, and we will discuss properties for each rate. We will end with some illustrative examples.

    Release date: 2007-03-02

  • Articles and reports: 12-001-X20060029551
    Description:

    To select a survey sample, it happens that one does not have a frame containing the desired collection units, but rather another frame of units linked in a certain way to the list of collection units. It can then be considered to select a sample from the available frame in order to produce an estimate for the desired target population by using the links existing between the two. This can be designated by Indirect Sampling.

    Estimation for the target population surveyed by Indirect Sampling can constitute a big challenge, in particular if the links between the units of the two are not one-to-one. The problem comes especially from the difficulty to associate a selection probability, or an estimation weight, to the surveyed units of the target population. In order to solve this type of estimation problem, the Generalized Weight Share Method (GWSM) has been developed by Lavallée (1995) and Lavallée (2002). The GWSM provides an estimation weight for every surveyed unit of the target population.

    This paper first describes Indirect Sampling, which constitutes the foundations of the GWSM. Second, an overview of the GWSM is given where we formulate the GWSM in a theoretical framework using matrix notation. Third, we present some properties of the GWSM such as unbiasedness and transitivity. Fourth, we consider the special case where the links between the two populations are expressed by indicator variables. Fifth, some special typical linkages are studied to assess their impact on the GWSM. Finally, we consider the problem of optimality. We obtain optimal weights in a weak sense (for specific values of the variable of interest), and conditions for which these weights are also optimal in a strong sense and independent of the variable of interest.

    Release date: 2006-12-21

  • Articles and reports: 11-522-X20030017594
    Description:

    This article describes the European Survey on Income and Living Conditions (SILC), which will replace the European Panel beginning in 2004. It also looks at the use of the weight share method in the longitudinal and cross-sectional weightings of the SILC.

    Release date: 2005-01-26

  • Articles and reports: 11-522-X20010016267
    Description:

    This paper discusses in detail issues dealing with the technical aspects of designing and conducting surveys. It is intended for an audience of survey methodologists.

    In practice, a list of the desired collection units is not always available. Instead, a list of different units that are somehow related to the collection units may be provided, thus producing two related populations, UA and UB. An estimate for UB needs to be created, however, the sampling frame provided is only for the UA population.

    One solution for this problem is to select a sample from UA (sA) and produce an estimate for UB using the existing relationship between the two populations. This process may be referred to as indirect sampling. To assign a selection probability, or an estimation weight, for the survey units, Lavallée (1995) developed the generalized weight share method (GWSM). The GWSM produces an estimation weight that basically constitutes an average of the sampling weights of the units in sA.

    This paper discusses the types of non-response associated with indirect sampling and the possible estimation problems that can occur in the application of the GWSM.

    Release date: 2002-09-12
Articles and reports (16)

Articles and reports (16) (0 to 10 of 16 results)

  • Articles and reports: 12-001-X202300200017
    Description: Jean-Claude Deville, who passed away in October 2021, was one of the most influential researchers in the field of survey statistics over the past 40 years. This article traces some of his contributions that have had a profound impact on both survey theory and practice. This article will cover the topics of balanced sampling using the cube method, calibration, the weight-sharing method, the development of variance expressions of complex estimators using influence function and quota sampling.
    Release date: 2024-01-03

  • Articles and reports: 12-001-X201900300004
    Description:

    Social or economic studies often need to have a global view of society. For example, in agricultural studies, the characteristics of farms can be linked to the social activities of individuals. Hence, studies of a given phenomenon should be done by considering variables of interest referring to different target populations that are related to each other. In order to get an insight into an underlying phenomenon, the observations must be carried out in an integrated way, in which the units of a given population have to be observed jointly with related units of the other population. In the agricultural example, this means that a sample of rural households should be selected that have some relationship with the farm sample to be used for the study. There are several ways to select integrated samples. This paper studies the problem of defining an optimal sampling strategy for this situation: the solution proposed minimizes the sampling cost, ensuring a predefined estimation precision for the variables of interest (of either one or both populations) describing the phenomenon. Indirect sampling provides a natural framework for this setting since the units belonging to a population can become carriers of information on another population that is the object of a given survey. The problem is studied for different contexts which characterize the information concerning the links available in the sampling design phase, ranging from situations in which the links among the different units are known in the design phase to a situation in which the available information on links is very poor. An empirical study of agricultural data for a developing country is presented. It shows how controlling the inclusion probabilities at the design phase using the available information (namely the links) is effective, can significantly reduce the errors of the estimates for the indirectly observed population. The need for good models for predicting the unknown variables or the links is also demonstrated.

    Release date: 2019-12-17

  • Articles and reports: 12-001-X201300111829
    Description:

    Indirect Sampling is used when the sampling frame is not the same as the target population, but related to the latter. The estimation process for Indirect Sampling is carried out using the Generalised Weight Share Method (GWSM), which is an unbiased procedure (see Lavallée 2002, 2007). For business surveys, Indirect Sampling is applied as follows: the sampling frame is one of establishments, while the target population is one of enterprises. Enterprises are selected through their establishments. This allows stratifying according to the establishment characteristics, rather than those associated with enterprises. Because the variables of interest of establishments are generally highly skewed (a small portion of the establishments covers the major portion of the economy), the GWSM results in unbiased estimates, but their variance can be large. The purpose of this paper is to suggest some adjustments to the weights to reduce the variance of the estimates in the context of skewed populations, while keeping the method unbiased. After a brief overview of Indirect Sampling and the GWSM, we describe the required adjustments to the GWSM. The estimates produced with these adjustments are compared to those from the original GWSM, via a small numerical example, and using real data originating from the Statistics Canada's Business Register.

    Release date: 2013-06-28

  • Articles and reports: 12-001-X200900211038
    Description:

    We examine overcoming the overestimation in using generalized weight share method (GWSM) caused by link nonresponse in indirect sampling. A few adjustment methods incorporating link nonresponse in using GWSM have been constructed for situations both with and without the availability of auxiliary variables. A simulation study on a longitudinal survey is presented using some of the adjustment methods we recommend. The simulation results show that these adjusted GWSMs perform well in reducing both estimation bias and variance. The advancement in bias reduction is significant.

    Release date: 2009-12-23

  • Articles and reports: 12-001-X200700210490
    Description:

    The European Union's Statistics on Income and Living Conditions (SILC) survey was introduced in 2004 as a replacement for the European Panel. It produces annual statistics on income distribution, poverty and social exclusion. First conducted in France in May 2004, it is a longitudinal survey of all individuals over the age of 15 in 16,000 dwellings selected from the master sample and the new-housing sample frame. All respondents are tracked over time, even when they move to a different dwelling. The survey also has to produce cross-sectional estimates of good quality.

    To limit the response burden, the sample design recommended by Eurostat is a rotation scheme consisting of four panels that remain in the sample for four years, with one panel replaced each year. France, however, decided to increase the panel duration to nine years. The rotating sample design meets the survey's longitudinal and cross-sectional requirements, but it presents some weighting challenges.

    Following a review of the inference context of a longitudinal survey, the paper discusses the longitudinal and cross-sectional weighting, which are designed to produce approximately unbiased estimators.

    Release date: 2008-01-03

  • Articles and reports: 12-001-X20070019851
    Description:

    To model economic depreciation, a database is used that contains information on assets discarded by companies. The acquisition and resale prices are known along with the length of use of these assets. However, the assets for which prices are known are only those that were involved in a transaction. While an asset depreciates on a continuous basis during its service life, the value of the asset is only known when there has been a transaction. This article proposes an ex post weighting to offset the effect of source of error in building econometric models.

    Release date: 2007-06-28

  • Articles and reports: 11-522-X20050019494
    Description:

    Traditionally, data quality indicators reported by surveys have been the sampling variance, coverage error, non-response rate and imputation rate. To obtain an imputation rate when combining survey data and administrative data, one of the problems is to compute the imputation rate itself. The presentation will discuss how to solve this problem. First, we will discuss the desired properties when developing a rate in a general context. Second, we will develop some concepts and definitions that will help us to develop combine rates. Third, we will propose different combined rates for the case of imputation. We will then present three different combined rates, and we will discuss properties for each rate. We will end with some illustrative examples.

    Release date: 2007-03-02

  • Articles and reports: 12-001-X20060029551
    Description:

    To select a survey sample, it happens that one does not have a frame containing the desired collection units, but rather another frame of units linked in a certain way to the list of collection units. It can then be considered to select a sample from the available frame in order to produce an estimate for the desired target population by using the links existing between the two. This can be designated by Indirect Sampling.

    Estimation for the target population surveyed by Indirect Sampling can constitute a big challenge, in particular if the links between the units of the two are not one-to-one. The problem comes especially from the difficulty to associate a selection probability, or an estimation weight, to the surveyed units of the target population. In order to solve this type of estimation problem, the Generalized Weight Share Method (GWSM) has been developed by Lavallée (1995) and Lavallée (2002). The GWSM provides an estimation weight for every surveyed unit of the target population.

    This paper first describes Indirect Sampling, which constitutes the foundations of the GWSM. Second, an overview of the GWSM is given where we formulate the GWSM in a theoretical framework using matrix notation. Third, we present some properties of the GWSM such as unbiasedness and transitivity. Fourth, we consider the special case where the links between the two populations are expressed by indicator variables. Fifth, some special typical linkages are studied to assess their impact on the GWSM. Finally, we consider the problem of optimality. We obtain optimal weights in a weak sense (for specific values of the variable of interest), and conditions for which these weights are also optimal in a strong sense and independent of the variable of interest.

    Release date: 2006-12-21

  • Articles and reports: 11-522-X20030017594
    Description:

    This article describes the European Survey on Income and Living Conditions (SILC), which will replace the European Panel beginning in 2004. It also looks at the use of the weight share method in the longitudinal and cross-sectional weightings of the SILC.

    Release date: 2005-01-26

  • Articles and reports: 11-522-X20010016267
    Description:

    This paper discusses in detail issues dealing with the technical aspects of designing and conducting surveys. It is intended for an audience of survey methodologists.

    In practice, a list of the desired collection units is not always available. Instead, a list of different units that are somehow related to the collection units may be provided, thus producing two related populations, UA and UB. An estimate for UB needs to be created, however, the sampling frame provided is only for the UA population.

    One solution for this problem is to select a sample from UA (sA) and produce an estimate for UB using the existing relationship between the two populations. This process may be referred to as indirect sampling. To assign a selection probability, or an estimation weight, for the survey units, Lavallée (1995) developed the generalized weight share method (GWSM). The GWSM produces an estimation weight that basically constitutes an average of the sampling weights of the units in sA.

    This paper discusses the types of non-response associated with indirect sampling and the possible estimation problems that can occur in the application of the GWSM.

    Release date: 2002-09-12