Keyword search

Filter results by

Search Help
Currently selected filters that can be removed

Keyword(s)

Geography

3 facets displayed. 0 facets selected.

Content

1 facets displayed. 0 facets selected.
Sort Help
entries

Results

All (186)

All (186) (30 to 40 of 186 results)

  • Articles and reports: 11F0027M2010060
    Geography: Canada
    Description:

    This paper asks whether synergies or managerial discipline operates in different ways across small versus large plants to affect the likelihood of mergers. Our findings indicate that those characteristics which provide the type of synergies upon which ownership changes rely are important factors leading to plant-ownership changes across most size classes. The magnitudes, however, are different across plant-size classes, with synergies generally being more important in larger plants.

    Foreign plants in all size classes are more likely to be taken over. The effective rates of control change differ much more in the small than in the larger size classes. Compared to domestic plants, multinational plants in the smaller size classes contain relatively more of the type of intangible capital that makes them attractive vehicles for the transmission of new knowledge via takeover.

    Release date: 2010-02-25

  • Articles and reports: 12-001-X200900211039
    Description:

    Propensity weighting is a procedure to adjust for unit nonresponse in surveys. A form of implementing this procedure consists of dividing the sampling weights by estimates of the probabilities that the sampled units respond to the survey. Typically, these estimates are obtained by fitting parametric models, such as logistic regression. The resulting adjusted estimators may become biased when the specified parametric models are incorrect. To avoid misspecifying such a model, we consider nonparametric estimation of the response probabilities by local polynomial regression. We study the asymptotic properties of the resulting estimator under quasi-randomization. The practical behavior of the proposed nonresponse adjustment approach is evaluated on NHANES data.

    Release date: 2009-12-23

  • Articles and reports: 12-001-X200900211045
    Description:

    In analysis of sample survey data, degrees-of-freedom quantities are often used to assess the stability of design-based variance estimators. For example, these degrees-of-freedom values are used in construction of confidence intervals based on t distribution approximations; and of related t tests. In addition, a small degrees-of-freedom term provides a qualitative indication of the possible limitations of a given variance estimator in a specific application. Degrees-of-freedom calculations sometimes are based on forms of the Satterthwaite approximation. These Satterthwaite-based calculations depend primarily on the relative magnitudes of stratum-level variances. However, for designs involving a small number of primary units selected per stratum, standard stratum-level variance estimators provide limited information on the true stratum variances. For such cases, customary Satterthwaite-based calculations can be problematic, especially in analyses for subpopulations that are concentrated in a relatively small number of strata. To address this problem, this paper uses estimated within-primary-sample-unit (within PSU) variances to provide auxiliary information regarding the relative magnitudes of the overall stratum-level variances. Analytic results indicate that the resulting degrees-of-freedom estimator will be better than modified Satterthwaite-type estimators provided: (a) the overall stratum-level variances are approximately proportional to the corresponding within-stratum variances; and (b) the variances of the within-PSU variance estimators are relatively small. In addition, this paper develops errors-in-variables methods that can be used to check conditions (a) and (b) empirically. For these model checks, we develop simulation-based reference distributions, which differ substantially from reference distributions based on customary large-sample normal approximations. The proposed methods are applied to four variables from the U.S. Third National Health and Nutrition Examination Survey (NHANES III).

    Release date: 2009-12-23

  • Articles and reports: 12-001-X200900211046
    Description:

    A semiparametric regression model is developed for complex surveys. In this model, the explanatory variables are represented separately as a nonparametric part and a parametric linear part. The estimation techniques combine nonparametric local polynomial regression estimation and least squares estimation. Asymptotic results such as consistency and normality of the estimators of regression coefficients and the regression functions have also been developed. Success of the performance of the methods and the properties of estimates have been shown by simulation and empirical examples with the Ontario Health Survey 1990.

    Release date: 2009-12-23

  • Articles and reports: 12-001-X200900211056
    Description:

    In this Issue is a column where the Editor biefly presents each paper of the current issue of Survey Methodology. As well, it sometimes contain informations on structure or management changes in the journal.

    Release date: 2009-12-23

  • Articles and reports: 11-522-X200800010961
    Description:

    Increasingly, children of all ages are becoming respondents in survey interviews. While juveniles are considered to be reliable respondents for many topics and survey settings it is unclear to what extend younger children provide reliable information in a face-to-face interview. In this paper we will report results from a study using video captures of 205 face-to-face interviews with children aged 8 through 14. The interviews have been coded using behavior codes on a question by question level which provides behavior-related indicators regarding the question-answer process. In addition, standard tests of cognitive resources have been conducted. Using visible and audible problems in the respondent behavior, we are able to assess the impact of the children's cognitive resources on respondent behaviors. Results suggest that girls and boys differ fundamentally in the cognitive mechanisms leading to problematic respondent behaviors.

    Release date: 2009-12-03

  • Articles and reports: 11-536-X200900110803
    Description:

    "Classical GREG estimator" is used here to refer to the generalized regression estimator extensively discussed for example in Särndal, Swensson and Wretman (1992). This paper summarize some recent extensions of the classical GREG estimator when applied to the estimation of totals for population subgroups or domains. GREG estimation was introduced for domain estimation in Särndal (1981, 1984), Hidiroglou and Särndal (1985) and Särndal and Hidiroglou (1989), and was developed further in Estevao, Hidiroglou and Särndal (1995). For the classical GREG estimator, fixed-effects linear model serves as the underlying working or assisting model, and aggregate-level auxiliary totals are incorporated in the estimation procedure. In some recent developments, an access to unit-level auxiliary data is assumed for GREG estimation for domains. Obviously, an access to micro-merged register and survey data involves much flexibility for domain estimation. This view has been adopted for GREG estimation for example in Lehtonen and Veijanen (1998), Lehtonen, Särndal and Veijanen (2003, 2005), and Lehtonen, Myrskylä, Särndal and Veijanen (2007). These extensions cover the cases of continuous and binary or polytomous response variables, use of generalized linear mixed models as assisting models, and unequal probability sampling designs. Relative merits and challenges of the various GREG estimators will be discussed.

    Release date: 2009-08-11

  • Articles and reports: 11-536-X200900110808
    Description:

    Let auxiliary information be available for use in designing of a survey sample. Let the sample selection procedure consist of selecting a probability sample, rejecting the sample if the sample mean of an auxiliary variable is not within a specified distance of the population mean, continuing until a sample is accepted. It is proven that the large sample properties of the regression estimator for the rejective sample are the same as those of the regression estimator for the original selection procedure. Likewise the usual estimator of variance for the regression estimator is appropriate for the rejective sample. In a Monte Carlo experiment, the large sample properties hold for relatively small samples. Also the Monte Carlo results are in agreement with the theoretical orders of approximation. The efficiency effect of the described rejective sampling is o(n-1) relative to regression estimation without rejection, but the effect can be important for particular samples.

    Release date: 2009-08-11

  • Articles and reports: 11-536-X200900110814
    Description:

    Calibration is the principal theme in many recent articles on estimation in survey sampling. Words such as "calibration approach" and "calibration estimators" are frequently used. As article authors like to point out, calibration provides a systematic way to incorporate auxiliary information in the procedure.

    Calibration has established itself as an important methodological instrument in large-scale production of statistics. Several national statistical agencies have developed software designed to compute weights, usually calibrated to auxiliary information available in administrative registers and other accurate sources.

    This paper presents a review of the calibration approach, with an emphasis on progress achieved in the past decade or so. The literature on calibration is growing rapidly; selected issues are discussed in this paper.

    The paper starts with a definition of the calibration approach. Its important features are reviewed. The calibration approach is contrasted with (generalized) regression estimation, which is an alternative but different way to take auxiliary information into account. The computational aspects of calibration are discussed, including methods for avoiding extreme weights. In the early sections of the paper, simple applications of calibration are examined: Estimation of a population total in direct, single phase sampling. Generalization to more complex parameters and more complex sampling designs are then considered. A common feature of more complex designs (sampling in two or more phases or stages) is that the available auxiliary information may consist of several components or layers. The uses of calibration in such cases of composite information are reviewed. In later sections of the paper, some examples are given to illustrate how the results of the calibration thinking may contrast with answers given by earlier established approaches. Finally, applications of calibration in the presence of nonsampling error are discussed, in particular methods for nonresponse bias adjustment.

    Release date: 2009-08-11

  • Articles and reports: 11-536-X200900110815
    Description:

    Regression estimation has only been used extensively by statistical organizations in recent years. The sample mean and the classical ratio estimator are particular cases of this estimator. Both have a long history that can be traced back to the ancient Greeks or earlier. For instance, the ratio estimator was used by the Egyptians and Babylonians to compute the circumference of a circle given its radius, given a constant approximating the famous pi value. Multiple examples that use the ratio estimator as a means to indirectly compute a measure of interest can be found in physics, and engineering. The ratio estimator was also used to estimate population census estimates in the past when exact counts were beyond the capabilities of the existing administrations (for example John Graunt 1662, and Laplace late eighteenth century).

    In this talk, we will trace the evolution of the regression estimation in survey sampling from the 1930's to the present time. We will outline its advantages and disadvantages in using it in survey sampling. Corresponding software development will also be presented

    Release date: 2009-08-11
Data (2)

Data (2) ((2 results))

  • Public use microdata: 99M0001X
    Description: The Individuals File, 2011 National Household Survey (Public Use Microdata Files) provides data on the characteristics of the Canadian population. The file contains a 2.7% sample of anonymous responses to the 2011 National Household Survey (NHS) questionnaire. The files have been carefully scrutinized to ensure the complete confidentiality of the individual responses and geographic identifiers have been restricted to provinces/territories and metropolitan areas. With 133 variables, this comprehensive tool is excellent for policy analysts, pollsters, social researchers and anyone interested in modelling and performing statistical regression analysis using National Household Survey data.

    Microdata files uniquely provide users access to non-aggregated data. The PUMFs user can group and manipulate these variables to suit data and research requirements. Tabulations excluded from other NHS products can be created or relationships between variables can be analyzed using different statistical tests. PUMFs provide quick access to a comprehensive social and economic database about Canada and its people.

    This product, offered on DVD-ROM, contains the data file (in ASCII format); user documentation and supporting information; all licence agreements; and SAS, SPSS and Stata program source codes to enable users to read the set of records. It is important to note that users will require knowledge of data manipulation packages (or software) such as SAS, SPSS or Stata to use this product.

    Release date: 2023-09-12

  • Table: 75-001-X19890022277
    Description:

    This study compares the earnings of bilingual and unilingual workers in three urban centres: Montreal, Toronto and Ottawa-Hull. Differences in the earnings of bilingual and unilingual workers are considered in the light of several demographic and job-related traits.

    Release date: 1989-06-30
Analysis (174)

Analysis (174) (0 to 10 of 174 results)

  • Articles and reports: 82-003-X202500800001
    Description: Data measuring life expectancy (LE) and health-adjusted life expectancy (HALE) in Canada are available for large geographical areas, such as provinces, territories, and health regions. However, to date, no study has analyzed LE and HALE at the municipal level. To address issues related to sparse administrative and survey data in small geographic areas, this study applies multilevel regression models and poststratification methods that have been shown to provide reliable estimates of population- and small area-level quantities from health surveys.
    Release date: 2025-08-20

  • Articles and reports: 12-001-X202300100002
    Description: We consider regression analysis in the context of data integration. To combine partial information from external sources, we employ the idea of model calibration which introduces a “working” reduced model based on the observed covariates. The working reduced model is not necessarily correctly specified but can be a useful device to incorporate the partial information from the external data. The actual implementation is based on a novel application of the information projection and model calibration weighting. The proposed method is particularly attractive for combining information from several sources with different missing patterns. The proposed method is applied to a real data example combining survey data from Korean National Health and Nutrition Examination Survey and big data from National Health Insurance Sharing Service in Korea.
    Release date: 2023-06-30

  • Articles and reports: 11-522-X202100100009
    Description:

    Use of auxiliary data to improve the efficiency of estimators of totals and means through model-assisted survey regression estimation has received considerable attention in recent years. Generalized regression (GREG) estimators, based on a working linear regression model, are currently used in establishment surveys at Statistics Canada and several other statistical agencies.  GREG estimators use common survey weights for all study variables and calibrate to known population totals of auxiliary variables. Increasingly, many auxiliary variables are available, some of which may be extraneous. This leads to unstable GREG weights when all the available auxiliary variables, including interactions among categorical variables, are used in the working linear regression model. On the other hand, new machine learning methods, such as regression trees and lasso, automatically select significant auxiliary variables and lead to stable nonnegative weights and possible efficiency gains over GREG.  In this paper, a simulation study, based on a real business survey sample data set treated as the target population, is conducted to study the relative performance of GREG, regression trees and lasso in terms of efficiency of the estimators.

    Key Words: Model assisted inference; calibration estimation; model selection; generalized regression estimator.

    Release date: 2021-10-29

  • Articles and reports: 89-657-X2018001
    Description:

    This study draws on data from the Longitudinal Immigration Database to examine participation in Canadian post-secondary education (PSE) among adult immigrants in the 2002-2005 landing cohort, with an explicit focus on resettled refugees. The study describes the demographic characteristics of participants, the qualities of participation, and the economic returns on investment in Canadian PSE. It also employs multivariate regression analysis to further examine the effects of participation in Canadian training on employment incidence and the income of those employed, while controlling for other factors associated with successful economic integration.

    Release date: 2018-11-14

  • Articles and reports: 12-001-X201600114541
    Description:

    In this work we compare nonparametric estimators for finite population distribution functions based on two types of fitted values: the fitted values from the well-known Kuo estimator and a modified version of them, which incorporates a nonparametric estimate for the mean regression function. For each type of fitted values we consider the corresponding model-based estimator and, after incorporating design weights, the corresponding generalized difference estimator. We show under fairly general conditions that the leading term in the model mean square error is not affected by the modification of the fitted values, even though it slows down the convergence rate for the model bias. Second order terms of the model mean square errors are difficult to obtain and will not be derived in the present paper. It remains thus an open question whether the modified fitted values bring about some benefit from the model-based perspective. We discuss also design-based properties of the estimators and propose a variance estimator for the generalized difference estimator based on the modified fitted values. Finally, we perform a simulation study. The simulation results suggest that the modified fitted values lead to a considerable reduction of the design mean square error if the sample size is small.

    Release date: 2016-06-22

  • Articles and reports: 12-001-X201600114543
    Description:

    The regression estimator is extensively used in practice because it can improve the reliability of the estimated parameters of interest such as means or totals. It uses control totals of variables known at the population level that are included in the regression set up. In this paper, we investigate the properties of the regression estimator that uses control totals estimated from the sample, as well as those known at the population level. This estimator is compared to the regression estimators that strictly use the known totals both theoretically and via a simulation study.

    Release date: 2016-06-22

  • Articles and reports: 12-001-X201600114545
    Description:

    The estimation of quantiles is an important topic not only in the regression framework, but also in sampling theory. A natural alternative or addition to quantiles are expectiles. Expectiles as a generalization of the mean have become popular during the last years as they not only give a more detailed picture of the data than the ordinary mean, but also can serve as a basis to calculate quantiles by using their close relationship. We show, how to estimate expectiles under sampling with unequal probabilities and how expectiles can be used to estimate the distribution function. The resulting fitted distribution function estimator can be inverted leading to quantile estimates. We run a simulation study to investigate and compare the efficiency of the expectile based estimator.

    Release date: 2016-06-22

  • Articles and reports: 11F0019M2016376
    Geography: Canada, Province or territory
    Description: The degree to which workers move across geographic areas in response to emerging employment opportunities or negative labour demand shocks is a key element in the adjustment process of an economy, and its ability to reach a desired allocation of resources.

    This study estimates the causal impact of real after-tax annual wages and salaries on the propensity of young men to migrate to Alberta or to accept jobs in that province while maintaining residence in their home province. To do so, it exploits the cross-provincial variation in earnings growth plausibly induced by increases in world oil prices that occurred during the 2000s.

    Release date: 2016-04-11

  • Articles and reports: 11F0019M2015371
    Description:

    This paper investigates whether registered pension plans (RPPs) help households prepare financially for retirement or simply substitute for other forms of private saving. This issue is addressed using a panel of 1.8 million Canadian households, from 1991 to 2010, which appear in the Longitudinal Administrative Databank. The analysis controls for correlations in savings across accounts due to unobserved tastes for saving by exploiting the fact that employer contribution rates increase discontinuously on earnings above the average industrial wage, a unique feature of occupational pensions in Canada, the effect being estimated in a Regression Kink Design.

    Release date: 2015-12-21

  • Articles and reports: 12-001-X201500214236
    Description:

    We propose a model-assisted extension of weighting design-effect measures. We develop a summary-level statistic for different variables of interest, in single-stage sampling and under calibration weight adjustments. Our proposed design effect measure captures the joint effects of a non-epsem sampling design, unequal weights produced using calibration adjustments, and the strength of the association between an analysis variable and the auxiliaries used in calibration. We compare our proposed measure to existing design effect measures in simulations using variables like those collected in establishment surveys and telephone surveys of households.

    Release date: 2015-12-17
Reference (10)

Reference (10) ((10 results))

  • Surveys and statistical programs – Documentation: 11-522-X20010016308
    Description:

    This paper discusses in detail issues dealing with the technical aspects of designing and conducting surveys. It is intended for an audience of survey methodologists.

    The Census Bureau uses response error analysis to evaluate the effectiveness of survey questions. For a given survey, questions that are deemed critical to the survey or considered problematic from past examination are selected for analysis. New or revised questions are prime candidates for re-interview. Re-interview is a new interview where a subset of questions from the original interview are re-asked to a sample of the survey respondents. For each re-interview question, the proportion of respondents who give inconsistent responses is evaluated. The "Index of Inconsistency" is used as the measure of response variance. Each question is labelled low, moderate, or high in response variance. In high response variance cases, the questions are put through cognitive testing, and modifications to the question are recommended.

    The Schools and Staffing Survey (SASS) sponsored by The National Center for Education Statistics (NCES), is also investigated for response error analysis and the possible relationships between inconsistent responses and characteristics of the schools and teachers in that survey. Results of this analysis can be used to change survey procedures and improve data quality.

    Release date: 2002-09-12

  • Surveys and statistical programs – Documentation: 11-522-X19990015656
    Description:

    Time series studies have shown associations between air pollution concentrations and morbidity and mortality. These studies have largely been conducted within single cities, and with varying methods. Critics of these studies have questioned the validity of the data sets used and the statistical techniques applied to them; the critics have noted inconsistencies in findings among studies and even in independent re-analyses of data from the same city. In this paper we review some of the statistical methods used to analyze a subset of a national data base of air pollution, mortality and weather assembled during the National Morbidity and Mortality Air Pollution Study (NMMAPS).

    Release date: 2000-03-02

  • Surveys and statistical programs – Documentation: 11-522-X19990015668
    Description:

    Following the problems with estimating underenumeration in the 1991 Census of England and Wales the aim for the 2001 Census is to create a database that is fully adjusted to net underenumeration. To achieve this, the paper investigates weighted donor imputation methodology that utilises information from both the census and census coverage survey (CCS). The US Census Bureau has considered a similar approach for their 2000 Census (see Isaki et al 1998). The proposed procedure distinguishes between individuals who are not counted by the census because their household is missed and those who are missed in counted households. Census data is linked to data from the CCS. Multinomial logistic regression is used to estimate the probabilities that households are missed by the census and the probabilities that individuals are missed in counted households. Household and individual coverage weights are constructed from the estimated probabilities and these feed into the donor imputation procedure.

    Release date: 2000-03-02

  • Surveys and statistical programs – Documentation: 11-522-X19990015682
    Description:

    The application of dual system estimation (DSE) to matched Census / Post Enumeration Survey (PES) data in order to measure net undercount is well understood (Hogan, 1993). However, this approach has so far not been used to measure net undercount in the UK. The 2001 PES in the UK will use this methodology. This paper presents the general approach to design and estimation for this PES (the 2001 Census Coverage Survey). The estimation combines DSE with standard ratio and regression estimation. A simulation study using census data from the 1991 Census of England and Wales demonstrates that the ratio model is in general more robust than the regression model.

    Release date: 2000-03-02

  • Surveys and statistical programs – Documentation: 11-522-X19990015684
    Description:

    Often, the same information is gathered almost simultaneously for several different surveys. In France, this practice is institutionalized for household surveys that have a common set of demographic variables, i.e., employment, residence and income. These variables are important co-factors for the variables of interest in each survey, and if used carefully, can reinforce the estimates derived from each survey. Techniques for calibrating uncertain data can apply naturally in this context. This involves finding the best unbiased estimator in common variables and calibrating each survey based on that estimator. The estimator thus obtained in each survey is always a linear estimator, the weightings of which can be easily explained and the variance can be obtained with no new problems, as can the variance estimate. To supplement the list of regression estimators, this technique can also be seen as a ridge-regression estimator, or as a Bayesian-regression estimator.

    Release date: 2000-03-02

  • Surveys and statistical programs – Documentation: 11-522-X19990015688
    Description:

    The geographical and temporal relationship between outdoor air pollution and asthma was examined by linking together data from multiple sources. These included the administrative records of 59 general practices widely dispersed across England and Wales for half a million patients and all their consultations for asthma, supplemented by a socio-economic interview survey. Postcode enabled linkage with: (i) computed local road density; (ii) emission estimates of sulphur dioxide and nitrogen dioxides, (iii) measured/interpolated concentration of black smoke, sulphur dioxide, nitrogen dioxide and other pollutants at practice level. Parallel Poisson time series analysis took into account between-practice variations to examine daily correlations in practices close to air quality monitoring stations. Preliminary analyses show small and generally non-significant geographical associations between consultation rates and pollution markers. The methodological issues relevant to combining such data, and the interpretation of these results will be discussed.

    Release date: 2000-03-02

  • Surveys and statistical programs – Documentation: 11-522-X19990015692
    Description:

    Electricity rates that vary by time-of-day have the potential to significantly increase economic efficiency in the energy market. A number of utilities have undertaken economic studies of time-of-use rates schemes for their residential customers. This paper uses meta-analysis to examine the impact of time-of-use rates on electricity demand pooling the results of thirty-eight separate programs. There are four key findings. First, very large peak to off-peak price ratios are needed to significantly affect peak demand. Second, summer peak rates are relatively effective compared to winter peak rates. Third, permanent time-or-use rates are relatively effective compared to experimental ones. Fourth, demand charges rival ordinary time-of-use rates in terms of impact.

    Release date: 2000-03-02

  • Surveys and statistical programs – Documentation: 11-522-X19980015017
    Description:

    Longitudinal studies with repeated observations on individuals permit better characterizations of change and assessment of possible risk factors, but there has been little experience applying sophisticated models for longitudinal data to the complex survey setting. We present results from a comparison of different variance estimation methods for random effects models of change in cognitive function among older adults. The sample design is a stratified sample of people 65 and older, drawn as part of a community-based study designed to examine risk factors for dementia. The model summarizes the population heterogeneity in overall level and rate of change in cognitive function using random effects for intercept and slope. We discuss an unweighted regression including covariates for the stratification variables, a weighted regression, and bootstrapping; we also did preliminary work into using balanced repeated replication and jackknife repeated replication.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015029
    Description:

    In longitudinal surveys, sample subjects are observed over several time points. This feature typically leads to dependent observations on the same subject, in addition to the customary correlations across subjects induced by the sample design. Much research in the literature has focussed on modeling the marginal mean of a response as a function of covariates. Liang and Zeger (1986) used generalized estimating equations (GEE), requiring only correct specification of the marginal mean, and obtained standard errors of regression parameter estimates and associated Wald tests, assuming a "working" correlation structure for the repeated measurements on a sample subject. Rotnitzky and Jewell (1990) developed quasi-score tests and Rao-Scott adjustments to "working" quasi-score tests under marginal models. These methods are asymptotically robust to misspecification of the within-subject correlation structure, but assume independence of sample subjects which is not satisfied for complex longitudinal survey data based on stratified multi-stage sampling. We proposed asymptotically valid Wald and quasi-score tests for longitudinal survey data, using the Taylor Linearization and jackknife methods. Alternative tests, based on Rao-Scott adjustments to naive tests that ignore survey design features and on Bonferroni-t, are also developed. These tests are particularly useful when the effective degrees of freedom, usually taken as the total number of sample primary units (clusters) minus the number of strata, is small.

    Release date: 1999-10-22

  • Surveys and statistical programs – Documentation: 11-522-X19980015035
    Description:

    In a longitudinal survey conducted for k periods some units may be observed for less than k of the periods. Examples include, surveys designed with partially overlapping subsamples, a pure panel survey with nonresponse, and a panel survey supplemented with additional samples for some of the time periods. Estimators of the regression type are exhibited for such surveys. An application to special studies associated with the National Resources Inventory is discussed.

    Release date: 1999-10-22