Filter results by

Search Help
Currently selected filters that can be removed

Keyword(s)

Year of publication

1 facets displayed. 1 facets selected.

Survey or statistical program

104 facets displayed. 0 facets selected.

Content

1 facets displayed. 0 facets selected.
Sort Help
entries

Results

All (386)

All (386) (0 to 10 of 386 results)

  • Articles and reports: 12-001-X202400200001
    Description: Cochran’s rule states that a standard (Wald) two-sided 95% confidence interval around a sample mean drawn from a population with positive skewness is reasonable when the sample size is greater than 25 times the square of the skewness coefficient of the population. We investigate whether a variant of this crude rule applies for a proportion estimated from a stratified simple random sample.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200002
    Description: This paper investigates whether survey data quality fluctuates over the day. After laying out the argument theoretically, panel data from the Survey of Unemployed Workers in New Jersey are analyzed. Several indirect indicators of response error are investigated, including item nonresponse, interview completion time, rounding, and measures of the quality of time diary data. The evidence that we assemble for a time of day of interview effect is weak or nonexistent. Item nonresponse and the probability that interview completion time is among the 5% shortest appear to increase in the evening, but a more thorough assessment requires instrumental variables.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200003
    Description: The optimum sample allocation in stratified sampling is one of the basic issues of survey methodology. It is a procedure of dividing the overall sample size into strata sample sizes in such a way that for given sampling designs in strata the variance of the stratified \pi estimator of the population total (or mean) for a given study variable assumes its minimum. In this work, we consider the optimum allocation of a sample, under lower and upper bounds imposed jointly on sample sizes in strata. We are concerned with the variance function of some generic form that, in particular, covers the case of the simple random sampling without replacement in strata. The goal of this paper is twofold. First, we establish (using the Karush-Kuhn-Tucker conditions) a generic form of the optimal solution, the so-called optimality conditions. Second, based on the established optimality conditions, we derive an efficient recursive algorithm, named RNABOX, which solves the allocation problem under study. The RNABOX can be viewed as a generalization of the classical recursive Neyman allocation algorithm, a popular tool for optimum allocation when only upper bounds are imposed on sample strata-sizes. We implement RNABOX in R as a part of our package stratallo which is available from the Comprehensive R Archive Network (CRAN) repository.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200004
    Description: While we avoid specifying the parametric relationship between the study variable and covariates, we illustrate the advantage of including a spatial component to better account for the covariates in our models to make Bayesian predictive inference. We treat each unique covariate combination as an individual stratum, then we use small area estimation techniques to make inference about the finite population mean of the continuous response variable. The two spatial models used are the conditional autoregressive and simple conditional autoregressive models. We include the spatial effects by creating the adjacency matrix via the Mahalanobis distance between covariates. We also show how to incorporate survey weights into the spatial models when dealing with probability survey data. We compare the results of two non-spatial models including the Scott-Smith model and the Battese, Harter, and Fuller model to the spatial models. We illustrate the comparison between the aforementioned models with an application using BMI data from eight counties in California. Our goal is to have neighboring strata yield similar predictions, and to increase the difference between strata that are not neighbors. Ultimately, using the spatial models shows less global pooling compared to the non-spatial models, which was the desired outcome.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200005
    Description: Adaptive survey designs (ASDs) tailor recruitment protocols to population subgroups that are relevant to a survey. In recent years, effective ASD optimization has been the topic of research and several applications. However, the performance of an optimized ASD over time is sensitive to time changes in response propensities. How adaptation strategies can adjust to such variation over time is not yet fully understood. In this paper, we propose a robust optimization approach in the context of sequential mixed-mode surveys employing Bayesian analysis. The approach is formulated as a mathematical programming problem that explicitly accounts for uncertainty due to time change. ASD decisions can then be made by considering time-dependent variation in conditional mode response propensities and between-mode correlations in response propensities. The approach is demonstrated using a case study: the 2014-2017 Dutch Health Survey. We evaluate the sensitivity of ASD performance to 1) the budget level and 2) the length of applicable historic time-series data. We find there is only a moderate dependence on the budget level and the dependence on historic data is moderated by the amount of seasonality during the year.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200006
    Description: As mixed-mode designs become increasingly popular, their effects on data quality have attracted much scholarly attention. Most studies focused on the bias properties of mixed-mode designs; few of them have investigated whether mixed-mode designs have heterogeneous variance structures across modes. While many characteristics of mixed-mode designs, such as varied interviewer usage, systematic differences in respondents, varying levels of social desirability bias, among others, may lead to heterogeneous variances in mode-specific point estimates of population means, this study specifically investigates whether interviewer variances remain consistent across different modes in mixed-mode studies. To address this research question, we utilize data collected from two distinct study designs. In the first design, when interviewers are responsible for either face-to-face or telephone mode, we examine whether there are mode differences in interviewer variances for 1) sensitive political questions, 2) international items, 3) and item missing indicators on international items, using the Arab Barometer wave 6 Jordan data. In the second design, we draw on Health and Retirement Study (HRS) 2016 core survey data to examine the question on three topics when interviewers are responsible for both modes. The topics cover 1) the CESD depression scale, 2) interviewer observations, and 3) the physical activity scale. To account for the lack of interpenetrated designs in both data sources, we include respondent-level covariates in our models. We find significant differences in interviewer variances on one item (twelve items in total) in the Arab Barometer study; whereas for HRS, the results are three out of eighteen. Overall, we find the magnitude of the interviewer variances larger in FTF than TEL on sensitive items. We conduct simulations to understand the power to detect mode effects in the typically modest interviewer sample sizes.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200007
    Description: The capture-recapture method can be applied to measure the coverage of administrative and big data sources, in official statistics. In its basic form, it involves the linkage of two sources while assuming a perfect linkage and other standard assumptions. In practice, linkage errors arise and are a potential source of bias, where the linkage is based on quasi-identifiers. These errors include false positives and false negatives, where the former arise when linking a pair of records from different units, and the latter arise when not linking a pair of records from the same unit. So far, the existing solutions have resorted to costly clerical reviews, or they have made the restrictive conditional independence assumption. In this work, these requirements are relaxed by modeling the number of links from a record instead. The same approach may be taken to estimate the linkage accuracy without clerical reviews, when linking two sources that each have some undercoverage.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200008
    Description: When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our approach adheres to the general strategy proposed by Rubin (1993). Specifically, we generate pseudo-populations by applying the weighted finite population Bayesian bootstrap to account for survey weights, take simple random samples from those pseudo-populations, estimate synthesis models using these simple random samples, and release simulated data drawn from the models as public use files. To facilitate variance estimation, we use the framework of multiple imputation with two data generation strategies. In the first, we generate multiple data sets from each simple random sample. In the second, we generate a single synthetic data set from each simple random sample. We present multiple imputation combining rules for each setting. We illustrate the repeated sampling properties of the combining rules via simulation studies, including comparisons with synthetic data generation based on pseudo-likelihood methods. We apply the proposed methods to a subset of data from the American Community Survey.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200009
    Description: Many studies face the problem of comparing estimates obtained with different survey methodology, including differences in frames, measurement instruments, and modes of delivery. The problem arises in multimode surveys and in surveys that are redesigned. Major redesign of survey processes could affect survey estimates systematically, and it is important to quantify and adjust for such discontinuities between the designs to ensure comparability of estimates over time. We propose a small area estimation approach to reconcile two sets of survey estimates, and apply it to two surveys in the Marine Recreational Information Program (MRIP), which monitors recreational fishing along the Atlantic and Gulf coasts of the United States. We develop a log-normal model for the estimates from the two surveys, accounting for temporal dynamics through regression on population size and state-by-wave seasonal factors, and accounting in part for changing coverage properties through regression on wireless telephone penetration. Using the estimated design variances, we develop a regression model that is analytically consistent with the log-normal mean model. We use the modeled design variances in a Fay-Herriot small area estimation procedure to obtain empirical best linear unbiased predictors of the reconciled estimates of fishing effort (requiring predictions at new sets of covariates), and provide an asymptotically valid mean square error approximation.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200010
    Description: Recent work in survey domain estimation has shown that incorporating a priori assumptions about orderings of population domain means reduces the variance of the estimators and provides smaller confidence intervals with good coverage. Here we show how partial ordering assumptions allow design-based estimation of sample means in domains for which the sample size is zero, with conservative variance estimates and confidence intervals. Order restrictions can also substantially improve estimation and inference in small-size domains. Examples with well-known survey data sets demonstrate the utility of the methods. Code to implement the examples using the R package csurvey is given in the appendix.
    Release date: 2024-12-20
Stats in brief (91)

Stats in brief (91) (0 to 10 of 91 results)

  • Stats in brief: 11-627-M2024058
    Description: This infographic provides an overview of the diversity and demographic characteristics of the Muslim population in Canada. Using data from the 2001 and 2021 Census of Population data (2001 and 2021), it explores topics such as the distribution of the Muslim population by province and territory and by age group, the main racialized groups, the top countries of birth, and the top languages most often spoken at home by the Muslim population in Canada.
    Release date: 2024-12-16

  • Stats in brief: 11-627-M2024059
    Description: This infographic presents the key findings of access to services in the minority official language in Canada, based on the 2022 Survey on the Official Language Minority Population.
    Release date: 2024-12-16

  • Stats in brief: 11-627-M2024060
    Description: This infographic presents the main results concerning parents' intentions to enroll their children in an official language minority or majority school, according to the 2022 Survey on the Official Language Minority Population.
    Release date: 2024-12-16

  • Stats in brief: 11-627-M2024048
    Description: Using police-reported data from the 2023 Homicide Survey, this infographic is a visual representation of some of these data. Findings include results at the national, provincial and territorial levels. Also included are findings related to the characteristics of victims as well as the prevalence of gang-related and firearm-related homicides.
    Release date: 2024-12-11

  • Stats in brief: 45-20-00032024007
    Description: The strategies used by cyber attackers are getting more complex. Many of us are inundated with what feels like never-ending phishing emails, scam text messages and fraudulent phone calls. It’s rare to talk to someone who hasn’t experienced some form of a cyber attack. The situation is no different for Canadian businesses. Identity theft, scams, fraud and ransomware are some of the ways bad actors are targeting businesses today. We wanted to know: Is cyber crime on the rise in Canada? What is the relatively new phenomenon of cyber risk insurance? And in what way are consumers affected when a business experiences an security breach? The Canadian Survey of Cyber Security and Cyber Crime has published new data and in this episode, we sat down with Howard Bilodeau, an economist at Statistics Canada to answer our questions about how cyber security is changing for businesses and what it means for the rest of us.
    Release date: 2024-12-09

  • Stats in brief: 11-627-M2024057
    Description: This infographic looks at the poverty rates of different groups of older women (aged 65 years and older) in Canada. Using the 2021 Census of Population, it looks specifically at the low-income and poverty rates of different groups of older women, including older immigrant and racialized women.
    Release date: 2024-12-04

  • Stats in brief: 11-627-M2024054
    Description: Utilizing data from the 2022 Canadian Survey on Disability, this infographic highlights the trends and experiences of persons with dexterity disabilities. This release is part of a series of infographics that focus on specific disability types.
    Release date: 2024-12-03

  • Stats in brief: 11-627-M2024055
    Description: Utilizing data from the 2022 Canadian Survey on Disability, this infographic highlights the trends and experiences of persons with flexibility disabilities. This release is part of a series of infographics that focus on specific disability types.
    Release date: 2024-12-03

  • Stats in brief: 11-627-M2024056
    Description: Utilizing data from the 2022 Canadian Survey on Disability, this infographic highlights the trends and experiences of persons with mobility disabilities. This release is part of a series of infographics that focus on specific disability types.
    Release date: 2024-12-03

  • Stats in brief: 11-621-M2024016
    Description: This paper leverages administrative data to examine the distribution of the workforce by gender, age, and full-time work status from 2015 to 2023 for the performing arts, spectator sports and related industries, and the amusement and recreation industries. 
    Release date: 2024-11-28
Articles and reports (290)

Articles and reports (290) (0 to 10 of 290 results)

  • Articles and reports: 12-001-X202400200001
    Description: Cochran’s rule states that a standard (Wald) two-sided 95% confidence interval around a sample mean drawn from a population with positive skewness is reasonable when the sample size is greater than 25 times the square of the skewness coefficient of the population. We investigate whether a variant of this crude rule applies for a proportion estimated from a stratified simple random sample.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200002
    Description: This paper investigates whether survey data quality fluctuates over the day. After laying out the argument theoretically, panel data from the Survey of Unemployed Workers in New Jersey are analyzed. Several indirect indicators of response error are investigated, including item nonresponse, interview completion time, rounding, and measures of the quality of time diary data. The evidence that we assemble for a time of day of interview effect is weak or nonexistent. Item nonresponse and the probability that interview completion time is among the 5% shortest appear to increase in the evening, but a more thorough assessment requires instrumental variables.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200003
    Description: The optimum sample allocation in stratified sampling is one of the basic issues of survey methodology. It is a procedure of dividing the overall sample size into strata sample sizes in such a way that for given sampling designs in strata the variance of the stratified \pi estimator of the population total (or mean) for a given study variable assumes its minimum. In this work, we consider the optimum allocation of a sample, under lower and upper bounds imposed jointly on sample sizes in strata. We are concerned with the variance function of some generic form that, in particular, covers the case of the simple random sampling without replacement in strata. The goal of this paper is twofold. First, we establish (using the Karush-Kuhn-Tucker conditions) a generic form of the optimal solution, the so-called optimality conditions. Second, based on the established optimality conditions, we derive an efficient recursive algorithm, named RNABOX, which solves the allocation problem under study. The RNABOX can be viewed as a generalization of the classical recursive Neyman allocation algorithm, a popular tool for optimum allocation when only upper bounds are imposed on sample strata-sizes. We implement RNABOX in R as a part of our package stratallo which is available from the Comprehensive R Archive Network (CRAN) repository.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200004
    Description: While we avoid specifying the parametric relationship between the study variable and covariates, we illustrate the advantage of including a spatial component to better account for the covariates in our models to make Bayesian predictive inference. We treat each unique covariate combination as an individual stratum, then we use small area estimation techniques to make inference about the finite population mean of the continuous response variable. The two spatial models used are the conditional autoregressive and simple conditional autoregressive models. We include the spatial effects by creating the adjacency matrix via the Mahalanobis distance between covariates. We also show how to incorporate survey weights into the spatial models when dealing with probability survey data. We compare the results of two non-spatial models including the Scott-Smith model and the Battese, Harter, and Fuller model to the spatial models. We illustrate the comparison between the aforementioned models with an application using BMI data from eight counties in California. Our goal is to have neighboring strata yield similar predictions, and to increase the difference between strata that are not neighbors. Ultimately, using the spatial models shows less global pooling compared to the non-spatial models, which was the desired outcome.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200005
    Description: Adaptive survey designs (ASDs) tailor recruitment protocols to population subgroups that are relevant to a survey. In recent years, effective ASD optimization has been the topic of research and several applications. However, the performance of an optimized ASD over time is sensitive to time changes in response propensities. How adaptation strategies can adjust to such variation over time is not yet fully understood. In this paper, we propose a robust optimization approach in the context of sequential mixed-mode surveys employing Bayesian analysis. The approach is formulated as a mathematical programming problem that explicitly accounts for uncertainty due to time change. ASD decisions can then be made by considering time-dependent variation in conditional mode response propensities and between-mode correlations in response propensities. The approach is demonstrated using a case study: the 2014-2017 Dutch Health Survey. We evaluate the sensitivity of ASD performance to 1) the budget level and 2) the length of applicable historic time-series data. We find there is only a moderate dependence on the budget level and the dependence on historic data is moderated by the amount of seasonality during the year.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200006
    Description: As mixed-mode designs become increasingly popular, their effects on data quality have attracted much scholarly attention. Most studies focused on the bias properties of mixed-mode designs; few of them have investigated whether mixed-mode designs have heterogeneous variance structures across modes. While many characteristics of mixed-mode designs, such as varied interviewer usage, systematic differences in respondents, varying levels of social desirability bias, among others, may lead to heterogeneous variances in mode-specific point estimates of population means, this study specifically investigates whether interviewer variances remain consistent across different modes in mixed-mode studies. To address this research question, we utilize data collected from two distinct study designs. In the first design, when interviewers are responsible for either face-to-face or telephone mode, we examine whether there are mode differences in interviewer variances for 1) sensitive political questions, 2) international items, 3) and item missing indicators on international items, using the Arab Barometer wave 6 Jordan data. In the second design, we draw on Health and Retirement Study (HRS) 2016 core survey data to examine the question on three topics when interviewers are responsible for both modes. The topics cover 1) the CESD depression scale, 2) interviewer observations, and 3) the physical activity scale. To account for the lack of interpenetrated designs in both data sources, we include respondent-level covariates in our models. We find significant differences in interviewer variances on one item (twelve items in total) in the Arab Barometer study; whereas for HRS, the results are three out of eighteen. Overall, we find the magnitude of the interviewer variances larger in FTF than TEL on sensitive items. We conduct simulations to understand the power to detect mode effects in the typically modest interviewer sample sizes.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200007
    Description: The capture-recapture method can be applied to measure the coverage of administrative and big data sources, in official statistics. In its basic form, it involves the linkage of two sources while assuming a perfect linkage and other standard assumptions. In practice, linkage errors arise and are a potential source of bias, where the linkage is based on quasi-identifiers. These errors include false positives and false negatives, where the former arise when linking a pair of records from different units, and the latter arise when not linking a pair of records from the same unit. So far, the existing solutions have resorted to costly clerical reviews, or they have made the restrictive conditional independence assumption. In this work, these requirements are relaxed by modeling the number of links from a record instead. The same approach may be taken to estimate the linkage accuracy without clerical reviews, when linking two sources that each have some undercoverage.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200008
    Description: When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our approach adheres to the general strategy proposed by Rubin (1993). Specifically, we generate pseudo-populations by applying the weighted finite population Bayesian bootstrap to account for survey weights, take simple random samples from those pseudo-populations, estimate synthesis models using these simple random samples, and release simulated data drawn from the models as public use files. To facilitate variance estimation, we use the framework of multiple imputation with two data generation strategies. In the first, we generate multiple data sets from each simple random sample. In the second, we generate a single synthetic data set from each simple random sample. We present multiple imputation combining rules for each setting. We illustrate the repeated sampling properties of the combining rules via simulation studies, including comparisons with synthetic data generation based on pseudo-likelihood methods. We apply the proposed methods to a subset of data from the American Community Survey.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200009
    Description: Many studies face the problem of comparing estimates obtained with different survey methodology, including differences in frames, measurement instruments, and modes of delivery. The problem arises in multimode surveys and in surveys that are redesigned. Major redesign of survey processes could affect survey estimates systematically, and it is important to quantify and adjust for such discontinuities between the designs to ensure comparability of estimates over time. We propose a small area estimation approach to reconcile two sets of survey estimates, and apply it to two surveys in the Marine Recreational Information Program (MRIP), which monitors recreational fishing along the Atlantic and Gulf coasts of the United States. We develop a log-normal model for the estimates from the two surveys, accounting for temporal dynamics through regression on population size and state-by-wave seasonal factors, and accounting in part for changing coverage properties through regression on wireless telephone penetration. Using the estimated design variances, we develop a regression model that is analytically consistent with the log-normal mean model. We use the modeled design variances in a Fay-Herriot small area estimation procedure to obtain empirical best linear unbiased predictors of the reconciled estimates of fishing effort (requiring predictions at new sets of covariates), and provide an asymptotically valid mean square error approximation.
    Release date: 2024-12-20

  • Articles and reports: 12-001-X202400200010
    Description: Recent work in survey domain estimation has shown that incorporating a priori assumptions about orderings of population domain means reduces the variance of the estimators and provides smaller confidence intervals with good coverage. Here we show how partial ordering assumptions allow design-based estimation of sample means in domains for which the sample size is zero, with conservative variance estimates and confidence intervals. Order restrictions can also substantially improve estimation and inference in small-size domains. Examples with well-known survey data sets demonstrate the utility of the methods. Code to implement the examples using the R package csurvey is given in the appendix.
    Release date: 2024-12-20
Journals and periodicals (5)

Journals and periodicals (5) ((5 results))

  • Journals and periodicals: 89-28-0001
    Description: Short and focused data tables related to current events.
    Release date: 2024-09-25

  • Journals and periodicals: 91-215-X
    Description: This publication presents annual estimates of the total population and annual estimates by age and gender for Canada, provinces and territories. It also presents estimates of the following components of population change: births, deaths, immigration, emigration, returning emigration, net non-permanent residents and inter-provincial migration, the latter by origin and destination. As in the case of population estimates, the components are also available for the total population and by age and gender.

    The Annual demographic estimates - Canada, provinces and territories publication contains the most recent estimates as well as an annual historical series. It also contains highlights and analysis of the most current demographic trends, as well as a brief description of the concepts, methods and data quality of the estimates.

    Release date: 2024-09-25

  • Journals and periodicals: 96-325-X
    Geography: Canada
    Description: This publication features short and accessible analytical articles that delve further into key findings and emerging trends identified in Census of Agriculture and other data sources related to agriculture. Subjects of analysis include matters related to farm land, crops, livestock, farm finances, technology, the environment and the farm population, as well as other economic and social aspects of Canada’s agriculture industry. Analytical articles are written in plain language and are intended to be a valuable source of information for a broad audience, including policy analysts, students, researchers, agricultural operators, the media and the public at large.
    Release date: 2024-03-07

  • Journals and periodicals: 75-004-M
    Geography: Canada
    Description: The papers in this series cover a variety of topics related to labour statistics. The studies are intended to show recent or historical trends observed with the surveys produced by the Centre for Labour Market Information, i.e. the Labour Force Survey, Survey of Employment Payrolls and Hours, Employment insurance Coverage Survey, Employment insurance statistics as well as administrative data sources. All the papers in this analytical series go through institutional and peer review to ensure that they conform to Statistics Canada's mandate as a government statistical agency and adhere to generally accepted standards of good professional practice.
    Release date: 2024-03-04

  • Journals and periodicals: 98-200-X
    Description: These short analytical articles, based on data from the Census of Population, provide analysis on specific topics of interest related to the Canadian population. They are available with each Census of Population major release.
    Release date: 2024-02-28