Keyword search
Filter results by
Search HelpCurrently selected filters that can be removed
Keyword(s)
Subject
- Government (3)
- Income, pensions, spending and wealth (20)
- International trade (4)
- Health (56)
- Labour (82)
- Languages (16)
- Manufacturing (2)
- Population and demography (11)
- Prices and price indexes (7)
- Statistical methods (69)
- Retail and wholesale (2)
- Business and consumer services and culture (5)
- Digital economy and society (5)
- Transportation (5)
- Travel and tourism (5)
- Energy (1)
- Science and technology (16)
- Agriculture and food (9)
- Business performance and ownership (27)
- Construction (2)
- Crime and justice (12)
- Economic accounts (15)
- Education, training and learning (22)
- Environment (15)
- Families, households and marital status (5)
- Indigenous peoples (17)
- Children and youth (13)
- Immigration and ethnocultural diversity (39)
- Older adults and population aging (7)
- Society and community (40)
- Housing (12)
Type
9 facets displayed. 0 facets selected.
Geography
3 facets displayed. 0 facets selected.
Survey or statistical program
107 facets displayed. 0 facets selected.
- Census of Population (57)
- Labour Force Survey (26)
- Canadian Survey on Disability (12)
- Canadian Survey on Business Conditions (10)
- Canadian System of Environmental-Economic Accounts - Ecosystem Accounts (8)
- Canadian Community Health Survey - Annual Component (7)
- Uniform Crime Reporting Survey (7)
- Labour Market Indicators (7)
- Postsecondary Student Information System (6)
- Census of Agriculture (5)
- Canadian Income Survey (5)
- Business Innovation and Growth Support (5)
- National Gross Domestic Product by Income and by Expenditure Accounts (4)
- Consumer Price Index (4)
- Integrated Criminal Court Survey (4)
- Annual Income Estimates for Census Families and Individuals (T1 Family File) (4)
- Canadian Health Measures Survey (4)
- Supply, Use and Input-Output Tables (3)
- Corporations Returns Act (3)
- Registered Apprenticeship Information System (3)
- Indigenous Peoples Survey (3)
- General Social Survey - Victimization (3)
- Frontier Counts (3)
- General Social Survey - Social Identity (3)
- Longitudinal Immigration Database (3)
- Canadian Housing Statistics Program (3)
- Canadian Health Survey on Seniors (3)
- Canadian Social Survey (3)
- Vital Statistics - Birth Database (2)
- Vital Statistics - Death Database (2)
- Homicide Survey (2)
- Survey of Residential Facilities for Victims of Abuse (2)
- Quarterly Demographic Estimates (2)
- Annual Demographic Estimates: Canada, Provinces and Territories (2)
- Canadian Internet Use Survey (2)
- General Social Survey - Family (2)
- Time Use Survey (2)
- Canadian System of Environmental-Economic Accounts - Natural Resource Asset Accounts (2)
- Survey of Innovation and Business Strategy (2)
- National Household Survey (2)
- Job Vacancy and Wage Survey (2)
- Canadian Employer-Employee Dynamics Database (2)
- Canadian Health Survey on Children and Youth (2)
- Survey of Safety in Public and Private Spaces (2)
- Visitor Travel Survey (2)
- The Open Database of Educational Facilities (2)
- Survey on the Official Language Minority Population (SOLMP) (2)
- Linkable File Environment (2)
- Gross Domestic Product by Industry - National (Monthly) (1)
- Gross Domestic Product by Industry - Provincial and Territorial (Annual) (1)
- Canada's International Transactions in Services (1)
- Canada's International Investment Position (1)
- National Tourism Indicators (1)
- Biennial Waste Management Survey (1)
- Monthly Wholesale Trade Survey (1)
- Annual Survey of Service Industries: Software Development and Computer Services (1)
- Survey of Service Industries: Film, Television and Video Production (1)
- Survey of Service Industries: Film and Video Distribution (1)
- Survey of Service Industries: Film, Television and Video Post-production (1)
- Survey of Service Industries: Motion Picture Theatres (1)
- Annual Survey of Service Industries: Accommodation Services (1)
- Annual Survey of Service Industries: Architectural Services (1)
- Annual Survey of Service Industries: Travel Arrangement Services (1)
- Annual Survey of Service Industries: Consumer Goods Rental (1)
- Annual Survey of Service Industries: Engineering Services (1)
- Annual Survey of Service Industries: Commercial and Industrial Machinery and Equipment Rental and Leasing (1)
- Quarterly Survey of Financial Statements (1)
- Employment Insurance Statistics - Monthly (1)
- Survey of Employment, Payrolls and Hours (1)
- Survey of Financial Security (1)
- Airport Activity Survey (1)
- Survey of Service Industries: Performing Arts (1)
- Survey of Service Industries: Sound Recording and Music Publishing (1)
- Annual Demographic Estimates : Subprovincial Areas (1)
- Households and the Environment Survey (1)
- Survey of Advanced Technology (1)
- General Social Survey - Caregiving and Care Receiving (1)
- Annual Survey of Service Industries: Accounting Services (1)
- Annual Survey of Service Industries: Consulting Services (1)
- Annual Survey of Service Industries: Employment Services (1)
- Annual Survey of Service Industries: Specialized Design (1)
- National Graduates Survey (1)
- Mental Health and Access to Care Survey (MHACS) (1)
- Aboriginal Children's Survey (1)
- Canadian System of Environmental-Economic Accounts - Physical Flow Accounts (1)
- Annual Survey of Service Industries: Spectator Sports, Event Promoters, Artists and Related Industries (1)
- Canada's Core Public Infrastructure Survey (1)
- General Social Survey: Canadians at Work and Home (1)
- National Travel Survey (1)
- Canadian Survey of Cyber Security and Cybercrime (1)
- Canadian Correctional Services Survey (1)
- Canadian Housing Survey (1)
- Indigenous Peoples Survey - Nunavut Inuit Supplement (1)
- Wastewater-based estimates of drug consumption (1)
- Gender Statistics (1)
- New Motor Vehicle Registration Survey (1)
- COVID-19 epidemiological reports (1)
- Canadian Survey on the Provision of Child Care Services (1)
- Survey on Access to Health Care and Pharmaceuticals During the Pandemic (1)
- House of Commons Canada (1)
- Band Governance Management System (5351) (1)
- Survey on Health Care Workers' Experiences During the Pandemic (1)
- Office of the Commissioner for Federal Judicial Affairs Canada (FJA) (1)
- Distributions of household economic accounts for wealth of Canadian households (1)
- Labour Market and Socio-economic Indicators (1)
- Survey Series on People and their Communities (1)
- Survey on Early Learning and Child Care Arrangements - Children with Long-term Conditions and Disabilities (SELCCA - CLCD) (1)
Results
All (358)
All (358) (0 to 10 of 358 results)
- Articles and reports: 12-001-X202400200001Description: Cochran’s rule states that a standard (Wald) two-sided 95% confidence interval around a sample mean drawn from a population with positive skewness is reasonable when the sample size is greater than 25 times the square of the skewness coefficient of the population. We investigate whether a variant of this crude rule applies for a proportion estimated from a stratified simple random sample.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200002Description: This paper investigates whether survey data quality fluctuates over the day. After laying out the argument theoretically, panel data from the Survey of Unemployed Workers in New Jersey are analyzed. Several indirect indicators of response error are investigated, including item nonresponse, interview completion time, rounding, and measures of the quality of time diary data. The evidence that we assemble for a time of day of interview effect is weak or nonexistent. Item nonresponse and the probability that interview completion time is among the 5% shortest appear to increase in the evening, but a more thorough assessment requires instrumental variables.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200003Description: The optimum sample allocation in stratified sampling is one of the basic issues of survey methodology. It is a procedure of dividing the overall sample size into strata sample sizes in such a way that for given sampling designs in strata the variance of the stratified \pi estimator of the population total (or mean) for a given study variable assumes its minimum. In this work, we consider the optimum allocation of a sample, under lower and upper bounds imposed jointly on sample sizes in strata. We are concerned with the variance function of some generic form that, in particular, covers the case of the simple random sampling without replacement in strata. The goal of this paper is twofold. First, we establish (using the Karush-Kuhn-Tucker conditions) a generic form of the optimal solution, the so-called optimality conditions. Second, based on the established optimality conditions, we derive an efficient recursive algorithm, named RNABOX, which solves the allocation problem under study. The RNABOX can be viewed as a generalization of the classical recursive Neyman allocation algorithm, a popular tool for optimum allocation when only upper bounds are imposed on sample strata-sizes. We implement RNABOX in R as a part of our package stratallo which is available from the Comprehensive R Archive Network (CRAN) repository.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200004Description: While we avoid specifying the parametric relationship between the study variable and covariates, we illustrate the advantage of including a spatial component to better account for the covariates in our models to make Bayesian predictive inference. We treat each unique covariate combination as an individual stratum, then we use small area estimation techniques to make inference about the finite population mean of the continuous response variable. The two spatial models used are the conditional autoregressive and simple conditional autoregressive models. We include the spatial effects by creating the adjacency matrix via the Mahalanobis distance between covariates. We also show how to incorporate survey weights into the spatial models when dealing with probability survey data. We compare the results of two non-spatial models including the Scott-Smith model and the Battese, Harter, and Fuller model to the spatial models. We illustrate the comparison between the aforementioned models with an application using BMI data from eight counties in California. Our goal is to have neighboring strata yield similar predictions, and to increase the difference between strata that are not neighbors. Ultimately, using the spatial models shows less global pooling compared to the non-spatial models, which was the desired outcome.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200005Description: Adaptive survey designs (ASDs) tailor recruitment protocols to population subgroups that are relevant to a survey. In recent years, effective ASD optimization has been the topic of research and several applications. However, the performance of an optimized ASD over time is sensitive to time changes in response propensities. How adaptation strategies can adjust to such variation over time is not yet fully understood. In this paper, we propose a robust optimization approach in the context of sequential mixed-mode surveys employing Bayesian analysis. The approach is formulated as a mathematical programming problem that explicitly accounts for uncertainty due to time change. ASD decisions can then be made by considering time-dependent variation in conditional mode response propensities and between-mode correlations in response propensities. The approach is demonstrated using a case study: the 2014-2017 Dutch Health Survey. We evaluate the sensitivity of ASD performance to 1) the budget level and 2) the length of applicable historic time-series data. We find there is only a moderate dependence on the budget level and the dependence on historic data is moderated by the amount of seasonality during the year.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200006Description: As mixed-mode designs become increasingly popular, their effects on data quality have attracted much scholarly attention. Most studies focused on the bias properties of mixed-mode designs; few of them have investigated whether mixed-mode designs have heterogeneous variance structures across modes. While many characteristics of mixed-mode designs, such as varied interviewer usage, systematic differences in respondents, varying levels of social desirability bias, among others, may lead to heterogeneous variances in mode-specific point estimates of population means, this study specifically investigates whether interviewer variances remain consistent across different modes in mixed-mode studies. To address this research question, we utilize data collected from two distinct study designs. In the first design, when interviewers are responsible for either face-to-face or telephone mode, we examine whether there are mode differences in interviewer variances for 1) sensitive political questions, 2) international items, 3) and item missing indicators on international items, using the Arab Barometer wave 6 Jordan data. In the second design, we draw on Health and Retirement Study (HRS) 2016 core survey data to examine the question on three topics when interviewers are responsible for both modes. The topics cover 1) the CESD depression scale, 2) interviewer observations, and 3) the physical activity scale. To account for the lack of interpenetrated designs in both data sources, we include respondent-level covariates in our models. We find significant differences in interviewer variances on one item (twelve items in total) in the Arab Barometer study; whereas for HRS, the results are three out of eighteen. Overall, we find the magnitude of the interviewer variances larger in FTF than TEL on sensitive items. We conduct simulations to understand the power to detect mode effects in the typically modest interviewer sample sizes.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200007Description: The capture-recapture method can be applied to measure the coverage of administrative and big data sources, in official statistics. In its basic form, it involves the linkage of two sources while assuming a perfect linkage and other standard assumptions. In practice, linkage errors arise and are a potential source of bias, where the linkage is based on quasi-identifiers. These errors include false positives and false negatives, where the former arise when linking a pair of records from different units, and the latter arise when not linking a pair of records from the same unit. So far, the existing solutions have resorted to costly clerical reviews, or they have made the restrictive conditional independence assumption. In this work, these requirements are relaxed by modeling the number of links from a record instead. The same approach may be taken to estimate the linkage accuracy without clerical reviews, when linking two sources that each have some undercoverage.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200008Description: When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our approach adheres to the general strategy proposed by Rubin (1993). Specifically, we generate pseudo-populations by applying the weighted finite population Bayesian bootstrap to account for survey weights, take simple random samples from those pseudo-populations, estimate synthesis models using these simple random samples, and release simulated data drawn from the models as public use files. To facilitate variance estimation, we use the framework of multiple imputation with two data generation strategies. In the first, we generate multiple data sets from each simple random sample. In the second, we generate a single synthetic data set from each simple random sample. We present multiple imputation combining rules for each setting. We illustrate the repeated sampling properties of the combining rules via simulation studies, including comparisons with synthetic data generation based on pseudo-likelihood methods. We apply the proposed methods to a subset of data from the American Community Survey.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200009Description: Many studies face the problem of comparing estimates obtained with different survey methodology, including differences in frames, measurement instruments, and modes of delivery. The problem arises in multimode surveys and in surveys that are redesigned. Major redesign of survey processes could affect survey estimates systematically, and it is important to quantify and adjust for such discontinuities between the designs to ensure comparability of estimates over time. We propose a small area estimation approach to reconcile two sets of survey estimates, and apply it to two surveys in the Marine Recreational Information Program (MRIP), which monitors recreational fishing along the Atlantic and Gulf coasts of the United States. We develop a log-normal model for the estimates from the two surveys, accounting for temporal dynamics through regression on population size and state-by-wave seasonal factors, and accounting in part for changing coverage properties through regression on wireless telephone penetration. Using the estimated design variances, we develop a regression model that is analytically consistent with the log-normal mean model. We use the modeled design variances in a Fay-Herriot small area estimation procedure to obtain empirical best linear unbiased predictors of the reconciled estimates of fishing effort (requiring predictions at new sets of covariates), and provide an asymptotically valid mean square error approximation.Release date: 2024-12-20
- 10. Design-based estimation of small and empty domains in survey data analysis using order constraintsArticles and reports: 12-001-X202400200010Description: Recent work in survey domain estimation has shown that incorporating a priori assumptions about orderings of population domain means reduces the variance of the estimators and provides smaller confidence intervals with good coverage. Here we show how partial ordering assumptions allow design-based estimation of sample means in domains for which the sample size is zero, with conservative variance estimates and confidence intervals. Order restrictions can also substantially improve estimation and inference in small-size domains. Examples with well-known survey data sets demonstrate the utility of the methods. Code to implement the examples using the R package csurvey is given in the appendix.Release date: 2024-12-20
- Previous Go to previous page of All results
- 1 (current) Go to page 1 of All results
- 2 Go to page 2 of All results
- 3 Go to page 3 of All results
- 4 Go to page 4 of All results
- 5 Go to page 5 of All results
- 6 Go to page 6 of All results
- 7 Go to page 7 of All results
- ...
- 36 Go to page 36 of All results
- Next Go to next page of All results
Data (22)
Data (22) (0 to 10 of 22 results)
- Table: 37-26-0001Description: The Open Database of Educational Facilities (ODEF) is a compilation of data from open and internet sources on the locations and types of educational facilities across Canada, originating from municipal, regional, and provincial governments. It is a centralized and harmonized repository of educational facility data made available under the Open Government License - Canada. The database is expected to be updated periodically as new open datasets from government sources become available. The database is made available for download as a zipped comma separated values (csv) file.Release date: 2024-12-13
- Table: 45-20-00042024004Description: Rural Canada Business Profiles is a database that provides financial profiles for Small and Medium-Sized Businesses in Canada with total annual revenues of $ 30,000 to $ 5,000,000 and $ 5,000,001 to $ 20,000,000 respectively. These data are available by industry, by province or territory, by legal status of businesses (incorporated and unincorporated) and the distinction of 'Rural and small town (RST)' or 'functional urban' location of businesses. Data released is for 2022.Release date: 2024-12-09
- Data Visualization: 71-607-X2022002Description: Environmental, social and governance (ESG) refers to three non-financial factors that can be used to inform the long-term risk and return of an investment. ESG are emerging as a priority for governments, businesses and international organisations. This experimental dashboard offers an overview of the performance over time of selected industries with respect to a collection of ESG indicators.Release date: 2024-11-20
- Table: 34-26-0003Description: The Open Database of Infrastructure (ODI) contains the locations of bridges, tunnels, solid waste facilities, pedestrian and cycling paths, public transit stops, and potable water, stormwater and wastewater infrastructure. The ODI is compiled from both open and publicly available data sources and is made available under the Open Government License - Canada. This database is a component of the Linkable Open Data Environment (LODE).Release date: 2024-11-13
- Table: 13-26-0003Description: In collaboration with the Public Health Agency of Canada (PHAC), this data file provides Canadians and researchers with data to monitor only the confirmed cases of coronavirus (COVID-19) in Canada.Release date: 2024-10-11
- Data Visualization: 71-607-X2024026Description: The Business Ownership Diversity Dashboard allows users to examine the distribution of businesses in Canada by equity and diversity indicators and by business variable dimensions. Equity and diversity indicators include visible minority status, age, immigrant status, Indigenous group, and gender. Business variable dimensions include the location of business operations, revenue size, business size (number of employees), and industry sector.Release date: 2024-09-12
- Data Visualization: 71-607-X2024021Description: This dashboard presents provisional monthly estimates of the levels of amphetamine, cannabis, cocaine (benzoylecgonine), codeine, fentanyl (norfentanyl), ecstasy, methadone, methamphetamine, morphine, and oxycodone in the wastewater of Halifax, Montréal, Toronto, Saskatoon, Prince Albert, Edmonton, and Metro Vancouver. The data that are relevant for monitoring the use of these substances in Canadian cities.Release date: 2024-09-06
- Table: 12-581-XDescription: Canada at a Glance presents current statistics on Canadian society, including subjects such as the population, education, health, prices and the economy, among others. Updated yearly, this booklet is a very useful reference for those who want quick access to a current statistical portrait of Canada.Release date: 2024-09-04
- 9. Age pyramidsData Visualization: 98-504-X2021001Description: Age pyramids are dynamic applications that allow users to see the evolution of the age structure of the Canadian population over a given time period and for selected geographies.Release date: 2024-06-26
- Data Visualization: 98-504-XDescription: Age pyramids are dynamic applications that allow users to see the evolution of the age structure of the Canadian population over a given time period and for selected geographies.Release date: 2024-06-26
Analysis (311)
Analysis (311) (0 to 10 of 311 results)
- Articles and reports: 12-001-X202400200001Description: Cochran’s rule states that a standard (Wald) two-sided 95% confidence interval around a sample mean drawn from a population with positive skewness is reasonable when the sample size is greater than 25 times the square of the skewness coefficient of the population. We investigate whether a variant of this crude rule applies for a proportion estimated from a stratified simple random sample.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200002Description: This paper investigates whether survey data quality fluctuates over the day. After laying out the argument theoretically, panel data from the Survey of Unemployed Workers in New Jersey are analyzed. Several indirect indicators of response error are investigated, including item nonresponse, interview completion time, rounding, and measures of the quality of time diary data. The evidence that we assemble for a time of day of interview effect is weak or nonexistent. Item nonresponse and the probability that interview completion time is among the 5% shortest appear to increase in the evening, but a more thorough assessment requires instrumental variables.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200003Description: The optimum sample allocation in stratified sampling is one of the basic issues of survey methodology. It is a procedure of dividing the overall sample size into strata sample sizes in such a way that for given sampling designs in strata the variance of the stratified \pi estimator of the population total (or mean) for a given study variable assumes its minimum. In this work, we consider the optimum allocation of a sample, under lower and upper bounds imposed jointly on sample sizes in strata. We are concerned with the variance function of some generic form that, in particular, covers the case of the simple random sampling without replacement in strata. The goal of this paper is twofold. First, we establish (using the Karush-Kuhn-Tucker conditions) a generic form of the optimal solution, the so-called optimality conditions. Second, based on the established optimality conditions, we derive an efficient recursive algorithm, named RNABOX, which solves the allocation problem under study. The RNABOX can be viewed as a generalization of the classical recursive Neyman allocation algorithm, a popular tool for optimum allocation when only upper bounds are imposed on sample strata-sizes. We implement RNABOX in R as a part of our package stratallo which is available from the Comprehensive R Archive Network (CRAN) repository.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200004Description: While we avoid specifying the parametric relationship between the study variable and covariates, we illustrate the advantage of including a spatial component to better account for the covariates in our models to make Bayesian predictive inference. We treat each unique covariate combination as an individual stratum, then we use small area estimation techniques to make inference about the finite population mean of the continuous response variable. The two spatial models used are the conditional autoregressive and simple conditional autoregressive models. We include the spatial effects by creating the adjacency matrix via the Mahalanobis distance between covariates. We also show how to incorporate survey weights into the spatial models when dealing with probability survey data. We compare the results of two non-spatial models including the Scott-Smith model and the Battese, Harter, and Fuller model to the spatial models. We illustrate the comparison between the aforementioned models with an application using BMI data from eight counties in California. Our goal is to have neighboring strata yield similar predictions, and to increase the difference between strata that are not neighbors. Ultimately, using the spatial models shows less global pooling compared to the non-spatial models, which was the desired outcome.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200005Description: Adaptive survey designs (ASDs) tailor recruitment protocols to population subgroups that are relevant to a survey. In recent years, effective ASD optimization has been the topic of research and several applications. However, the performance of an optimized ASD over time is sensitive to time changes in response propensities. How adaptation strategies can adjust to such variation over time is not yet fully understood. In this paper, we propose a robust optimization approach in the context of sequential mixed-mode surveys employing Bayesian analysis. The approach is formulated as a mathematical programming problem that explicitly accounts for uncertainty due to time change. ASD decisions can then be made by considering time-dependent variation in conditional mode response propensities and between-mode correlations in response propensities. The approach is demonstrated using a case study: the 2014-2017 Dutch Health Survey. We evaluate the sensitivity of ASD performance to 1) the budget level and 2) the length of applicable historic time-series data. We find there is only a moderate dependence on the budget level and the dependence on historic data is moderated by the amount of seasonality during the year.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200006Description: As mixed-mode designs become increasingly popular, their effects on data quality have attracted much scholarly attention. Most studies focused on the bias properties of mixed-mode designs; few of them have investigated whether mixed-mode designs have heterogeneous variance structures across modes. While many characteristics of mixed-mode designs, such as varied interviewer usage, systematic differences in respondents, varying levels of social desirability bias, among others, may lead to heterogeneous variances in mode-specific point estimates of population means, this study specifically investigates whether interviewer variances remain consistent across different modes in mixed-mode studies. To address this research question, we utilize data collected from two distinct study designs. In the first design, when interviewers are responsible for either face-to-face or telephone mode, we examine whether there are mode differences in interviewer variances for 1) sensitive political questions, 2) international items, 3) and item missing indicators on international items, using the Arab Barometer wave 6 Jordan data. In the second design, we draw on Health and Retirement Study (HRS) 2016 core survey data to examine the question on three topics when interviewers are responsible for both modes. The topics cover 1) the CESD depression scale, 2) interviewer observations, and 3) the physical activity scale. To account for the lack of interpenetrated designs in both data sources, we include respondent-level covariates in our models. We find significant differences in interviewer variances on one item (twelve items in total) in the Arab Barometer study; whereas for HRS, the results are three out of eighteen. Overall, we find the magnitude of the interviewer variances larger in FTF than TEL on sensitive items. We conduct simulations to understand the power to detect mode effects in the typically modest interviewer sample sizes.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200007Description: The capture-recapture method can be applied to measure the coverage of administrative and big data sources, in official statistics. In its basic form, it involves the linkage of two sources while assuming a perfect linkage and other standard assumptions. In practice, linkage errors arise and are a potential source of bias, where the linkage is based on quasi-identifiers. These errors include false positives and false negatives, where the former arise when linking a pair of records from different units, and the latter arise when not linking a pair of records from the same unit. So far, the existing solutions have resorted to costly clerical reviews, or they have made the restrictive conditional independence assumption. In this work, these requirements are relaxed by modeling the number of links from a record instead. The same approach may be taken to estimate the linkage accuracy without clerical reviews, when linking two sources that each have some undercoverage.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200008Description: When seeking to release public use files for confidential data, statistical agencies can generate fully synthetic data. We propose an approach for making fully synthetic data from surveys collected with complex sampling designs. Our approach adheres to the general strategy proposed by Rubin (1993). Specifically, we generate pseudo-populations by applying the weighted finite population Bayesian bootstrap to account for survey weights, take simple random samples from those pseudo-populations, estimate synthesis models using these simple random samples, and release simulated data drawn from the models as public use files. To facilitate variance estimation, we use the framework of multiple imputation with two data generation strategies. In the first, we generate multiple data sets from each simple random sample. In the second, we generate a single synthetic data set from each simple random sample. We present multiple imputation combining rules for each setting. We illustrate the repeated sampling properties of the combining rules via simulation studies, including comparisons with synthetic data generation based on pseudo-likelihood methods. We apply the proposed methods to a subset of data from the American Community Survey.Release date: 2024-12-20
- Articles and reports: 12-001-X202400200009Description: Many studies face the problem of comparing estimates obtained with different survey methodology, including differences in frames, measurement instruments, and modes of delivery. The problem arises in multimode surveys and in surveys that are redesigned. Major redesign of survey processes could affect survey estimates systematically, and it is important to quantify and adjust for such discontinuities between the designs to ensure comparability of estimates over time. We propose a small area estimation approach to reconcile two sets of survey estimates, and apply it to two surveys in the Marine Recreational Information Program (MRIP), which monitors recreational fishing along the Atlantic and Gulf coasts of the United States. We develop a log-normal model for the estimates from the two surveys, accounting for temporal dynamics through regression on population size and state-by-wave seasonal factors, and accounting in part for changing coverage properties through regression on wireless telephone penetration. Using the estimated design variances, we develop a regression model that is analytically consistent with the log-normal mean model. We use the modeled design variances in a Fay-Herriot small area estimation procedure to obtain empirical best linear unbiased predictors of the reconciled estimates of fishing effort (requiring predictions at new sets of covariates), and provide an asymptotically valid mean square error approximation.Release date: 2024-12-20
- 10. Design-based estimation of small and empty domains in survey data analysis using order constraintsArticles and reports: 12-001-X202400200010Description: Recent work in survey domain estimation has shown that incorporating a priori assumptions about orderings of population domain means reduces the variance of the estimators and provides smaller confidence intervals with good coverage. Here we show how partial ordering assumptions allow design-based estimation of sample means in domains for which the sample size is zero, with conservative variance estimates and confidence intervals. Order restrictions can also substantially improve estimation and inference in small-size domains. Examples with well-known survey data sets demonstrate the utility of the methods. Code to implement the examples using the R package csurvey is given in the appendix.Release date: 2024-12-20
- Previous Go to previous page of Analysis results
- 1 (current) Go to page 1 of Analysis results
- 2 Go to page 2 of Analysis results
- 3 Go to page 3 of Analysis results
- 4 Go to page 4 of Analysis results
- 5 Go to page 5 of Analysis results
- 6 Go to page 6 of Analysis results
- 7 Go to page 7 of Analysis results
- ...
- 32 Go to page 32 of Analysis results
- Next Go to next page of Analysis results
Reference (25)
Reference (25) (0 to 10 of 25 results)
- Geographic files and documentation: 16-510-X2024005Description: This product contains specifications intended for users of the ocean and coastal ecosystem extent geospatial files. This document provides important technical information for users and links to methodology.Release date: 2024-12-16
- Geographic files and documentation: 16-510-X2024006Description: This product contains gridded datasets of ocean and coastal ecosystem extent for ocean and coastal areas of Canada. Mapping the extent of ocean ecosystems is the first stage in creating spatially explicit accounts, to help understand the ocean and coastal ecosystems of Canada. The files cover the area from the coastline, defined by the 2021 Statistics Canada Census of Population, to the outer boundary of Canada’s exclusive economic zone (EEZ). Coastal areas of salt marsh are also included in the files.Release date: 2024-12-16
- Surveys and statistical programs – Documentation: 37-20-00012024003Description: This technical reference guide is intended for users of the Education and Labour Market Longitudinal Platform (ELMLP). The data for the products associated with this issue are based on the longitudinal Postsecondary Student Information System (PSIS) administrative data files. Statistics Canada has derived a series of annual indicators of public postsecondary students including persistence rates, graduation rates, and average time to graduation by educational qualification, field of study, age group and gender for Canada, the provinces, and the three combined Territories.Release date: 2024-12-11
- Surveys and statistical programs – Documentation: 11-633-X2024004Description: The Longitudinal Immigration Database (IMDB) is a comprehensive source of data that plays a key role in the understanding of the economic behaviour of immigrants. It is the only annual Canadian dataset that allows users to study the characteristics of immigrants to Canada at the time of admission and their economic outcomes and regional (inter-provincial) mobility over a time span of more than 40 years.Release date: 2024-12-09
- Surveys and statistical programs – Documentation: 11-633-X2024005Description: The Analytical Studies and Modelling Branch is the research (ASMB), modelling, training and access hub of Statistics Canada. It focuses on leveraging the agency’s vast data holdings to generate in-depth insights that support evidence-based policy making and to enable others to do so through analytical training and data access. The ASMB, like other program areas in the agency, works to support Statistics Canada’s overall mission of delivering insights through data for a better Canada.Release date: 2024-12-06
- Notices and consultations: 95-635-XDescription: To stay relevant, preparing for a new Census of Agriculture requires a thorough evaluation of data requirements. Before each census, Statistics Canada conducts consultations to solicit input and feedback on the Census of Agriculture's content. This report describes those consultations and the process that was followed to test and determine which topics could be potentially retained for the next census.Release date: 2024-11-27
- Geographic files and documentation: 16-510-X2024004Description: This product contains specifications intended for users of the urban greenness geospatial files. This document provides important technical information for users and links to methodology.Release date: 2024-11-21
- Surveys and statistical programs – Documentation: 98-303-XDescription: The Coverage Technical Report will present the errors included in census data that result from persons who are either missed (not enumerated) or enumerated more than once. The population coverage error is one of the most important types of errors because it affects the accuracy of not only population counts, but also all the census data results that describe the characteristics of the population universe.Release date: 2024-10-23
- Geographic files and documentation: 16-510-X2024001Description: This product contains gridded datasets of annual and 30-year average estimates of water yield. Tracking water yield—an estimate of renewable water supply derived from National Water Data Archive (HYDAT) streamflow data—provides information to help understand water resources available for human use and ecosystem needs. Annual datasets are available for the years 1971 to 2021 and cover southern Canada. Thirty-year averages are available for 1971-2000, 1981-2010, and 1991-2020. They cover the terrestrial and freshwater extent of Canada, except for the Arctic Archipelago.Release date: 2024-09-19
- Geographic files and documentation: 16-510-X2024002Description: This product contains specifications intended for users of the water yield geospatial files. This document provides important technical information for users and links to methodology.Release date: 2024-09-19