Statistical methods
Key indicators
Selected geographical area:Canada
-
$5,106.5 million-2.2%
(12-month change) -
$36,023.7 million7.8%
(year-over-year change)
Subject
- Limit subject index to Administrative data
- Limit subject index to Collection and questionnaires
- Limit subject index to Data analysis
- Limit subject index to Disclosure control and data dissemination
- Limit subject index to Editing and imputation
- Limit subject index to Frames and coverage
- Limit subject index to History and context
- Limit subject index to Inference and foundations
- Limit subject index to Quality assurance
- Limit subject index to Response and nonresponse
- Limit subject index to Simulations
- Limit subject index to Statistical techniques
- Limit subject index to Survey design
- Limit subject index to Time series
- Limit subject index to Weighting and estimation
- Limit subject index to Other content related to Statistical methods
Results
All (2,478)
All (2,478) (50 to 60 of 2,478 results)
- Articles and reports: 11-522-X202500100019Description: Accurate and efficient record linkage is crucial for maintaining a comprehensive and current Statistical Business Register (SBR) at Statistics Canada. Linking external business lists to the SBR by name presents computational and methodological challenges, especially as data volumes grow. This paper describes a scalable methodology that employs blocking techniques to constrain the computational search space and integrates multiple similarity measures—from edit distances and n-gram overlaps to embedding-based methods using Sentence-BERT (SBERT)—to identify likely matches. By combining simple character-level comparisons with more advanced semantic embedding methods, the approach can adapt to various naming conventions and complexities. While it does not guarantee superior accuracy in all circumstances, it offers a pragmatic balance between computational feasibility and linkage quality.Release date: 2025-09-08
- Articles and reports: 11-522-X202500100020Description: At Statistics Canada, many data sets are linked with quasi-identifiers such as the first name, last name, or address. In such cases, linkage errors are a potential concern and must be measured. In that regard, previous studies have shown that the evaluation may be based on modeling the number of links from a given record while accounting for all the interactions among the linkage variables and dispensing with clerical reviews, so long as the decision to link two records does not involve other records. In this communication, the methodology is adapted for a class of practical strategies, which violate this constraint by linking the records in consecutive waves, where a given wave links a subset of the records that are not linked in previous waves. In particular, the linkage may be based on a deterministic wave followed by a probabilistic one.Release date: 2025-09-08
- Articles and reports: 11-522-X202500100021Description: Optimal threshold selection is a critical challenge in probabilistic linkage, with significant implications for the accuracy and reliability of linked datasets. This paper analyzes the performance of the neighbour model, a recently proposed error model which models linkage errors by the number of links from each record. Three threshold selection algorithms utilizing the neighbour model were assessed, highlighting the strengths and limitations of each. Their performance was assessed through simulation studies, which demonstrated that methods using the neighbour model achieved lower relative bias compared to two established methods for threshold selection. Additionally, the practical utility was validated through goodness-of-fit tests conducted on four agricultural datasets, showing the potential of the model for use in real-world applications.Release date: 2025-09-08
- 54. T1 Redesign: T1 Partnership Identification Process ArchivedArticles and reports: 11-522-X202500100022Description: In Canada, T1 Tax forms are used to report personal income, whether earned as an employee or through self-employment. Income from self-employment, or "T1 Business Income" is reported by sole proprietorships or partnerships. A T1 partnership involves two or more legal entities jointly filing for a shared business. T1 business data is received as individual filings, meaning partnerships are received separately for each partner. Internal record linkage within the T1 business database is performed to identify partnerships and prevent overcoverage within the final population of T1 businesses. This new T1 partnership identification process takes advantage of newer algorithms, such as DBSCAN numerical clustering fuzzy matching, to identify internal linkages. Graph theory is used to construct the list of partnerships from the row-pairs identified in the linkage process.Release date: 2025-09-08
- Articles and reports: 11-522-X202500100023Description: The latest Canadian Census Health and Environment Cohort (CanCHEC) continues a series of population-based microdata linkages focused on population health research by demographic, social and economic characteristics. The 2021 CanCHEC consists of 95.5% of the 2021 Census long-form sample survey records. The records of survey respondents that could not be linked to the Derived Record Depository and those presumed to be duplicates account for the remaining 4.5%. Linkage-adjusted main and replicate weights allow researchers to estimate and evaluate the variance of summary measures about population health in the presence of missed linked pairs to better understand the experiences of diverse population groups.Release date: 2025-09-08
- 56. The Future of National Statistical Organisations: The Longer-Term Role and Shape of NSOs ArchivedArticles and reports: 11-522-X202500100024Description: This paper explores a vision for the future of National Statistics Offices (NSOs). It analyses the history and role of NSOs before exploring current and future challenges and opportunities for NSOs, before finally outlining a future where NSOs become more agile, open, and collaborative while maintaining their high level of trust in the community, thereby allowing them to fulfil their new role as data stewards in a rapidly evolving data landscape.Release date: 2025-09-08
- 57. Statistical Inference for a Finite Population Mean with Machine Learning-Based Imputation for Missing Survey Data ArchivedArticles and reports: 11-522-X202500100025Description: National statistical offices have increasingly adopted machine learning (ML) for its potential to improve survey estimates. ML techniques offer significant advantages, notably the ability to manage high-dimensional data and to capture complex, nonlinear relationships, thereby enhancing the overall quality of survey statistics. In this article, following the approach of Chernozhukov et al. (2018), we describe a double debiased machine learning framework that enables valid statistical inference when imputed estimators are derived from ML procedures. Simulation results suggest that the proposed framework performs well in a wide range of scenarios.Release date: 2025-09-08
- 58. A Safe and Inclusive Approach to Disseminating Statistical Information about the Non-binary Population in Canada ArchivedArticles and reports: 11-522-X202500100026Description: In 2022, Canada became the first country to release statistical information about its transgender and non-binary populations based on census data. Moreover, following a 2018 government-wide policy direction, Statistics Canada's surveys have been collecting and disseminating information about gender by default rather than sex at birth. Due to the small size of the transgender and non-binary populations, disseminating safe statistical information about them at detailed geographical levels poses a challenge.Release date: 2025-09-08
- Articles and reports: 11-522-X202500100027Description: Several challenges encountered when constructing U.S. administrative record-based (AR-based) population estimates for 2020 are identified. They include locational accuracy, person coverage and its consistency over time, filtering out non-residents and people not alive on the reference date, uncovering missing links across person and address records, and predicting demographic characteristics. Several ways to address these issues are discussed. Regression results illustrate how the challenges and solutions affect the AR-based county population estimates.Release date: 2025-09-08
- Articles and reports: 11-522-X202500100028Description: The United Nations Sustainable Development Goals require detailed, disaggregated data, typically obtained through household surveys. However, surveys alone cannot meet these needs for granular statistics. To address this, National Statistical Institutes adopt small area methods, but these face challenges as auxiliary variables, often derived from surveys, introduce measurement errors into the models. The aim is the application of measurement error correction in classic Fay-Herriot area-level model. The results demonstrate the robustness of the standard approach and ignoring measurement error but show there are specific scenarios where correction for measurement errors is beneficial. The approach is applied to a case study utilizing Indonesian household survey data.Release date: 2025-09-08
- Previous Go to previous page of All results
- 1 Go to page 1 of All results
- 2 Go to page 2 of All results
- 3 Go to page 3 of All results
- 4 Go to page 4 of All results
- 5 Go to page 5 of All results
- 6 (current) Go to page 6 of All results
- 7 Go to page 7 of All results
- ...
- 248 Go to page 248 of All results
- Next Go to next page of All results
Data (10)
Data (10) ((10 results))
- Public use microdata: 89F0002XDescription: The SPSD/M is a static microsimulation model designed to analyse financial interactions between governments and individuals in Canada. It can compute taxes paid to and cash transfers received from government. It is comprised of a database, a series of tax/transfer algorithms and models, analytical software and user documentation.Release date: 2026-02-12
- Profile of a community or region: 46-26-0002Description: The National Address Register (NAR) is a list of commercial and residential addresses in Canada that are extracted from Statistics Canada's Building Register and deemed non-confidential.Release date: 2025-12-19
- Table: 89-26-0006Description: PASSAGES is an open-source dynamic microsimulation model aimed at supporting policy analysis and research relating to Canadian retirement income system outcomes at the individual and family level. The publicly available version includes a synthetic starting database, a model, and documentation. A confidential starting database is also available.Release date: 2025-03-12
- 4. Canadian Statistical Geospatial Explorer Hub ArchivedData Visualization: 71-607-X2020010Description: The Canadian Statistical Geospatial Explorer empowers users to discover geo enabled data holdings of Statistics Canada at various levels of geography including at the neighbourhood level. Users are able to visualize, thematically map, spatially explore and analyze, export and consume data in various formats. Users can also view the data superimposed on satellite imagery, topographic and street layers.Release date: 2024-08-21
- Table: 11-10-0074-01Geography: Census tractFrequency: OccasionalDescription:
The divergence index (D-index) describes the degree that families with different income levels are mixing together in neighbourhoods. It compares neighbourhood (census tract, CT) discrete income distributions to a base distribution, which is the income quintiles of the neighbourhood’s census metropolitan area (CMA).
Release date: 2020-06-22 - 6. Housing Data Viewer ArchivedData Visualization: 71-607-X2019010Description: The Housing Data Viewer is a visualization tool that allows users to explore Statistics Canada data on a map. Users can use the tool to navigate, compare and export data.Release date: 2019-10-30
- Table: 53-500-XDescription:
This report presents the results of a pilot survey conducted by Statistics Canada to measure the fuel consumption of on-road motor vehicles registered in Canada. This study was carried out in connection with the Canadian Vehicle Survey (CVS) which collects information on road activity such as distance traveled, number of passengers and trip purpose.
Release date: 2004-10-21 - Table: 13-220-XDescription: In the 1997 edition, new and revised benchmarks were introduced for 1992 and 1988. The indicators are used to monitor supply, demand and employment for tourism in Canada on a timely basis. The annual tables are derived using the National Income and Expenditure Accounts (NIEA) and various industry and travel surveys. Tables providing actual data and percentage changes, for seasonally adjusted current and constant price estimates are included. In addition, an analytical section provides graphs, and time series of first differences, percentage changes, and seasonal factors for selected indicators. Data are published from 1987 and the publication will be available on the day of release. New data are included in the demand tables for non-tourism commodities produced by non-tourism industries and in the employment tables covering direct tourism employment generated by non-tourism industries. This product was commissioned by the Canadian Tourism Commission to provide annual updates for the Tourism Satellite Account.Release date: 2003-01-08
- 9. Historical Statistics of Canada ArchivedTable: 11-516-XDescription:
The second edition of Historical statistics of Canada was jointly produced by the Social Science Federation of Canada and Statistics Canada in 1983. This volume contains about 1,088 statistical tables on the social, economic and institutional conditions of Canada from the start of Confederation in 1867 to the mid-1970s. The tables are arranged in sections with an introduction explaining the content of each section, the principal sources of data for each table, and general explanatory notes regarding the statistics. In most cases, there is sufficient description of the individual series to enable the reader to use them without consulting the numerous basic sources referenced in the publication.
The electronic version of this historical publication is accessible on the Internet site of Statistics Canada as a free downloadable document: text as HTML pages and all tables as individual spreadsheets in a comma delimited format (CSV) (which allows online viewing or downloading).
Release date: 1999-07-29 - 10. National Population Health Survey Overview ArchivedTable: 82-567-XDescription:
The National Population Health Survey (NPHS) is designed to enhance the understanding of the processes affecting health. The survey collects cross-sectional as well as longitudinal data. In 1994/95 the survey interviewed a panel of 17,276 individuals, then returned to interview them a second time in 1996/97. The response rate for these individuals was 96% in 1996/97. Data collection from the panel will continue for up to two decades. For cross-sectional purposes, data were collected for a total of 81,000 household residents in all provinces (except people on Indian reserves or on Canadian Forces bases) in 1996/97.
This overview illustrates the variety of information available by presenting data on perceived health, chronic conditions, injuries, repetitive strains, depression, smoking, alcohol consumption, physical activity, consultations with medical professionals, use of medications and use of alternative medicine.
Release date: 1998-07-29
Analysis (2,036)
Analysis (2,036) (2,010 to 2,020 of 2,036 results)
- 2,011. Double frame Ontario Pilot Hog Surveys ArchivedArticles and reports: 12-001-X197600200007Description: Three Ontario pilot hog surveys were conducted in 1975 to test a sampling method based on the simultaneous use of two list frames. This paper describes the different aspects of the experience. Particular emphasis is given to the double frame methodology such as discussed by Hartley [11]. Optimal allocation of the sample between frames is considered, with revision for each following survey based on all the accumulated results.Release date: 1976-12-13
- 2,012. Methodology and analysis of the Pilot Auto Exit Survey, 1974 ArchivedArticles and reports: 12-001-X197600100001Description: The 1974 Pilot Auto Exit Survey tested three variants of a handout, mailback questionnaire for U.S. visitors leaving Canada, in a search for a reliable low-cost method of data collection. The sample design was based on a personal interview survey done in Ontario in 1973/74. This design, results of the Pilot, comparison with the Ontario results and some conclusions are presented in this paper.Release date: 1976-06-14
- 2,013. Typical survey data: Estimation and imputation ArchivedArticles and reports: 12-001-X197600100002Description: A special class of missing data problems is discussed, namely that of typical survey data whereby zeros dominate the multivariate response space. Here, techniques which impute means (whether conditional or unconditional) distort rather than improve the quality of the data. A probabilistic model is described which provides reasonable estimates, but also upholds the integrity of the data base. Results are given from a comparative study of the proposed methodology with other estimation/imputation models.Release date: 1976-06-14
- 2,014. Raking ratio estimators ArchivedArticles and reports: 12-001-X197600100003Description: This paper presents large sample results for the bias and variance of raking-ratio estimators for up to four iterations. Estimators of the bias and variance are also presented. An expression for the asymptotic covariance matrix of the maximum likelihood estimators of the cell proportions in a two-way table with known marginals is also given.Release date: 1976-06-14
- 2,015. Methodology of the Labour Force Survey re-interview program ArchivedArticles and reports: 12-001-X197600100004Description: With the recent review of the Labour Force Survey, several periphexal projects have been redesigned. This is the case with the LFS re-interview program which will for the coming years be oriented toward the measurement of response errors. This paper describes the new design of the program and discusses how data will be analysed to achieve the objectives.Release date: 1976-06-14
- Articles and reports: 12-001-X197600100005Description: This paper presents the Behrens-Fisher problem and gives an overview of the major solutions brought forward to this date. The aim of the paper is to use the most appropriate approach to the problem for testing sets of six month Labour Force Survey data against those of a pilot study. This is done since in many cases (such as Methods Test Panel studies) studies are conducted for six consecutive months and comparisons are required on the basis of those sets of six month data. Empirical results are also given by testing Methods Test Panel Phase III data against corresponding Labour Force Survey data.Release date: 1976-06-14
- Articles and reports: 12-001-X197600100006Description: Multi-stage statistical surveys as a means of obtaining socioeconomic characteristics for the population have been in use for many years. Each survey requires an extensive and precise sample design which is governed by the cost structure for obtaining the data and the variance of the characteristic data between units at various stages of sampling. The authors analyzed variance components derived from one month's data of the Canadian Labour Force Survey and examined the variance that would have resulted under different allocation strategies in Table 6 and for different average sizes of units in Table 7. The percentage components of variance, the design effects by stage of sampling and population variances between units of the various stages, as well as measures of homogeneity for households within stages, are derived and shown in Tables 2 to 5. The analysis was carried out for the Canadian Labour Force Survey, but the methodology of component of variance estimation (Gray [4]) and the methods used to analyze the results of a particular survey are readily applied to any multi-stage statistical sample survey, where Horvitz-Thompsen estimators and ratio estimation are applied.Release date: 1976-06-14
- 2,018. Estimation of process average in attribute sampling plans ArchivedArticles and reports: 12-001-X197500254826Description: Exact formulae for bias and mean square error of an estimator of process average in single sampling with rectification for finite lots are obtained. Efficiency of the estimator as compared to an unbiased estimator based on the first sample is obtained for a number of values of lot size, sample size, acceptance number and process average used in sampling plans in quality control of data processing.Release date: 1975-12-15
- 2,019. Method Test Panel Phase II ‒ Data analysis ArchivedArticles and reports: 12-001-X197500254827Description: In the Methods Test Panel Phase II it was required to do analysis of variance on proportions. Since such analysis gives only approximate results, two models were used in order to be able to draw safe conclusions. Analysis of variance was performed with the proportions as variable and also with the arc sine of the square root of the proportions. The two models are outlined in the present paper and empirical comparisons are made using the MTP Phase II data.Release date: 1975-12-15
- 2,020. The methodology of the Canadian Travel Survey, 1971 ArchivedArticles and reports: 12-001-X197500254828Description: The Canadian Travel Survey, 1971 was the largest survey on travel of Canadian residents. This paper describes some important aspects of the methodology. Particular emphasis is given to the development of definitions in relation to the methodology, the sampling technique and interview strategy.Release date: 1975-12-15
- Previous Go to previous page of Analysis results
- 1 Go to page 1 of Analysis results
- ...
- 198 Go to page 198 of Analysis results
- 199 Go to page 199 of Analysis results
- 200 Go to page 200 of Analysis results
- 201 Go to page 201 of Analysis results
- 202 (current) Go to page 202 of Analysis results
- 203 Go to page 203 of Analysis results
- 204 Go to page 204 of Analysis results
- Next Go to next page of Analysis results
Reference (380)
Reference (380) (20 to 30 of 380 results)
- Surveys and statistical programs – Documentation: 84-538-XGeography: CanadaDescription: This electronic publication presents the methodology underlying the production of the life tables for Canada, provinces and territories.Release date: 2023-08-28
- Surveys and statistical programs – Documentation: 32-26-0006Description: This report provides data quality information pertaining to the Agriculture–Population Linkage, such as sources of error, matching process, response rates, imputation rates, sampling, weighting, disclosure control methods and data quality indicators.Release date: 2023-08-25
- Surveys and statistical programs – Documentation: 98-20-00032021011Description: This video explains the key concepts of different levels of aggregation of income data such as household and family income; income concepts derived from key income variables such as adjusted income and equivalence scale; and statistics used for income data such as median and average income, quartiles, quintiles, deciles and percentiles.Release date: 2023-03-29
- Surveys and statistical programs – Documentation: 98-20-00032021012Description: This video builds on concepts introduced in the other videos on income. It explains key low-income concepts - Market Basket Measure (MBM), Low income measure (LIM) and Low-income cut-offs (LICO) and the indicators associated with these concepts such as the low-income gap and the low-income ratio. These concepts are used in analysis of the economic well-being of the population.Release date: 2023-03-29
- Surveys and statistical programs – Documentation: 11-633-X2022009Description: The Longitudinal Immigration Database (IMDB) is a comprehensive source of data that plays a key role in the understanding of the economic behaviour of immigrants. It is the only annual Canadian dataset that allows users to study the characteristics of immigrants to Canada at the time of admission and their economic outcomes and regional (inter-provincial) mobility over a time span of more than 35 years.
This report will discuss the IMDB data sources, concepts and variables, record linkage, data processing, dissemination, data evaluation and quality indicators, comparability with other immigration datasets, and the analyses possible with the IMDB.
Release date: 2022-12-05 - Surveys and statistical programs – Documentation: 32-26-0002Description: This reference guide may be useful to both new and experienced users who wish to familiarize themselves with and find specific information about the Census of Agriculture.
It provides an overview of the Census of Agriculture communications, content determination, collection, processing, data quality evaluation and dissemination activities. It also summarizes the key changes to the census and other useful information.
Release date: 2022-04-14 - Geographic files and documentation: 12-572-XDescription:
The Standard Geographical Classification (SGC) provides a systematic classification structure that categorizes all of the geographic area of Canada. The SGC is the official classification used in the Census of Population and other Statistics Canada surveys.
The classification is organized in two volumes: Volume I, The Classification and Volume II, Reference Maps.
Volume II contains reference maps showing boundaries, names, codes and locations of the geographic areas in the classification. The reference maps show census subdivisions, census divisions, census metropolitan areas, census agglomerations, census metropolitan influenced zones and economic regions. Definitions for these terms are found in Volume I, The Classification. Volume I describes the classification and related standard geographic areas and place names.
The maps in Volume II can be downloaded in PDF format from our website.
Release date: 2022-02-09 - Surveys and statistical programs – Documentation: 11-633-X2021008Description: The Longitudinal Immigration Database (IMDB) is a comprehensive source of data that plays a key role in the understanding of the economic behaviour of immigrants. It is the only annual Canadian dataset that allows users to study the characteristics of immigrants to Canada at the time of admission and their economic outcomes and regional (inter-provincial) mobility over a time span of more than 35 years. The IMDB includes Immigration, Refugees and Citizenship Canada (IRCC) administrative records which contain exhaustive information about immigrants who were admitted to Canada since 1952. It also includes data about non-permanent residents who have been issued temporary resident permits since 1980. This report will discuss the IMDB data sources, concepts and variables, record linkage, data processing, dissemination, data evaluation and quality indicators, comparability with other immigration datasets, and the analyses possible with the IMDB.Release date: 2021-12-06
- Surveys and statistical programs – Documentation: 12-004-XDescription:
Statistics: Power from Data! is a web resource that was created in 2001 to assist secondary students and teachers of Mathematics and Information Studies in getting the most from statistics. Over the past 20 years, this product has become one of Statistics Canada most popular references for students, teachers, and many other members of the general population. This product was last updated in 2021.
Release date: 2021-09-02 - 30. Multi-year Consolidated Plan for Research, Modelling and Data Development, 2021 to 2023 ArchivedSurveys and statistical programs – Documentation: 11-633-X2021005Description:
The Analytical Studies and Modelling Branch (ASMB) is the research arm of Statistics Canada mandated to provide high-quality, relevant and timely information on economic, health and social issues that are important to Canadians. The branch strategically makes use of expert knowledge and a broad range of data sources and modelling techniques to address the information needs of a broad range of government, academic and public sector partners and stakeholders through analysis and research, modeling and predictive analytics, and data development. The branch strives to deliver relevant, high-quality, timely, comprehensive, horizontal and integrated research and to enable the use of its research through capacity building and strategic dissemination to meet the user needs of policy makers, academics and the general public.
This Multi-year Consolidated Plan for Research, Modelling and Data Development outlines the priorities for the branch over the next two years.
Release date: 2021-08-12
- Previous Go to previous page of Reference results
- 1 Go to page 1 of Reference results
- 2 Go to page 2 of Reference results
- 3 (current) Go to page 3 of Reference results
- 4 Go to page 4 of Reference results
- 5 Go to page 5 of Reference results
- 6 Go to page 6 of Reference results
- 7 Go to page 7 of Reference results
- ...
- 38 Go to page 38 of Reference results
- Next Go to next page of Reference results