Concepts and Methods Guide
Appendix D – Glossary of survey terms

Warning View the most recent version.

Archived Content

Information identified as archived is provided for reference, research or recordkeeping purposes. It is not subject to the Government of Canada Web Standards and has not been altered or updated since it was archived. Please "contact us" to request a format other than those available.

Skip to text

Text begins

A

Aboriginal peoples

Aboriginal peoples in Canada include three distinct groups: First Nations (North American Indian), Métis and Inuk (Inuit), each recognized in the Constitution Act. There are many cultural, historical, regional, political and socio-economic differences between these groups, as well as within each of the three groups.

Administrative data

Administrative data are information that is collected by other government agencies and private sector companies for their own purposes, which is then used by Statistics Canada to efficiently accomplish its mandated objectives. Statistics Canada treats all data that can identify a person, a business or an organization with strict confidentiality.

Analytical file

A Statistics Canada microdata set for a given survey, available for use in Research Data Centres (RDCs) across Canada. RDCs provide researchers with access, in a secure setting, to microdata from population and household surveys. The centres are staffed by Statistics Canada employees. They are operated under the provisions of the Statistics Act in accordance with all the confidentiality rules and are accessible only to researchers with approved projects who have been sworn in under the Statistics Actas ‘deemed employees.’

B

Bootstrap method

The bootstrap method is an approach for estimating error in a dataset related to sampling. Sampling introduces error because data are not taken from the entire population, but only a sub-section, called a sample, which is then used to make estimates for the whole population. There are several methods for estimating the level of sampling error. The bootstrap method usually selects a number of subsamples from the main sample and produces estimates for each subsample. The sampling error is estimated as a function of the observed differences between estimates from the different subsamples and estimates from the complete sample.

C

Census metropolitan area (CMA)

A census metropolitan area (CMA) is formed by one or more adjacent municipalities centred on a population centre (known as the core). A CMA must have a total population of at least 100,000 of which 50,000 or more must live in the cored based on adjusted data from the previous Census of Population Program.

Census of population

A census is the collection of information about all units in a population, sometimes also called a 100% sample survey. Under the Statistics Act of 1971, it is a statutory requirement to conduct a nationwide census every five years. The Census of Population provides information needed to plan community services such as schools, day care, police services and fire protection, to forecast consumer demand and to conduct market research studies.

Census subdivision (CSD)

Census subdivision (CSD) is the general term for municipalities (as determined by provincial/territorial legislation) or areas treated a municipal equivalents for statistical purposes (e.g., Indian reserves, Indian settlements and unorganized territories).

Coefficient of variation (CV)

In a sample survey, results from the sample are used to estimate what the findings would be if the whole population were to be measured. In this process of estimation, some level of error is inevitable. The coefficient of variation (CV) is a way of expressing the sampling error associated with an estimate. First a standard error or ‘average’ error of the estimate is calculated. The CV is obtained by dividing the standard error of the estimate by the estimate itself and expressing the resulting fraction as a percentage. The lower the CV, the higher the data quality (see margin of error).

Confidential information

This is a term used within Statistics Canada to describe information that is subject to the secrecy provisions of the Statistics Act. Information is deemed confidential either because it directly identifies a responding unit, for example, by name, or because it could permit specific responding units to be identified, even when the data is stripped of identifiers, due to the information’s detail or its geographical structure or format.

Confidentiality

Confidentiality denotes an implied trust relationship between the person providing the information and the individual or organization collecting it. This relationship is built on the assurance that the information will not be disclosed without the person’s permission. Under the Statistics Act, information that would identify an individual, business or institution cannot be disclosed without their knowledge or consent.

Coverage

Coverage is the extent to which every person or unit intended for inclusion in a survey or census is in fact counted and counted only once. Coverage errors refer to when persons or units of the survey or census are missed (under-coverage) or over-counted (over-coverage). Studies are often conducted by Statistics Canada to provide estimates of under-coverage and over-coverage of a given survey or census or to examine related issues. For example, Statistics Canada has studied and analyzed the extent to which cell-phone use affects coverage for telephone surveys.

D

Data

Observations and measurements collected during a survey, census or other study. Facts or figures from which conclusions can be drawn.

Data quality

A degree or level of confidence that the data and statistical information are “fit for use”. The particular issues of quality or fitness for use that must be addressed by Statistics Canada are relevance, accuracy, timeliness, accessibility, interpretability and coherence.

Dataset, database

An organized and sorted list of facts or information about a set of individuals, households, businesses, or other relevant units. A Statistics Canada dataset is usually generated by a survey or administrative data, stored on a computer, and organized in such a way that it may be accessed easily by a wide variety of statistical application programs.

Derived variable

A new variable constructed by applying logical or mathematical operations to one or more existing variables in order to meet particular data needs. For example, an age variable can be derived from date of birth information. As another example, a derived variable could be obtained called ‘presence of a chronic health condition’ based on whether or not a respondent answered ‘yes’ at least once to a series of questions asking about specific chronic health conditions such as asthma, diabetes, heart disease, etc.

Dissemination

The process of providing statistical products and services to the general public and to specific data users. Statistics Canada disseminates data and analysis in the form of survey results, research reports, technical papers, periodical magazines, census products, and research compendia. Online products date from 1996 to the present. Historical material can be located using the Library Catalogue. Statistics Canada information is also distributed to an approved network of depository libraries. The objective of dissemination activities is to provide relevant information in a timely fashion, in useful formats, and through accessible channels. Activities in place to support the dissemination of products include client consultation services, marketing, promotions, user-training and other client services.

E

Editing

Editing is a process that ensures survey data are accurate, complete and consistent. A set of editing rules or conditions is applied to a dataset. Data which do not meet the conditions are examined and corrected where appropriate.

Errors

In a sample survey, results from the sample are used to estimate what the findings would be if the whole population were to be measured. The accuracy of such an estimate is a measure of how much the estimate differs from the correct or “true” figure. Departures from true figures are known as errors. Errors can arise from many sources, but can be grouped into a few broad categories: coverage errors, non-response errors, response errors, processing errors and sampling errors.

Coverage errors

Coverage errors refer to when persons or units of the survey are missed (under-coverage) or over-counted (over-coverage).

Non-response errors

Non-response errors occur when it proves impossible to obtain a complete questionnaire from a person, household, or organization. Although certain adjustments for missing data can be made during processing, non-response means that some loss of accuracy is inevitable.

Processing errors

Processing errors include mistakes made during data entry, coding, tabulation or other forms of data manipulation.

Response errors

Response errors indicate that a response may not be entirely accurate. The respondent may have misinterpreted the question or may not know the answer, especially if it is given for an absent household member, for example.

Sampling error

Sampling error refers to the fact that the results of the weighted sample differ somewhat from the results that would have been obtained from the total population. The difference is known as sampling error. The actual sampling error is of course unknown, but it is possible to calculate an “average” value, known as the “standard error”.

Estimate, estimation

Using results of the weighted sample to estimate the characteristics of the total population.

F

First Nations, First Nations people

A term that came into common usage in the 1970s to replace the word “Indian,” which many people found offensive. Although the term First Nations is widely used, no legal definition of it exists. Among its uses, the term “First Nations people” refers to the North American Indian people in Canada, both Status and Non-Status. Many people have also adopted the term “First Nation” to replace the word “band” in the name of their community.

Frame

A list, map, or conceptual specification of the units comprising the survey population from which persons can be selected. For example, a telephone or city directory, or a list of members of a particular association or group.

Frequency

The number of times an event or item occurs in a dataset.

G - H - I

Health regions

Health regions are legislated administrative areas defined by provincial ministries of health. These administrative areas represent geographic areas of responsibility for hospital boards or regional health authorities. Health regions, being provincial administrative areas, are subject to change.

Imputation

Imputation involves replacing either missing or invalid data with valid data. This is normally performed using predetermined rules or with the use of data from a ‘statistical neighbour ’–another responding unit who has similar characteristics. Imputation is often combined with data editing.

Indian Act

The Canadian federal legislation, first passed in 1876, that sets out certain federal government obligations, and regulates the management of Indian reserve lands. The act has been amended several times, most recently in 2017.

Indian band

A group of North American Indian people for whom lands have been set apart and money is held by the Crown. Each band has its own governing band council, usually consisting of one or more chiefs, and several councillors. Community members choose the chief and councillors by election, or sometimes through traditional custom. The members of a band generally share common values, traditions and practices rooted in their ancestral heritage. Today, many bands prefer to be known as First Nations.

Information

Data that have been recorded, classified, organized, related or interpreted within a framework so that meaning emerges.

Information product

Organization of results from Statistics Canada activities, including data files, databases, tables, graphs, maps, and text. This organization can be either pre-defined (standard information product) or made in response to special requests (customized information product). Information products can be made available on either print or electronic media.

Interpretability

Interpretability reflects the ease with which the user may understand, properly use and analyze the data or information. The degree of interpretability is largely determined by: the adequacy of definitions on concepts, target populations and variables; terminology underlying the data; and information on any limitations of the data.

Inuit

“Inuit” means “people” in Inuktitut, one of the languages of Inuit people. Most Inuit live in the Northwest Territories, Nunavut, Northern Quebec and Labrador.

Inuit Nunangat

Inuit Nunangat is the homeland of Inuit of Canada. It includes the communities located in the four Inuit regions: Nunatsiavut (Northern coastal Labrador), Nunavik (Northern Quebec), the territory of Nunavut and the Inuvialuit region of the Northwest Territories. These regions collectively encompass the area traditionally occupied by Inuit in Canada.

Inuk

The singular form of the word Inuit (i.e. ‘a person’).

J - K - L

Logistic regression

A form of regression analysis used when the response variable is a binary variable (a variable having two possible values).

M

Margin of error

In a sample survey, results from the sample are used to estimate what the findings would be if the whole population were to be measured. In this process of estimation, some level of error is inevitable. The margin of error, a measure used to build confidence intervals, serves as a rough indicator of the precision of an estimate. For example, pollsters often say that a certain percentage of the population, plus or minus the margin of error (expressed in percentage points), is likely to vote for a certain candidate, 19 times out of 20. To calculate the margin of error, which in this example corresponds to a 95% confidence interval, the pollster would use the equivalent of plus or minus two standard errors of the estimate (see Standard error).

Methodology

A set of research methods and techniques applied to a particular field of study. At Statistics Canada, methodology refers to survey methodology.

Métis

There is no single definition of Métis.  To some respondents, Métis refers to the Métis Nation; to others, it might refer to a person of mixed Aboriginal and European ancestry who self-identifies as Métis.

Microdata

Files of records pertaining to individual responding units.

N - O

National Household Survey (NHS)

This survey took place in 2011 as a replacement for the long census questionnaire, more widely known as Census Form 2B/2D. The NHS was designed to collect social and economic data about the Canadian population. The objective of the NHS was to provide data for small geographic areas and small population groups. For more information please visit the 2011 National Household Survey.

North American Indian

A term that describes all Aboriginal people in Canada who are not Inuit or Métis. North American Indian peoples are one of three groups of people recognized as Aboriginal in the Constitution Act, 1982. This also refers to First Nations people consisting of Status and non-Status Indians.

Observation

Data collected for a given variable about a particular responding unit. Examples include the specific values for a responding unit on characteristics such as age, gender or marital status—the observation might be ‘77’, ‘woman’ and ‘widowed’.

P

Population centre

The term population centre replaces the term urban area (as used in the Census of Population until 2006). A population centre is defined as an area with a population of at least 1,000 and no fewer than 400 persons per square kilometre. Population centres are classified into three groups, depending on the size of their population:

Postcensal survey

A postcensal survey is one where surveyed units are selected based upon their responses to the Census of Population. These surveys are generally conducted shortly after the Census.

Proportion

A proportion refers to how many responses fall into a given response category in relation to the total responses. It is calculated by dividing the frequency of the response category by the total number of responses to the question.

Public use microdata file (PUMF)

Public use microdata files provide access to responding units so that users can conduct their own research or analysis. They involve a non-identifiable data set containing characteristics pertaining to the units of the survey (e.g., individuals, households or businesses). All such datasets have been authorized for release to the public by the Statistics Canada Microdata Release Committee. The dataset contains no confidential information in that individual identifiers have been removed and any data combination or geography which could potentially reveal the identity of a responding unit has been modified.

Q - R

Record

A record is the data for an individual responding unit in a file containing data for all of a survey’s responding units.

Regression

A statistical method which tries to predict the value of a characteristic by studying its relationship with one or more other characteristics. This relationship is expressed through the means of a regression equation.

Research Data Centres (RDCs)

The Research Data Centre Program provides researchers with access, in a secure Statistics Canada governed setting, to microdata from population and household surveys. The RDC program is part of an initiative by Statistics Canada, the Social Sciences and Humanities Research Council (SSHRC) and university consortia to help strengthen Canada’s social research capacity and to support the policy research community. The program is also supported by the Canadian Foundation for Innovation (CFI) and the Canadian Institutes of Health Research (CIHR).

Respondent, responding unit

The respondent is the person providing the information for the surveyed unit, which could be a person, household, business or institution. In the case of the 2017 Aboriginal Peoples Survey, in general, the respondent is the selected person aged 18 and older. For youth between the ages of 15 and 17, the prior approval of the individual‘s parent or guardian is required in order to conduct the interview directly with the youth. Thus, for the age group 15 to 17, the respondent is either the youth or his or her parent or guardian.

Response rate

The proportion of a sample for which a response to a questionnaire is obtained, usually expressed as a percentage. Non-response covers those who refused to participate as well as persons whom the survey was unable to reach.

S

Sample design, Sampling design

A set of specifications that describe the sampling elements of a survey in detail. These elements include population, frame, surveyed units, sample size, sample selection and estimation method. 

Sampling

The process of selecting some part of a population to observe so as to estimate something of interest about the whole population. Examples of different sampling methods include simple random sampling, stratified random sampling, cluster sampling, multiple-phase sampling and multi-stage sampling.

Sampled unit

The unit selected by the sample design and from which measurements are taken for a survey. Examples include persons, households, families or businesses. For APS, the sampling unit is the person.

Sampling fraction

Sample size divided by the population size.

Standard deviation

Standard deviation measures the dispersion of a data set around the mean. It is the most widely-used measure of dispersion. Mathematically, the standard deviation is the square root of the variance.

Standard error

In a sample survey, results from the sample are used to estimate what the findings would be if the whole population were to be measured. Sampling error refers to the fact that the results of the weighted sample differ somewhat from the results that would have been obtained from the total population. The difference is known as sampling error. The actual sampling error is of course unknown, but it is possible to calculate an “average” value, known as the “standard error”.

Statistics Act

An Act regarding statistics of Canada. Includes the definition of Statistics Canada’s mandate: ‘’There shall continue to be a statistics bureau under the Minister, to be known as Statistics Canada, the duties of which are:

Status Indian

"Status Indians" include Registered and Treaty Indians.

Registered Indian

Registered Indians are persons who are registered under the Indian Act of Canada.

Treaty Indian

Treaty Indians are persons who belong to a First Nation or Indian band that signed a treaty with the Crown.

Stratification

A sampling procedure in which the population is divided into homogeneous subgroups or strata and the selection of samples is done independently in each stratum.

Suppress, Suppression

The process by which particular data are prevented from being released based on criteria designed to protect confidentiality. ‘Cell’ suppression refers to procedures used to protect sensitive tabular data from disclosure; a cell being an individual entry in a table. For the APS, data were also suppressed for reasons of data quality (CV larger than 33.3%).

Surveyed unit

The selected unit from which measurements are taken for a sample survey or a census. Examples include persons, households, families or businesses. For APS, the surveyed unit (which is also the sampled unit since the APS is a sample survey) includes persons aged 15 and older.

T - U - V

Target population

The complete group of units to which survey results are to apply. These units may be persons, households, businesses, institutions, etc. This is the population for which information is wanted.

User guides

These guides accompany Statistics Canada survey datasets, such as analytical files and Public use microdata files (PUMF), providing the detailed technical information required to use the data appropriately. The guide typically contains important information to know prior to data analysis: weighting variables to use, procedures related to the estimate of variance, and precautions to take in the dissemination of the data.

Variable

A characteristic that may assume more than one value to which a numerical measure can be assigned (e.g. income, age and weight).

Variance

A measure of dispersion for a given characteristic or variable in a dataset. It indicates how much variability exists for that characteristic. Technically, it is calculated as the average squared deviation from the mean of each observation in the data set for a particular variable.

W - X - Y - Z

Weight

A weight is the average number of units in the population that a unit in the survey represents. Examples of a unit include a person or a household. Weights are applied to responding units in a sample database in order to ensure that, when making inferences from the survey data to population parameters, estimates of characteristics for the total population are obtained.


Date modified: