5. Data quality

Lineage
Positional accuracy
Attribute accuracy
Logical consistency
Consistency with other products
Completeness

Spatial data quality elements provide information on the fitness-for-use of a spatial database by describing why, when and how the data are created, and how accurate the data are. The quality elements include an overview reporting on the lineage, positional accuracy, attribute accuracy, logical consistency and completeness. This information is provided to users for all spatial data products disseminated for the census.


Top of Page

Lineage

Lineage describes the history of the spatial data, including descriptions of the source material from which the data were derived, and the methods of derivation. It also contains the dates of the source material, and all transformations involved in producing the final digital files.

The National Geographic Database (NGD) is a joint Statistics Canada-Elections Canada initiative to develop and maintain a spatial database which serves the needs of both organizations. The focus of the NGD is the continual improvement of quality and currency of spatial coverage using updates from provinces, territories and local sources. The source files used for the creation of the boundary file reside on Statistics Canada's Spatial Data Infrastructure (SDI) which was derived directly from data stored on the NGD.

For digital boundary file creation, spatial and attribute information were extracted from the SDI using the lowest level of geography, the dissemination block. Primary data manipulation of the product files included preserving the geographic hierarchy of the attributes inherent within a geographic level. The dissemination block file was copied into a File Geo Database to facilitate geo-processing (e.g., projecting, joins, transforming and verification operations). The spatial component of the file was reprojected from Lambert conformal conic into latitude and longitude coordinates (NAD83) using the ArcGIS® ArcCatalog Feature-Project tool.

All of the higher level digital boundary files were created from the dissemination block level. The files were verified for their spatial and attribute content, translated into French and English, and appropriately named according to the file naming convention. The 2011 Census standard geographic area unique identifier, name, type, and the relationships among the various geographic levels are found on the SDI.

To create the cartographic boundary files, a subset of the full hydrography, the coastal file, was created. This subset of coastal hydrographic features was then used to erase portions of forward sortation areas that are covered by coastal waters.

The inland lakes and rivers file was created by selecting hydrographic features from the National Geographic Database's hydrographic reference layer. These reference data were sourced from the National Topographic Data Base (1:50,000 and the 1:250,000 maps) and the Digital Chart of the World. In British Columbia, information was supplemented with data from the National Hydro Network.

The inland lakes and rivers polygon file contains a selection of hydrographic bodies not found in the coastal file. The inland rivers file contains a selection of linear hydrographic features such as rivers and streams.

Final data processing consisted of the conversion from the File Geo Database format, using FME® (Safe Software), into the following GIS file formats: ArcGIS® (.shp), Geography Markup Language (.gml) and MapInfo® (.tab).

Sources

The product was derived from the 2011 Census postal codeOM variable and the National Geographic Database. The postal codeOM is captured for all households from the address information provided or accepted by the respondent on the front page of the census questionnaire. For the 2011 Census, held on May 10, 2011, some census questionnaires contained a pre-printed postal codeOM that the respondents could either accept or correct; however, other census questionnaires did not contain a pre-printed postal codeOM and respondents were asked to provide a postal codeOM by writing it on the questionnaire. These reported postal codesOM were processed through a series of edit operations that identified missing or invalid responses and replaced them with a valid response to produce the 2011 Census postal codeOM variable. At the end of this process, a final postal codeOM was associated with each census household.


Top of Page

Positional accuracy

Positional accuracy refers to the absolute and relative accuracy of the positions of geographic features. Absolute accuracy is the closeness of the coordinate values in a dataset to values accepted as or being true. Relative accuracy is the closeness of the relative positions of features to their respective relative positions accepted as or being true. Descriptions of positional accuracy include the quality of the final file or product after all transformations.

The Spatial Data Infrastructure is not Global Positioning Systems (GPS)-compliant. However, every possible attempt is made to ensure that the 2011 Census standard geographic area boundaries maintained in the Spatial Data Infrastructure respect the limits of the administrative entities that they represent (e.g., forward sortation areas) or on which they are based (e.g., dissemination blocks). The positional accuracy of these limits is dependent upon source materials used by Statistics Canada to identify the location of limits. In addition, due to the importance placed on relative positional accuracy, the positional accuracy of other geographic data (e.g., road network data and hydrographic data) that are stored within the Spatial Data Infrastructure is considered when positioning the limits of the 2011 Census standard geographic areas.


Top of Page

Attribute accuracy

Attribute accuracy refers to the accuracy of the quantitative and qualitative information attached to each feature (e.g., forward sortation area unique identifier).

The attribute data associated with the polygons in the 2011 Census FSA Boundary File are derived from postal codesOM captured from the 2011 Census of Population questionnaires. Edit procedures verify that a reported postal codeOM was valid and consistent with neighbouring postal codesOM. Postal codesOM which failed these checks were imputed, thus ensuring that 100% of the reported postal codesOM were valid postal codesOM according to Canada Post Corporation as of the postal codeOM reference month.

It is important to note that postal codesOM were not verified against Canada Post Corporation's address information, merely that the postal codeOM was considered valid by Canada Post Corporation.


Top of Page

Logical consistency

Logical consistency describes the fidelity of relationships encoded in the data structure of the digital spatial data.

Boundaries found in this product are compatible with those found in other spatial products produced as part of the suite of 2011 Census Geography products. FSA boundaries are derived from the dissemination block level of the 2011 dissemination block boundaries and as such are inherently consistent with those features.

The FSA Boundary File is derived from the 2011 Census responses and not from address-based data from Canada Post Corporation. Whole dissemination blocks are assigned only one FSA in the FSA Boundary File. Furthermore, since whole dissemination blocks are assigned one and only one FSA, the population and dwelling counts derived by aggregating dissemination blocks assigned to an FSA will not match the aggregations based on each household's reported FSA.


Top of Page

Consistency with other products

Topology checks were performed with the road network file and the FSA boundary file to measure the degree of integration amongst these products. The results indicated the degree of integration was within the default tolerance parameters as defined below.

XY resolution: 0.000000001 degrees
XY tolerance: 0.000000008983153 degrees

The 2011 Census Forward Sortation Area Boundary File and the associated hydrographic reference files are not necessarily compatible with files available from other sources.


Top of Page

Completeness

Completeness refers to the degree to which geographic features, their attributes and their relationships are included or omitted in a dataset. It also includes information on selection criteria, definitions used, and other relevant mapping rules.

The product contains boundaries for 1,621 FSAs. In total, 1,638 FSAs were reported by at least one household in the 2011 Census.

The reasons why a reported FSA may not be represented in the 2011 Census FSA Boundary File includes cases where the FSAs did not meet the criteria for minimum number of responses and were eliminated due to this constraint. As well, an FSA may not be the most frequently reported on any dissemination block thus not appearing in the product. Finally, an FSA may not have appeared in the Census Response Database.

It is important to note that in the digital boundary file and cartographic boundary file, a 2011 FSA may be depicted by more than one polygon. In the digital boundary file, there are some 2011 FSAs that have two or more parts. The cartographic boundary file contains additional polygons as a result of removing the coastal water area from the digital boundary file, thus creating several polygons for one 2011 FSA.

Below is a list of the seventeen forward sortation areas which are not included in the boundary file because they failed to meet the minimum number of response criterion and/or were not the dominant FSA in a dissemination block.

E2R
G1A
H0M
H4Y
H5B
K1A
L0V
L5P
M5K
M5L
M5W
M5X
M7A
M7R
M7Y
T1Z
V7X

Date modified: