Data quality, concepts and methodology: Analytical techniques
Archived Content
Information identified as archived is provided for reference, research or recordkeeping purposes. It is not subject to the Government of Canada Web Standards and has not been altered or updated since it was archived. Please "contact us" to request a format other than those available.
Subjects
Cancer incidence data are from the October 2011 version of the Canadian Cancer Registry (CCR), a dynamic, person-oriented, population-based database maintained by Statistics Canada. The CCR contains information on cases diagnosed from 1992 onward, compiled from reports from every provincial/territorial cancer registry. A file containing records of invasive cancer cases and in situ bladder cancer cases (the latter are reported for each province/territory except Ontario) was created using the multiple primary coding rules of the International Agency for Research on Cancer. 1 Cancer cases were classified based on the International Classification of Diseases for Oncology Third Edition 2 and grouped using Surveillance, Epidemiology, and End Results (SEER) Program groups with mesothelioma and Kaposi's sarcoma as separate groups. 3
Analyses were based on all primary cancers—an approach that is becoming standard practice. 4 , 5 The effect of including multiple cancers in survival analyses has been studied internationally 4 , 5 and in Canada. 6 Data from the province of Quebec were excluded from the analysis primarily because of issues in correctly ascertaining the vital status of cases. Records were also excluded if: age at diagnosis was younger than 15 or older than 99; diagnosis was established through autopsy only or death certificate only (DCO); or the year of birth or death was unknown. The majority of exclusions were autopsy only or DCO cases; these were left out because the date of diagnosis, and hence survival time, was unknown. The "true" survival of cases registered by DCO is generally poorer than that of those in the registry population. 7 The necessity of excluding DCO cases may have led to increases in survival estimates, particularly in provinces with proportionately more DCO cases. However, the magnitude of such increases is generally minor. 7
Mortality follow-up through December 31, 2008 was determined by record linkage to the Canadian Vital Statistics Death Database (excluding deaths registered in the province of Quebec), and from information reported by provincial/territorial cancer registries. For deaths reported by a provincial registry but not confirmed by record linkage, the date of death was assumed to be that submitted by the reporting registry. Survival time was calculated as the difference in days between the date of diagnosis and the date of last observation (date of death or December 31, 2008, whichever was earliest) to a maximum of five years. For a small percentage of subjects with missing information on day/month of diagnosis and/or day/month of death, the survival time was estimated. 8
Analysis
Survival analyses were based on a publicly available algorithm, 9 to which minor adaptations were made. The algorithm uses a life table (actuarial) approach in which survival estimates are calculated at discrete points in the follow-up, generally by taking the product of interval-specific (conditional) estimates over sub-intervals of the follow-up. Observation time for each individual is split into multiple observations, one for each sub-interval of follow-up time. Observations are collapsed over calendar year(s) at time of diagnosis. Three month sub-intervals were used for the first year of follow-up and then 6 month sub-intervals for the remaining 4 years for a total of 12 intervals. More intervals were used in the first year of follow-up because the actuarial method assumes an approximately even distribution of deaths within each interval and mortality is often highest during the first year. With the exception of cases previously excluded because they were diagnosed through autopsy only or death certificate only, persons with the same date of diagnosis and death were assigned one day of survival because the program automatically excludes cases with zero days of survival.
Expected survival proportions were derived, from sex-specific complete provincial life tables produced by Statistics Canada, using the Ederer II approach. 10 With this approach, expected survival proportions are estimated for each interval, based on only those cases alive at the start of the interval. Data from the 1990/1992 life tables 11 were used for case follow-up in 1992 and 1993, data from 1995/1997 life tables 12 were used for follow-up from 1994 to 1998, and data from the 2000/2002 life tables 13 were used for follow-up from 1999 to 2008. As complete life tables were not available for Prince Edward Island nor for the territories, expected survival proportions for these areas were derived from abridged life tables for Canada, Prince Edward Island, and the territories, using a method suggested by Dickman et al. 14 Where this was not possible (i.e., territories 1990-92), Canadian complete life table values were used. The aforementioned method of Dickman et al. was also used to extend, by single year of age, the 1990-1992 set of provincial life tables from 85 to 99 years.
Age-, sex-, and province-specific five-year relative survival ratios were estimated for each selected cancer as the ratio of the observed survival of the group of individuals diagnosed with cancer to the expected survival for the corresponding general population of the same age, sex, province of residence, and time period. Survival estimates for the territories were not presented due to the small number of cases for analysis. Cases from these areas were, however, used in the calculation of national estimates. As an indication of the level of statistical uncertainty in the survival estimates, confidence intervals formed from standard errors estimated using Greenwood's method 15 are provided. To avoid implausible lower limits less than zero and/or upper limits greater than one for observed survival estimates, asymmetric confidence intervals based on the log (-log) transformation were constructed. Relative survival ratio confidence limits were derived by dividing the observed survival limits by the corresponding expected survival proportion.
Age-standardized estimates were calculated using the direct method. Age-specific estimates for a given cancer were weighted to the age distribution of persons diagnosed with that cancer from 2001 to 2005 (based on the July 2010 tabulation master file). The age categories used in the weighting depended on the cancer under study and were the same as those that were used in the presentation of age-specific survival estimates for Canada. Confidence intervals for age-standardized relative survival ratios were formed by multiplying the corresponding age-standardized observed upper and lower limits by the ratio of the age-standardized relative survival point estimate to the age-standardized observed survival point estimate. While the choice of a standard population is ultimately an arbitrary one, the chosen population has the advantage that it leads to age-standardized survival estimates that are not widely different from the corresponding non-standardized estimates. 16
- Date modified: