Maximum entropy classification for record linkage
Section 5. Discussion
Below we discuss and
compare two other approaches in the unsupervised setting, including the ways by
which some of their elements can be incorporated into the MEC approach. Other
less practical approaches are discussed in the supplementary material.
5.1 The classical approach
Recall Problems I and II
of the classical approach mentioned in Section 2.
From a practical point of
view, Problem I can be dealt with by any deduplication method of the set of classified
records pairs, where is above a
threshold value for all As “an advance
over previous ad hoc assignment methods”, Jaro (1989) chooses the linked set
which maximises
the sum of subject to the
constraint of one-one link. Since is a monotonic
function of this amounts to
choose which maximises
the expected number of matches in it, denoted by
But is still not
connected to the probabilities of false links and non-links defined by (2.1). As illustrated below,
neither does it directly control the errors of the linked
Consider linking two
files with 100 records each. Suppose Jaro’s assignment method yields on one
occasion, where 80 links have and 20 links
have such that Suppose it
yields 90 links with and 10 links
with on another
occasion, where Clearly, does not
directly control the linkage errors in Moreover, there
is no compelling reason to accept 100 links on both these occasions, simply
because 100 one-one links are possible.
In forming the MEC set
one deals with Problem I directly, based on the concept of maximum entropy that
has relevance in many areas of scientific investigation. The implementation is
simple and fast for large datasets. The estimated error rates FLR (4.5) and MMR
in (4.6) are directly defined for a given MEC set.
Problem II concerns the
parameter estimation. As explained earlier, applying the EM algorithm based on
the objective function (2.2)
proposed by Winkler (1988) and
Jaro (1989) is not a valid
approach of maximum likelihood estimation (MLE). One may easily compare this
WJ-procedure to that given in Section 4.1, where both adopt the same model
(3.3) and the same estimator of via given by (4.4).
It is then clear that the same formula is used for updating at each
iteration, but a different formula is used for
where the numerator is derived from all the pairs in whereas given by (4.2) uses only the pairs in the MEC
set Notice that the two differ only in the
unsupervised setting, but they would become the same in the supervised setting,
where one can use the observed binary instead of the estimated fractional
Thus, one may incorporate
the WJ-procedure as a variation of the unsupervised MEC algorithm, where the
formulae (5.1) and (4.4) are chosen specifically. This is the reason why it can
give reasonable parameter estimates in many situations, despite its
misconception as the MLE. Simulations will be used later to compare empirically
the two formulae (4.2) and (5.1) for
5.2 An approach of MLE
Below we derive another
estimator of by the ML
approach, which can be incorporated into the proposed MEC algorithm, instead of
(4.4). This requires a model of the key variables, which explicates the
assumptions of key-variable errors. Let be the key variable
which takes value Copas and Hilton (1990) envisage a non-informative hit-miss generation
process, where the observed can take the
true value despite the perturbation. Copas and Hilton (1990)
demonstrate that the hit-miss model is plausible in the SL (Supervised
Learning) setting based on labelled datasets.
We adapt the hit-miss
model to the unsupervised setting as follows. First, for any let where if the
associated pair of key variables are subjected to any form of perturbation
that could potentially cause disagreement of the key variable,
and otherwise. Let
where we assume that must be positive for some and
for or Next, for any record in either or let if it has a match in the other file and otherwise. Given with or without perturbation, let We have if is non-informative. A slightly more
relaxed assumption is that is only non-informative in one of the two
files. To be more resilient against its potential failure, one can assume to hold for all the records in the smaller
file, and allow to differ for the records with in the larger file. Suppose Let
be the probability that a record in has a match in One may assume to be independent over giving
where
.
The complete-data log-likelihood
based on is
where
and
, based on an assumption of independent across the entities in
Under separate modelling
of and let be the MLE
based on given which an
EM-algorithm for estimating and follows from (5.2)
by treating as the missing
data. However, the estimation is feasible only if and are not exactly
the same; whereas the MLE of has a large
variance, when and are close to
each other, even if they are not exactly equal.
Meanwhile, the closeness
between and does not affect
the MEC approach, where is obtained
from solving (3.7) given where is indeed most
reliably estimated when Moreover, one
can incorporate a profile EM-algorithm, based on (5.2) given to update in the
unsupervised MEC algorithm of Section 4.1. At the iteration,
where given and estimated from
the smaller file obtain by
ISSN : 1492-0921
Editorial policy
Survey Methodology publishes articles dealing with various aspects of statistical development relevant to a statistical agency, such as design issues in the context of practical constraints, use of different data sources and collection techniques, total survey error, survey evaluation, research in survey methodology, time series analysis, seasonal adjustment, demographic studies, data integration, estimation and data analysis methods, and general survey systems development. The emphasis is placed on the development and evaluation of specific methodologies as applied to data collection or the data themselves. All papers will be refereed. However, the authors retain full responsibility for the contents of their papers and opinions expressed are not necessarily those of the Editorial Board or of Statistics Canada.
Submission of Manuscripts
Survey Methodology is published twice a year in electronic format. Authors are invited to submit their articles in English or French in electronic form, preferably in Word to the Editor, (statcan.smj-rte.statcan@canada.ca, Statistics Canada, 150 Tunney’s Pasture Driveway, Ottawa, Ontario, Canada, K1A 0T6). For formatting instructions, please see the guidelines provided in the journal and on the web site (www.statcan.gc.ca/SurveyMethodology).
Note of appreciation
Canada owes the success of its statistical system to a long-standing partnership between Statistics Canada, the citizens of Canada, its businesses, governments and other institutions. Accurate and timely statistical information could not be produced without their continued co-operation and goodwill.
Standards of service to the public
Statistics Canada is committed to serving its clients in a prompt, reliable and courteous manner. To this end, the Agency has developed standards of service which its employees observe in serving its clients.
Copyright
Published by authority of the Minister responsible for Statistics Canada.
© His Majesty the King in Right of Canada as represented by the Minister of Industry, 2022
Use of this publication is governed by the Statistics Canada Open Licence Agreement.
Catalogue No. 12-001-X
Frequency: Semi-annual
Ottawa