Maximum entropy classification for record linkage
Section 7. Final remarks

We have developed an approach of maximum entropy classification to record linkage. This provides a unified probabilistic record linkage framework both in the supervised and unsupervised settings, where a coherent classification set of links are chosen explicitly with respect to the associated uncertainty. The theoretical formulation overcomes some persistent flaws of the classical approaches. Furthermore, the proposed MEC algorithm is fully automatic, unlike the classical approach that generally requires clerical review to resolve the undecided cases.

An important issue that is worth further research concerns the estimation of relevant parameters in the model of key-variable errors that cause problems for record linkage. First, as pointed out earlier, treating record linkage as a classification problem allows one to explore many modern machine learning techniques. A key challenge in this respect is the fact that the different record pairs are not distinct “units”, such that any powerful supervised learning technique needs to be adapted to the unsupervised setting, where it is impossible to estimate the relevant parameters based on the true matches and non-matches, including the number of matched entities. Next, the model of the key-variable errors or the comparison scores can be refined. Once these issues are resolved together, further improvements on the parameter estimation can hopefully be made, which will benefit both the classification of the set of links and the assessment of the associated uncertainty.

Another issue that is interesting to explore in practice is the various possible forms of informative key-variable errors, insofar as the model pertaining to the matched entities in one way or another differs to that of the unmatched entities. Suitable variations of the MEC approach may need to be configured in different situations.

Acknowledgements

The authors thank the associate editor and the reviewers for their constructive comments. Dr. Kim is partially supported by NSF grant MMS 1733572.

Supplementary material

In the supplementary material (arXiv:2009.14797), we present the theoretical convergence property of the proposed algorithm and some special cases of MEC sets for record linkage, and discuss two less practical approaches that can be incorporated into the MEC algorithm. An additional simulation study with low levels of the files’ overlap is also presented.

References

Armstrong, J.B., and Mayda, J.E. (1993). Model-based estimation of record linkage error rates. Survey Methodology, 19, 2, 137-147. Paper available at https://www150.statcan.gc.ca/n1/en/pub/12-001-x/1993002/article/14459-eng.pdf.

Berger, A.L., Della Pietra, S.A. and Della Pietra, V.J. (1996). A maximum entropy approach to natural language processing. Computational Linguistics, 22, 39-71.

Binette, O., and Steorts, R.C. (2020). (almost) all of entity resolution. arXiv preprint arXiv:2008.04443.

Christen, P. (2007). A two-step classification approach to unsupervised record linkage. In Proceedings of the Sixth Australasian Conference on Data Mining and Analytics, Citeseer, 70, 111-119.

Christen, P. (2008). Automatic record linkage using seeded nearest neighbour and support vector machine classification. In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 151-159.

Christen, P. (2012). A survey of indexing techniques for scalable record linkage and deduplication. IEEE Transactions on Knowledge and Data Engineering, 24(9), 1537-1555.

Copas, J., and Hilton, F. (1990). Record linkage: Statistical models for matching computer records. Journal of the Royal Statistical Society, Series A, (Statistics in Society), 153(3), 287-312.

Enamorado, T., Fifield, B. and Imai, K. (2019). Using a probabilistic model to assist merging of large-scale administrative records. American Political Science Review, 113(2), 353-371.

Fellegi, I.P., and Sunter, A.B. (1969). A theory for record linkage. Journal of the American Statistical Association, 64(328), 1183-1210.

Gull, S.F., and Daniell, G.J. (1984). Maximum entropy method in image processing. IEE Proceedings 131F, 646-659.

Hand, D., and Christen, P. (2018). A note on using the f-measure for evaluating record linkage algorithms. Statistics and Computing, 28(3), 539-547.

Herzog, T.N., Scheuren, F.J. and Winkler, W.E. (2007). Data Quality and Record Linkage Techniques. Springer Science & Business Media.

Jaro, M.A. (1989). Advances in record-linkage methodology as applied to matching the 1985 census of Tampa, Florida. Journal of the American Statistical Association, 84(406), 414-420.

Larsen, M.D., and Rubin, D.B. (2001). Iterative automated record linkage using mixture models. Journal of the American Statistical Association, 96(453), 32-41.

McLeod, P., Heasman, D. and Forbes, I. (2011). Simulated data for the on the job training. Essnet DI, 70. Available at http://www.crosportal.eu/content/job-training.

Newcombe, H.B., Kennedy, J.M., Axford, S. and James, A.P. (1959). Automatic linkage of vital records. Science, 130(3381), 954-959.

Nguyen, X., Wainwright, M.J. and Jordan, M.I. (2010). Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory, 56(11), 5847-5861.

Nigam, K., Lafferty, J. and McCallum, A. (1999). Using maximum entropy for text classification. In IJCAI-99 Workshop on Machine Learning for Information Filtering, Stockholom, Sweden, 1, 61-67.

Owen, A., Jones, P. and Ralphs, M. (2015). Large-scale linkage for total populations in official statistics. Methodological Developments in Data Linkage, 170-200.

Sadinle, M. (2017). Bayesian estimation of bipartite matchings for record linkage. Journal of the American Statistical Association, 112(518), 600-612.

Sarawagi, S., and Bhamidipaty, A. (2002). Interactive deduplication using active learning. In Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 269-278.

Steorts, R.C. (2015). Entity resolution with empirically motivated priors. Bayesian Analysis, 10, 849-875.

Stringham, T. (2021). Fast Bayesian record linkage with record-specific disagreement parameters. Journal of Business & Economic Statistics, 0(0), 1-14.

Tancredi, A., and Liseo, B. (2011). A hierarchical Bayesian approach to record linkage and population size problems. The Annals of Applied Statistics, 5(2B), 1553-1585.

Winkler, W.E. (1988). Using the EM algorithm for weight computation in the Fellegi-Sunter model of record linkage. In Proceedings of the Section on Survey Research Methods, American Statistical Association, 667-671.

Winkler, W.E. (1993). Improved decision rules in the Fellegi-Sunter model of record linkage. In Proceedings of the Section on Survey Research Methods, American Statistical Association, 274-270.

Winkler, W.E. (1994). Advanced methods for record linkage. In Proceedings of the Section on Survey Research Methods, American Statistical Association, 467-472.

Winkler, W.E., and Thibaudeau, Y. (1991). An Application of the Fellegi-Sunter Model of Record Linkage to the 1990 US Decennial Census. Citeseer.

Xu, H., Li, X., Shen, C., Hui, S.L. and Grannis, S. (2019). Incorporating conditional dependence in latent class models for probabilistic record linkage: Does it matter? Annals of Applied Statistics, 13(3), 1753-1790.

Zhang, G., and Campbell, P. (2012). Data survey: Developing the statistical longitudinal census dataset and identifying its potential uses. Australian Economic Review, 45(1), 125-133.


Date modified: