Integrating Data Cleaning and Robustification: A Study on Linear Regression
Abstract
In this study, we propose a general framework that integrates both optimistic and pessimistic optimization approaches in solving the regression problem to address outlier cleaning and robustification in a unified fashion. Although data cleaning aims to down-weight the outliers, robustification renders the regression models to heavily rely on extreme data. The main objective of this framework is to construct a new optimization scheme capable of withstanding the influence of outliers without harming the robustness level, by combining these two rather contrasting concepts and operations. In addition to showing its generalization to a few well-known regression models, a set of structural properties of our framework is derived to ensure its statistical significance and to understand its computational demand. Then, we develop solution methods, including mixed integer formulations, alternating direction method of multipliers algorithms, and computation enhancement techniques, that can be applied to handle data sets of different scales. Numerical results on both synthetic and benchmark data sets from the University of California, Irvine (UCI) Machine Learning Repository verify the superiority of our new framework and demonstrate the unified strength to handle complex data sets.
History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning.
Funding: X. Qian received financial support from the U.S. National Science Foundation (NSF) [Grants SHF-2215573 and IIS-2212419].
Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information (https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2024.0884) as well as from the IJOC GitHub software repository (https://github.com/INFORMSJoC/2024.0884). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/.

