From Data-Model Fit to Model-World Fit: A Framework for Integrating Domain Knowledge into Artificial Intelligence
Abstract
Machine-learning-based artificial intelligence (AI) systems often perform well in controlled settings but fail when deployed in the real world. These failures stem from a fundamental problem: ML/AI models optimize data-model fit by increasing statistical performance on training data, whereas real-world deployment demands a new paradigm in ML/AI: model-world fit that preserves domain semantics and contextual understanding. We propose a Machine Learning Model-World Fit framework that provides both a conceptual structure for understanding how domain knowledge shapes ML/AI systems and a research roadmap for investigating these relationships. The framework identifies how domain knowledge can be integrated into ML/AI development process via external knowledge representations, such as data, value, process models or domain ontologies. Incorporating external knowledge representations can help overcome three types of knowledge deficiencies in AI: know-what, know-how, and know-why. To illustrate the framework, we apply it to a child welfare placement case, a domain where contextual understanding is critical and statistical accuracy is an insufficient evaluation criterion. Our insights lead to several directions for scholars working at the intersection of machine learning, design science, and knowledge representation.

