Using Representation Learning and Website Text to Identify Competitor Networks
Abstract
This paper introduces a new approach to identify competitors using company websites to map competitive relationships among public and private firms. We apply representation learning techniques to create embeddings of companies based on website content, emphasizing information about industry-specific products and services. By placing all firms, public and private, within the same embedding space, we construct a peer competitor network that supports analysis of company interactions and market evolution. We label this new approach and set of network competitors as the website text-based network industry classification (WTNIC). We evaluate the quality of the network using multiple ground truth datasets and benchmarks and show that it offers a robust alternative for understanding competition at scale. Despite using noisy website data, WTNIC matches or outperforms prior approaches that rely on curated data. We demonstrate the value of WTNIC through a case study examining how competition from China affects the public and private peers of a focal firm, highlighting the importance of placing both public and private companies within the same embedding space.
Funding: This work was supported by the National Science Foundation [Grants 1561068 and 1937153], the Tuck School of Business at Dartmouth College, the USC Marshall Institute for Outlier Research in Business (iORB), and the Smith School of Business at the University of Maryland.
Supplemental Material: The data files are available at https://doi.org/10.1287/mnsc.2024.06516.

