Near-Optimal Topic Models for Large Scale Text Data
Adam Breuer
- Abstract
- In this paper, we introduce a new class of topic models called Exemplar Topic Models (ETMs) that give results that are provably more accurate, interpretable, and scalable than the current state of the art. Topic models have rapidly become our workhorse method for inferring the key themes that characterize the content of text datasets. Yet despite their massive popularity, all standard topic models (e.g. LDA, STM, etc.) share two key shortcomings: First, unlike other core methods in the political methodology toolkit, standard topic models have no theoretic guarantees. This means they find topics that are arbitrarily mismeasured (even on ideal data) according to their own models’ objective functions; their results are unstable over multiple runs; and they often miss important topics and return spurious ones. Second, standard topic models find topics that lack a rigorous interpretation, so interpreting a topic requires a researcher to manually inspect each topic’s distribution over ~20,000 dictionary words and make a subjective judgement (e.g. claim ‘this topic is about war’).
The ETMs we introduce share the same rich probabilistic model as conventional topic models such as LDA. However, unlike conventional topic models, ETMs find the topics associated with each document in the text dataset that, even in the worst-case, nearly maximize the model’s a-posteriori probability, guaranteeing near-optimal results. Moreover, each topic found by an ETM has a concise and rigorous interpretation related to a single word (e.g. ‘war’). Leveraging recent results from theoretic machine learning, we show that these results are surprisingly always achievable in sublinear computation time in parallel, which means that ETMs can be applied even to massive new political text datasets containing billions of documents. Beyond their theoretic guarantees, ETM’s also consistently outperform standard LDA topic models in terms of standard measures of topic quality on a variety of canonical datasets as well as new original datasets, such as all online political ads in the US election cycle and all posts on Parler.com during the Jan. 6th attack.
- Presented by
- Adam Breuer
- Institution
- Harvard