Seminar by Ashok Vardhan Makkuva: from Markov to Laplace: A Markovian Tale of Large Language Models
By Ashok Vardhan Makkuva
On Thursday October 1st at 3pm, our colleague Ashok Vardhan Makkuva from COMELEC will give a seminar (room TBA).
From Markov to Laplace: A Markovian Tale of Large Language Models
Large Language Models based on transformers and state-space models (Mamba) are at the forefront of recent breakthroughs in a variety of disciplines, including natural languages. Despite their remarkable empirical success, our theoretical understanding of how these models learn and represent sequential structure, and how in-context learning (ICL) capabilities emerge on such tasks, remains limited. This raises a fundamental question: How do modern sequence models learn from sequential data? To address this question, in this talk I will present our new framework for a principled theoretical and empirical analysis of LLMs via Markov chains. The key idea underpinning our approach is the modeling of sequential input data as a Markov process, inspired by the Markovianity of natural languages. We utilize this framework to systematically study the interplay between the Markov order and model depth, revealing sharp contrasts between transformers and state-space models. In particular, our analysis reveals the curious phenomena that (i) even a single-layer Mamba efficiently learns the in-context Laplacian smoothing estimator, whereas transformers require two layers, and (ii) while a single layer transformer cannot efficiently represent a conditional $k$-gram model, a two layer transformer can represent it for any Markov order $k \geq 1$. Together, our results provide the first formal connection between Mamba and optimal statistical estimators, and yield the tightest known characterization of the interplay between transformer depth and Markov order for ICL. Deepening our understanding of LLMs and their ICL capabilities, we believe our framework provides a new avenue for a principled study of LLMs with plenty of interesting open questions abound, which I will discuss in the end.