I am a senior research scientist at the Allen Institute for Artificial Intelligence in Seattle, where I work on natural language processing and machine learning. My research focuses on core language model development and applications, with a recent focus on formal and probabilistic methods for LLMs and generative agents for scientific discovery.
Previously, I was a researcher at the Institute for Natural Language Processing at the University of Stuttgart in Germany, where I received my PhD in 2018.
See my publications and selected talks below for more details.
NeurIPS: Two papers accepted: Operadic Consistency: A Label-Free Signal for Compositional Reasoning Failures in LLMs and CoTs as Probabilistic Programs: A Programmatic View of Thinking Step-by-Step in Language Models.
EMNLP: From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models was accepted.
TPM @ UAI: CoTs as Tractable Probabilistic Programs was accepted. I also gave a keynote, Tractable Language Model Programming: Themes and Prospects (video).
COLM: Artifact Linker, a benchmark and environment for LLM-driven automated scientific discovery, was accepted.
We released a new preprint on modeling question decomposition with operads, along with a shorter technical report presented at the Combining Theory and Benchmarks workshop at ICML.
We released Probabilistic Programs of Thought, joint work with the UCLA StarAI Lab.
ICLR: Two papers accepted: AstaBench: Rigorous Benchmarking of AI Agents with a Holistic Scientific Research Suite and Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis.
I gave an invited keynote at the International Workshop on Symbolic-Neural Learning in Osaka (slides).
NeurIPS: Language Modeling by Language Models, our work on research agents for autonomous machine learning, was accepted as a spotlight paper.
EMNLP: TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents was accepted.
We released the AstaBench leaderboard and accompanying technical paper for evaluating LLM agents across scientific tasks.
I taught an updated version of our Language Model Programming course at ESSLLI, including new lectures on probabilistic programming for prompting and loss-function decompilation.
I presented Language Modeling by Language Models, a talk on automated scientific discovery, at the NAACL Workshop on AI for Scientific Discovery (preprint).
ICML: Two papers accepted: Understanding the Logic of Direct Preference Alignment through Logic (brief overview) and ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.
For earlier work and a complete list, see my Google Scholar profile.