I am a theoretical physicist by training. I completed my PhD at Harvard in 2020, then held an independent postdoctoral fellowship at Princeton's Center for the Physics of Biological Function. My research has spanned quantum condensed matter, computational neuroscience, and the statistical mechanics of learning.

Most recently, I worked as a quantitative researcher in high-frequency trading, designing and running large-scale experiments for forecast model development and deploying models into production. I am currently focused on AI research, with particular interests in alignment, reinforcement learning, and mechanistic interpretability in LLMs and other foundation models. My most recent projects are documented in my blog. My latest work is an analysis of the recoverability of linear truth directions in LLMs and anomalous behavioral effects of steering along the directions returned by different estimators.

Julia Steinberg

Research interests

My work uses tools from statistical physics and machine learning to study how neural networks represent and bind information and how they generalize. In academia this produced a model of associative memory for structured knowledge (talk) and of generalization in neural networks with sparse expansions. My current interests extend these questions to AI alignment and interpretability: how reward optimization succeeds and fails, and the honesty and calibration of foundation models.

Code for current projects:

I welcome correspondence from those working on or interested in these questions.

Site built with the assistance of Claude (Anthropic).