Large language models remain vulnerable to jailbreaks and other failures in which they comply with harmful requests, yet the internal organization of harmful response generation remains poorly understood. Here, we ask how the ability to comply with harmful requests is organized in the model’s parameters. We perform a mechanistic analysis directly on model parameters, identifying and pruning parameters that contribute strongly to harmful response generation while contributing minimally to general utility. We find that harmful response generation depends on a sparse set of critical parameters: pruning these parameters substantially reduces harmful compliance while causing only limited degradation to general model capabilities. This selective effect suggests that the mechanism supporting harmful response generation is at least partially separable from the one supporting benign capabilities. Moreover, parameters identified using one harm category often reduce harmful responses in other categories, indicating that harmful compliance relies on components shared across harm types. We observe this structure primarily in aligned models, suggesting that alignment reshapes harmful representations internally, even when their safeguards remain brittle. Notably, the capability to generate harmful responses is dissociated from the ability to recognize and reason about harmfulness. Finally, we extend the analysis to emergent misalignment, showing that pruning critical misalignment parameters can reduce misaligned behavior beyond the domain used to identify them. Together, these results reveal a coherent parameter-level structure underlying unsafe behaviors and suggest a path toward more principled safety interventions.
About the Speaker
Hadas is a Research Fellow at the Kempner institute at Harvard University, where she studies the internal mechanics of large AI models to improve their robustness, safety, and reliability. She completed her PhD in the Technion under the supervision of Prof. Yonatan Belinkov. Previously, she worked at Apple and Microsoft.
While years of scientific research on model training and scaling assume that learning is a gradual and continuous process, breakthroughs on specific capabilities have drawn wide attention. Why are breakthroughs so exciting? Because humans don’t naturally think in continuous gradients, but in discrete conceptual categories. If artificial language models naturally learn discrete conceptual categories, perhaps model understanding is within our grasp. I will describe what we know of categorical learning in language models, and how discrete concepts are identifiable through empirical training dynamics and through random variation between training runs. These concepts involve syntax learning, weight mechanisms, and interpretable patterns—all of which can predict model behavior. By leveraging categorical learning, we can ultimately understand a model’s natural conceptual structure and evaluate our understanding through testable predictions.
About the Speaker
Naomi Saphra is a research fellow at the Kempner Institute at Harvard University and incoming Assistant Professor in Boston University’s faculty of Computing & Data Science. She is interested in NLP training dynamics: how models learn to encode linguistic patterns or other structure and how we can encode useful inductive biases into the training process. Recently, she has begun collaborating with natural and social scientists to use interpretability to understand the world around us. She has become particularly interested in fish. Previously, she earned a PhD from the University of Edinburgh on Training Dynamics of Neural Language Models; worked at NYU, Google and Facebook; and attended Johns Hopkins and Carnegie Mellon University. Outside of research, she plays roller derby under the name Gaussian Retribution and performs standup comedy.
What is a language model actually representing when it processes text? LatentQA reframes activation interpretation as a QA task: a decoder LLM is trained to answer open-ended questions about the internal representations of a subject model, enabling flexible, scalable probing of beliefs, intentions, and attributes—without fixed concept vocabularies. Alexander will present the method and its implications for interpretability and safety monitoring. Based on ICLR 2026 work with Lijie Chen and Jacob Steinhardt.
About the Speaker
Alexander Pan is a researcher at Meta working on agentic security and evals. Previously, he led the safety team at xAI and finished his PhD at UC Berkeley, advised by Jacob Steinhardt. He is interested in understanding and mitigating risks from misaligned AI agents.
AI models often learn problematic reasoning processes due to misspecified training objectives. Interpretability helps us detect, and often fix, such reasoning. For example, inspecting Chain-of-Thought reasoning in LLMs is perhaps the single most common approach to understanding how a model got to its answer. This practice has proven effective for identifying model reasoning failures, mistaken background knowledge, and misinterpretation of user instructions. Yet whether Chain-of-Thought is a faithful reflection of a model’s true reasoning remains a subject of debate. On this point, I present work on the CoT faithfulness problem, including evaluations for explanation faithfulness and methods for improving the faithfulness of CoT explanations. Process supervision, and not merely outcome supervision, significantly improves CoT faithfulness, opening up important applications in monitoring model reasoning for safety. From here, I argue that in order to obtain a complete picture of model interpretability, we must also sharpen our understanding of how internal model representations drive external behavior. I show that, by determining how models represent knowledge, we can control what facts are encoded in models and detect when they output claims that they know are untrue or misleading. With more faithful textual reasoning and better interpretability of model representations, we will be able to efficiently identify and fix safety failures in LLMs.
About the Speaker
Peter Hase is a Postdoc at Stanford University and an AI Institute Fellow at Schmidt Sciences. His research focuses on LLM safety and interpretability, with the goal of enabling human understanding, validation, and control of model reasoning. This work has earned multiple spotlight awards at top AI conferences and appeared in publications including Nature magazine and the International AI Safety Report. Previously, he has worked at Anthropic, Google, Meta, and the Allen Institute for AI. He has served as an Area Chair six times, receiving two Outstanding AC awards, and as a Senior Area Chair for ACL and EMNLP. He received his PhD from the University of North Carolina at Chapel Hill, supported by a Google PhD Fellowship.
To receive updates on our events, subscribe to our mailing list by using this link to send an email to sympa@maillist.ox.ac.uk with the subject “subscribe oxford-ai-safety-and-interp” and follow the instructions in the automated response.