TSG Lab studies the technical and institutional questions raised by increasingly capable AI systems. Our research spans four areas: interpretability, AI safety and alignment, technical AI governance, and the effects of AI on society. This allows us to connect work on the mechanisms inside neural networks with work on evaluation, oversight, and deployment.
Work by members of the Lab has appeared at NeurIPS, ICML, ICLR, ACL, EMNLP, FAccT, and other peer-reviewed venues.
The outputs of an AI system, including its stated reasoning, do not necessarily explain how it reached a decision. We study the representations, features, and circuits inside neural networks to identify the mechanisms responsible for particular behaviours. We also develop automated methods for tracking how these mechanisms change during training and after deployment.
Tests conducted before deployment may not reveal how a system will behave in new settings. We study failure modes including deception, reward manipulation, jailbreaking, high-confidence hallucinations, and the re-emergence of behaviour after fine-tuning or unlearning. We develop evaluations and interventions to test whether safeguards continue to work when models or their environments change.
Developers, auditors, and regulators need reliable evidence about the capabilities and risks of advanced AI systems. We develop evaluation methods, auditing protocols, safety cases, and risk and reliability benchmarks. We also study how technical evidence can support regulatory oversight, verify safety claims, and clarify responsibility when systems fail.
AI systems can change who holds information and decision-making power. We study how their deployment affects human agency, institutional accountability, and the concentration of power, including the risk of AI-enabled authoritarianism. We are interested in technical and policy approaches that keep consequential decisions subject to effective human oversight.