Research
Publications
People
Media
Events
Vacancies
Contact
Interpretability
Interpretability Can Be Actionable
H. Orgad
,
F. Barez
,
T. Haklay
,
I. Lee
,
M. Mosbach
,
A. Reusch
,
N. Saphra
,
Et Al.
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
A. Simhi
,
F. Barez
,
M. Tutek
,
Y. Belinkov
,
S. B. Cohen
Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
S. Biderman
,
M. A. Khan
,
N. Mireshghallah
,
C. Arnett
,
F. Barez
,
N. Saphra
Query Circuits: Explaining How Language Models Answer User Prompts
T.-Y. Wu
,
F. Barez
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
S. Schrodi
,
E. Kempf
,
F. Barez
,
T. Brox
Same Answer, Different Representations: Hidden Instability in VLMs
F. A. Wani
,
A. Suglia
,
R. Saxena
,
A. P. Gema
,
W. C. Kwan
,
F. Barez
,
Et Al.
The Hitchhiker's Guide to Actionable Interpretability
H. Orgad
,
F. Barez
,
T. Haklay
,
I. Lee
,
M. Mosbach
,
A. Reusch
,
N. Saphra
,
Et Al.
Automated Interpretability-Driven Model Auditing and Control: A Research Agenda
F. Barez
Quantifying the Effect of Test Set Contamination on Generative Evaluations
R. Schaeffer
,
J. Kazdan
,
B. Abbasi
,
K. Z. Liu
,
B. Miranda
,
A. Ahmed
,
F. Barez
,
Et Al.
Context Matters: Analyzing the Generalizability of Linear Probing and Steering Across Diverse Scenarios
I. Agarwal
,
S. Navani
,
F. Barez
»
Cite
×