Research
Publications
People
Media
Events
Vacancies
Contact
Article
Same Answer, Different Representations: Hidden Instability in VLMs
F. A. Wani
,
A. Suglia
,
R. Saxena
,
A. P. Gema
,
W. C. Kwan
,
F. Barez
,
Et Al.
The Hitchhiker's Guide to Actionable Interpretability
H. Orgad
,
F. Barez
,
T. Haklay
,
I. Lee
,
M. Mosbach
,
A. Reusch
,
N. Saphra
,
Et Al.
Automated Interpretability-Driven Model Auditing and Control: A Research Agenda
F. Barez
Quantifying the Effect of Test Set Contamination on Generative Evaluations
R. Schaeffer
,
J. Kazdan
,
B. Abbasi
,
K. Z. Liu
,
B. Miranda
,
A. Ahmed
,
F. Barez
,
Et Al.
The Capability Frontier: Benchmarks Miss 82% of Model Performance
B. Fowler
,
R. Smith
,
D. T. Graviet
,
W. Myers
,
J. Greaves
,
N. F. Oozeer
,
A. García
,
Et Al.
When AI Systems Learn During Deployment, Our Safety Evaluations Break
F. Barez
Chain-of-Thought Hijacking
J. Zhao
,
T. Fu
,
R. Schaeffer
,
M. Sharma
,
F. Barez
HACK: Hallucinations Along Certainty and Knowledge Axes
A. Simhi
,
J. Herzig
,
I. Itzhak
,
D. Arad
,
Z. Gekhman
,
R. Reichart
,
F. Barez
,
Et Al.
Val-Bench: Measuring Value Alignment in Language Models
A. Gupta
,
D. O'Shea
,
F. Barez
Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
P. Quirke
,
N. Oozeer
,
C. Bandi
,
A. Abdullah
,
J. Hoelscher-Obermaier
,
F. Barez
,
Et Al.
»
Cite
×