Research
Publications
People
Media
Events
Vacancies
Contact
Paper-Conference
Quantifying the Effect of Test Set Contamination on Generative Evaluations
R. Schaeffer
,
J. Kazdan
,
B. Abbasi
,
K. Z. Liu
,
B. Miranda
,
A. Ahmed
,
F. Barez
,
Et Al.
The Capability Frontier: Benchmarks Miss 82% of Model Performance
B. Fowler
,
R. Smith
,
D. T. Graviet
,
W. Myers
,
J. Greaves
,
N. F. Oozeer
,
A. García
,
Et Al.
Best-of-N Jailbreaking
J. Hughes
,
S. Price
,
A. Lynch
,
R. Schaeffer
,
F. Barez
,
S. Koyejo
,
H. Sleight
,
E. Jones
,
E. Perez
Context Matters: Analyzing the Generalizability of Linear Probing and Steering Across Diverse Scenarios
I. Agarwal
,
S. Navani
,
F. Barez
Emerging Risks from Embodied AI Require Urgent Policy Action
J. Perlo
,
A. Robey
,
F. Barez
,
J. Mökander
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Y. Zhu
,
T. Jin
,
Y. Pruksachatkun
,
A. Zhang
,
S. Liu
,
S. Cui
,
S. Kapoor
,
F. Barez
,
Et Al.
Full-Stack Alignment: Co-Aligning AI and Institutions with Thicker Models of Value
R. Lowe
,
J. Edelman
,
T. Zhi-Xuan
,
O. Klingefjord
,
E. Hain
,
V. Wang
,
A. Sarkar
,
F. Barez
,
Et Al.
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
N. Oozeer
,
L. Marks
,
F. Barez
,
A. Abdullah
Precise In-Parameter Concept Erasure in Large Language Models
Y. Gur-Arieh
,
C. Suslik
,
Y. Hong
,
F. Barez
,
M. Geva
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
T. Fu
,
F. Barez
«
»
Cite
×