Research

Trustworthy & Interpretable AI

We make models understandable and trustworthy — extracting human-interpretable concepts with sparse autoencoders, debiasing representations and generative models, characterizing dataset bias, and improving robustness and domain generalization so models behave reliably beyond their training distribution.

12 publications 822 citations

Selected Publications — 5 most representative of 12