We make models understandable and trustworthy — extracting human-interpretable concepts with sparse autoencoders, debiasing representations and generative models, characterizing dataset bias, and improving robustness and domain generalization so models behave reliably beyond their training distribution.
Selected Publications — 5 most representative of 12
A sparse autoencoder (PatchSAE) extracts interpretable visual concepts from CLIP, showing prompt-based adaptation mostly reuses concepts already present in the model