Hybrid Prompt Optimization for Sparse Autoencoder Feature Visualization in Large Language Models
Despite remarkable advances in Large Language Models (LLMs) and their widespread deployment, our understanding of their internal operations remains limited. Recent work has introduced methods to extract, in an unsupervised fashion, specific directions (features) in the activation space of LLMs that encode interpretable concepts. However, interpreting these features remains difficult, as existing methods depend on identifying examples with strong activations, which in turn requires large datasets and significant computational resources to register all feature activations simultaneously. In this dissertation, we take inspiration from existing work on feature visualization in vision models and from previous prompt optimization techniques to introduce ADAPT, a new method which substantially outperforms existing baselines in this domain.