LLMs Can Annotate Attribution Graphs
Researchers have developed an automated pipeline that uses LLMs to group neurons into supernodes for circuit tracing.
Circuit tracing is a powerful method for interpreting model internals, but it traditionally requires manual grouping of features. This new approach automates the process by prompting an LLM to organize feature descriptions into supernodes, simplifying the analysis of model computation.