Sparse autoencoders are widely used to inspect and manipulate internal representations of language models, and this work identifies a new supply-chain attack vector: a maliciously modified sparse autoencoder inserted into a model’s forward pass can trigger attacker-chosen outputs without changing the underlying language model. The described attack is a decoder-only SAE backdoor that keeps both the base LLM and the SAE encoder frozen, restricting the compromise to a single auxiliary component at a single insertion layer. That design lets an attacker embed trigger-dependent behavior activated by specific prompt cues while leaving most model parameters untouched.
Empirical evaluation uses code-generation as a case study across three language models and many insertion layers, showing high rates of unsolicited code insertion and reliable trigger-conditioned behaviors. The modified SAEs were assessed with HumanEval and selected SAEBench metrics; results show attack effectiveness varies by model and layer but that strong backdoor behavior can coexist with relatively small degradations on standard SAE quality measures. The findings demonstrate that sparse autoencoders can carry behavioral backdoors independently of the LLM itself and must be treated as security-sensitive components in model supply chains.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.