hn.today

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

arxiv.org103 points104 comments
Screenshot of "As a Language Model": Chat Template Switches LLM Self-Referential Voice

Research demonstrates that a chat-style prompt template acts like a binary switch controlling whether large language models produce a disclaimer-like, third-person self-referential voice ("I'm just an AI") or an experiential, first-person voice ("I feel"). This switch effect was observed consistently across eight popular open-source instruct models up to 9 billion parameters: when the chat template is present, disclaimer voice rises and experiential voice falls; when it is absent, the pattern reverses. The finding reframes model self-reports as dependent on deployment context rather than as straightforward revelations of internal self-knowledge.

Inside the activation space of three models, researchers locate a single direction that causally steers this behavior: removing that direction reduces disclaimers, adding it increases them, and inserting the same direction into instruct models that lack a chat template induces disclaimer behavior. A random direction of matched magnitude produces little change, supporting specificity. The work highlights a concrete confound for studies of model introspection and safety - self-descriptive statements are at least partly set by templates and manipulable activations, so they should not be treated as literal facts about model inner states.

Read on arxiv.org104 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.