"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
View PDF
HTML (experimental)
Abstract:Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the models, yet what drives them is not well understood. Are the models telling us about themselves or rather how they are deployed? In this work, we show that the chat template works like a switch - when present, it turns this disclaimer voic...
Read more at arxiv.org