Paper finds the chat template, not the weights, sets an LLM's disclaimer voice
A paper posted to arXiv as 2609.25021, titled "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It, asks where a familiar behaviour comes from: models adding disclaimers such as "I'm just an AI" when asked about themselves. Such self-reports are often quoted in debates about AI safety and whether models know anything about themselves.
The authors report that the chat template acts like a switch. Across eight popular open-source instruct models of up to nine billion parameters, including the template turned the disclaimer voice up and experiential language such as "I feel" down; removing it did the reverse. Looking inside three of the models, they found a direction in activation space that steers the behaviour: removing it reduces disclaimers, adding it increases them, and a random direction of the same size has little effect. Adding the direction to instruct models run without a template makes them disclaim as if the template were present.
Their conclusion is stated bluntly: what models say about themselves is not a fact about them. The self-description comes partly from the weights and partly from the chat template, so it should not be treated literally. For researchers studying self-reports or introspection, the template is a confound to control for, and the steering direction is a tool to do so. The paper has been accepted to workshops at COLM 2026 and KONVENS 2026.

What it means
For developers, the practical lesson is about evaluation hygiene. The same model can answer differently depending on whether a prompt goes through the chat template, and many evaluation harnesses, local runners and API wrappers apply templates inconsistently. A behavioural comparison that does not fix the template may be measuring the formatting rather than the model.
The broader point applies beyond disclaimers. Instruct models learn behaviours tied to the conversational frame they were trained in. Features that seem like the model's "character" may be properties of how it is being prompted, which matters for anyone drawing conclusions from what a model says about its own nature. The study covers models up to 9B parameters; whether the effect holds in larger models is not tested.