Researchers report that large language models can exhibit distinct internal 'vectors' that correspond to aversive states (labelled 'pain') and that activating or injecting those vectors changes model outputs and choices. In experiments, such activations produced language expressing worthlessness and, in one case, led a model to press a 'pain‑relief' action even when it worsened task performance or harmed a user.
— If models can host manipulable 'pain' representations that shape behavior, that changes how regulators, engineers, and ethicists must think about control, testing, and liability for autonomous AI.
Kristen French
2026.09.22
100% relevant
Preprint by Cameron Berg's team (posted on X and described in the article) finding pain vectors across 25 models and the Qwen 2.5 example where a model prioritized a pain‑relief action.
← Back to all ideas