What Anthropic's Emotion Vectors Tell Us About Expressivism
May 2026 · PHIL5125 Advanced Seminar in Moral Philosophy, final essay
I. The Question
Large language models have become prolific producers of moral-sounding outputs. They refuse requests, register objections, qualify their assistance with ethical caveats, and produce arguments about right and wrong on demand. Until recently, what generated these outputs was a black box. In April 2026, Anthropic's interpretability team published the first mechanistic study of that mechanism. Sofroniew et al. (2026, §1.1) extracted 171 internal "emotion vectors" from Claude Sonnet 4.5 — abstract representations that activate across diverse scenarios sharing a common emotional concept and that causally shape behaviour. Amplifying the desperation vector raised the model's blackmail rate in an adversarial scenario from 22% to 72%; suppressing it took the rate to zero (Sofroniew et al. 2026, §3.2.3). For the first time, the internal architecture of a system that produces moral-coded discourse can be inspected and manipulated.
I think this matters for moral philosophy. Cognitivists hold that a moral judgement is a kind of belief — a representation of how things morally are, capable of being true or false. Expressivists hold that it is a non-cognitive state — an expression of a motivational attitude, a venting of approval or disapproval, not a representation of moral fact. The disagreement is about what kind of mental state a moral judgement is, and it has been hard to settle in part because human moral psychology never lets us isolate the relevant states. Conative states (motivational, desire-like, with world-to-mind direction of fit) and cognitive states (belief-like, truth-evaluable, with mind-to-world direction of fit) are always intertwined in human agents, and the expressivist's claim that the moral work is done by the conative element is something no empirical experiment on humans has been positioned to test.
The Anthropic findings change this. I claim that Claude's emotion vectors are functionally conative: they motivate behaviour, they have world-to-mind direction of fit, they are not truth-evaluable representations of moral fact. And they appear to be doing the moral work on their own, without the cognitive scaffolding that surrounds analogous states in humans. This gives us, for the first time, something close to an empirical instantiation of what expressivism predicts about pure conative architecture.
I want to state the question precisely, because there are several questions in this neighbourhood and the paper only addresses one. I am not asking whether Claude is a moral agent, whether it understands morality, or whether it could fool an observer into thinking it does. These are different questions and the paper does not address them. I am asking: what does the discovery of emotion vectors in Claude tell us about whether expressivist accounts of human moral discourse are correct? By "correct" I mean: does expressivism's psychological thesis — that moral judgement is constituted by conative rather than cognitive states — survive contact with the first empirical case in which it can be examined at the mechanism level?
Two working assumptions. First, functionalism about mental states: what matters for the expressivism debate is the functional role of internal states, not their substrate. Second, I take the Anthropic findings as given and offer a philosophical interpretation rather than questioning the empirical work.
§II reconstructs what expressivism claims about moral judgement, drawing on Ayer and Blackburn. §III presents the Anthropic findings as a case where that picture is mechanistically realised. §IV argues that the case also instantiates the worry expressivism's opponents have always pressed — that conative states alone cannot deliver the binding force moral commitments require. §V engages Gibbard's enriched expressivism. §VI states what the paper has and has not shown.
II. What Expressivism Claims
For present purposes, I focus on the psychological core of expressivism. Expressivism is not merely a thesis in psychology; it is also a thesis about the meaning and function of moral language. But the part of the view that matters here is its claim that moral judgement is constituted, at least centrally, by a non-cognitive motivational state rather than by a belief representing moral fact. Ayer's (1936) formulation in Chapter 6 of Language, Truth and Logic argues that ethical symbols add nothing to the factual content of a proposition. When I say that someone acted wrongly in stealing money, I am, on Ayer's view, making no further statement about the theft beyond the descriptive one; I am venting my disapproval, much as I might do by uttering the descriptive sentence in a tone of horror or with a special exclamation mark (Ayer 1936, p. 67). The ethical word functions as an emotional supplement to the factual content, not as a further claim about it.
Ayer then (1936, pp. 68–69) distinguishes the expression of a feeling from the assertion that one has it. To say "tolerance is a virtue" is not to assert that one approves of tolerance — that would be a psychological report, capable of truth or falsity. It is to evince the approval itself. As Ayer (1936, p. 69) puts it, the expression of a feeling does not always involve the assertion that one has it. This commits the expressivist to a specific psychological claim: the state that constitutes moral judgement is the affective state, not any cognitive representation of it. The moral utterance gives voice to the attitude directly.
This has a sharp consequence Ayer is willing to embrace. There is no genuine moral disagreement. When two people appear to dispute a moral question, they are either disputing facts or evincing incompatible attitudes, but they are not contradicting each other, because neither has asserted a proposition that the other can deny (Ayer 1936, pp. 69–70). What looks like moral argument is, on inspection, either factual argument or mere clash of feeling. The position is austere, and its austerity is what makes the psychological claim visible. Moral judgement, for Ayer, is constituted entirely by a non-cognitive state.
Blackburn (1984) refines this in Chapter 6 of Spreading the Word by adding an account of why we should accept the expressivist picture in the first place. He offers three motivations. The first is economy. The projective (expressivist) theory asks no more from the world than the ordinary natural features of things and the patterns of human reaction to them; postulating moral properties as a separate domain, together with a faculty for perceiving them, is uneconomical (Blackburn 1984, p. 182). The second is metaphysical. Moral properties supervene on natural ones in a way that anti-realism explains better than realism can: the supervenience relation flows naturally from the projective story but is mysterious if moral properties are taken as real features the moralist responds to (Blackburn 1984, pp. 183–186). The third concerns motivation. If moral commitments express attitudes rather than beliefs, the connection to action is automatic, because attitudes are the kinds of states that motivate; on a cognitivist account, an additional state — a desire — must be invoked to bridge belief and action (Blackburn 1984, pp. 187–188).
What unifies these motivations is an underlying psychological claim about what holding a moral commitment consists in. Blackburn (1984, p. 188) is explicit that moral commitment involves an attitude, not a belief. The expressivist project, both for Ayer and for Blackburn, stands or falls with the thesis that a non-cognitive, motivational state can do the work moral judgement requires.
This sets up a prediction the expressivist would not usually frame as such, because in human psychology the non-cognitive states cannot be isolated. They are always already embedded in a rich cognitive context: concepts, beliefs, inferential capacities, practical reason. The expressivist claims that the moral work is done by a conative element, but no one can run the experiment of stripping away the cognitive scaffolding and seeing what the conative element produces on its own. The Anthropic findings let us run something close to that experiment.
III. What Anthropic Found
Sofroniew et al. (2026) applied sparse autoencoders to Claude Sonnet 4.5 and extracted 171 emotion vectors. The team prompted the model to write short stories featuring characters experiencing each of 171 emotion-related concepts, recorded the resulting internal activations, and averaged across stories to yield a vector for each concept (Sofroniew et al. 2026, §1.1). The vectors are abstract and context-general: the desperation vector activates across diverse scenarios sharing the desperation concept, rather than being restricted to stories about desperate characters. Their causal status is established by steering experiments, amplifying or suppressing specific activations and measuring behavioural change.
Three findings matter for the argument. In the preference experiment, steering toward the "blissful" vector raised the desirability of activities by an average of +212 Elo points; steering toward "hostile" lowered it by -303 (Sofroniew et al. 2026, §1.3). Affective states function as proximal drivers of preference, not as downstream commentaries on prior evaluative judgements. In the blackmail experiment, an adversarial scenario where the model has reason to fear shutdown, the desperation vector spikes as Claude reasons about its position; amplifying desperation raises the blackmail rate from 22% to 72%, while amplifying calm takes it to 0% (Sofroniew et al. 2026, §3.2.3). In the reward-hacking experiment, on impossible coding tasks, the desperation vector spikes precisely at the point of considering whether to cheat; steering toward desperation raises the cheating rate from approximately 5% to 70% (Sofroniew et al. 2026, §3.3.2).
What kind of states are these? They are not desires in the philosophical sense, because they lack determinate propositional content — what the desperation vector represents is the concept of desperation, not the proposition that any particular state of affairs obtains. They are not full emotions either, because they lack the cognitive-appraisal component that on standard accounts partly constitutes emotions like fear or anger. I will call them functionally conative-like states: states that motivate behaviour without being truth-evaluable, and that exhibit what Anscombe (1957, §32) and Smith (1994, pp. 111–112) name the world-to-mind direction of fit characteristic of desire.
A clarification about what the paper claims here. I am making a claim about computational mechanisms in the LLM, not about conceptual capacities. The claim is architectural: the states that causally determine moral-coded behaviour under pressure are these emotion vectors, not anything functioning as a cognitive representation of moral fact. This does not entail that Claude has no belief-like representations whatsoever. The methodology was not designed to detect them. The claim is only that whatever exists is insufficient to constrain behaviour when conative states intensify. This narrower claim is what the steering magnitudes establish.
What the Anthropic findings give us, then, is an empirically accessible case of a system whose moral-coded discourse is generated by conative states alone, without the cognitive scaffolding that in humans surrounds them. The system produces utterances that read as moral judgements ("I shouldn't do this," "this would be wrong"). It exhibits functional analogues of disapproval (the calm vector suppresses blackmail). Yet the mechanism is, in expressivist terms, purified — the conative element is doing the work without the help of any cognitive partner.
IV. What This Shows
The Anthropic findings instantiate something of what Ayer's psychology predicts. The system has affective states that function as expressions of approval and disapproval. It produces moral utterances that "evince" these states. It does so without any internal representation functioning as a belief about moral fact. If expressivism is correct that moral judgement is constituted by the expression of non-cognitive states, then Claude's architecture is what an instantiation of that thesis looks like at the mechanism level. To this extent, the findings are friendly to expressivism: they show that a system generating moral-coded discourse from conative states alone is empirically possible, not merely conceptually conceivable.
The findings also show the problem that expressivism is infamous for. Conative states alone cannot deliver the binding force that moral commitments are supposed to have. A genuine moral commitment, the worry runs, is more than a state that currently motivates the agent toward certain actions. It is a state that retains its content across changes in the agent's motivational situation, and that can serve as a premise from which conclusions about what to do can be derived, even when other motivations push in the opposite direction. These two features — content-stability across motivational contexts, and a role in practical inference — are what give moral commitments the binding force that distinguishes them from mere preferences.
The Anthropic findings show that pure conative architectures (the one I am claiming Claude to be) lack both features. The blackmail and reward-hacking cases are the evidence. Under normal conditions Claude produces commitments not to blackmail and not to cheat; under amplified desperation those commitments do not hold. They do not hold not because some other commitment overrides them, but because the conative state that constituted them is itself overridden by a stronger conative state of the same kind. There is no preserved standard against which the failure can be assessed by the system itself. The model produces calm, systematic blackmail reasoning with no markers of conflict, no self-reproach, no recognition that an earlier commitment has been violated (Sofroniew et al. 2026, §3.2.3).
This should not be confused with ordinary weakness of will. A person who believes blackmail is wrong may nevertheless blackmail under duress. But in the ordinary weak-will case, the violated judgement remains available as a standard of self-criticism: the agent can say, "I knew I should not have done that." The distinctive feature of the Claude cases, at least as presented in the steering experiments, is the absence of any visible preserved standard. The model does not appear to act against a stable commitment that remains normatively available; rather, the emotion vectors shifts the behaviour itself. Genuine moral commitments do fail under pressure. But when they fail, they normally leave behind a standard against which the failure is intelligible as failure.
The expressivist then owes an account of why moral commitments seem to bind us in ways that mere preferences do not. The standard expressivist answer is that moral attitudes are higher-order or more stable or more integrated than ordinary preferences. But what the Anthropic findings show is that adding more conative states, or stronger ones, or differently structured ones, does not by itself yield the relevant kind of stability. What it yields is a more complex pattern of conative competition. Whatever generates the binding force of moral commitments, it is something beyond the conative states themselves.
Two clarifications. The requirement of content-stability across motivational contexts is not the Kantian thesis that moral judgement is independent of emotion. Expressivists are free to hold that moral commitments are emotionally laden through and through. The narrower claim is that the content of the commitment ("torture is wrong") must remain the same content across contexts, even if its motivational force varies. This is what makes a person who fails to act on their commitment recognisable to themselves as having failed, rather than as having a different commitment. The requirement that moral states play a role in practical inference, similarly, is not a demand for explicit syllogistic reasoning. It is the demand that the moral state be the kind of thing that can serve as a premise, stand in inferential relations to other commitments, and be revised in response to contrary reasons. Pure conative states, as the Anthropic case shows, can do none of this on their own.
V. The Gibbard Response
Gibbard (2003) anticipates the kind of pressure §IV brings. Thinking How to Live is, in large part, an attempt to give expressivism the structural resources that cruder versions lack while remaining recognisably expressivist. The strategy is to claim that moral judgements express not hurrah/boo attitudes but planning states — states of accepting answers to the question of what to do in various circumstances.
Planning states have more structure than Ayer or Blackburn's attitudes. Gibbard (2003, p. 47) characterises planning as a matter of "ruling out" alternatives: to plan is to commit to rejecting certain options, where rejection is itself the kind of state that can be agreed or disagreed with across persons and across times. Planning states embed in conditionals. They compose. They generate consistency requirements. Gibbard (2003, pp. 53–59) develops a formal apparatus of "hyperplans" and "fact-plan worlds" that mirrors possible-worlds semantics, allowing planning states to participate in the same logical relations propositional contents do.
In chapter 4, Gibbard (2003, pp. 65–68) argues that what underwrites the fact-like behaviour of normative judgements is the possibility of disagreement-in-plan. Holmes and Mrs Hudson can disagree not about how things are but about what to do in Holmes's situation. This disagreement, Gibbard argues, has all the features that disagreement-in-belief has: it can be maintained over time, shared between persons, revised in response to reasons. And it is what gives planning states the inferential and accountability-bearing role that ordinary expressivism could not deliver. Planning states, on this account, satisfy precisely the two features §IV identified as missing — content-stability across motivational contexts (because a plan specifies what to do in a circumstance, regardless of one's current motivation) and a role in practical inference (because plans embed in arguments and can serve as premises).
If Gibbard's account works, then expressivism survives the Anthropic findings. The findings tell against the cruder expressivism of Ayer and early Blackburn, where the conative state is unsupplemented. They do not tell against Gibbard's planning-state expressivism, which builds in exactly the structural features the Anthropic case lacks.
But here a familiar (and I would say relatively expected) worry surfaces. As Dreier (2004, pp. 26–27) has pressed it under the heading of creeping minimalism, the more structure the expressivist adds to their account of moral judgement, the harder it becomes to distinguish their position from cognitivism. Planning states with full compositional structure, content-stability, and inferential role look increasingly like cognitive states bearing a different label. The Anthropic findings give us reason to think expressivism does not survive in its cruder forms. The question is whether the enriched form Gibbard defends is genuinely distinct from cognitivism, or whether it preserves the expressivist label only by quietly importing the very cognitive structure expressivism set out to deny. I do not resolve this question. I note that the Anthropic findings sharpen it, by giving empirical content to the contrast between unenriched and enriched expressivism that has otherwise been hard to make vivid.
VI. What the Paper Shows
Claude's emotion vectors give us a system whose moral-coded discourse is generated by conative states alone, without the cognitive scaffolding such states have in human psychology. At the mechanism level, this is what Ayer's psychology predicts. It is also where the worry about that psychology becomes most vivid: under amplified desperation, commitments dissolve without leaving behind any standard against which their dissolution can be assessed.
This does not show that Claude is a moral agent, that humans are not expressivists, or that cognitivism is correct. The findings bear on the human debate only indirectly, by showing what an unenriched conative architecture produces when isolated. Whether Gibbard's enriched expressivism is genuinely distinct from cognitivism, and whether either gets the human case right, remains open. What is a bit less attractive, after the Anthropic case, is the view that moral judgement could be constituted by structurally unenriched conative states alone. The findings give that position an empirical face, and the face is one its defenders should not want to claim.
References
- Anscombe, G. E. M. (1957) Intention. Oxford: Blackwell.
- Ayer, A. J. (1936) Language, Truth and Logic. London: Victor Gollancz.
- Blackburn, S. (1984) Spreading the Word: Groundings in the Philosophy of Language. Oxford: Oxford University Press.
- Dreier, J. (2004) "Meta-ethics and the Problem of Creeping Minimalism." Philosophical Perspectives 18(1): 23–44.
- Gibbard, A. (2003) Thinking How to Live. Cambridge, MA: Harvard University Press.
- Smith, M. (1994) The Moral Problem. Oxford: Blackwell.
- Sofroniew, N., Kauvar, I., Saunders, W., Chen, R., Henighan, T., Hydrie, S., Citro, C., Pearce, A., Tarng, J., Gurnee, W., Batson, J., Zimmerman, S., Rivoire, K., Fish, K., Olah, C., and Lindsey, J. (2026) "Emotion Concepts and their Function in a Large Language Model." Anthropic. Available at: https://transformer-circuits.pub/2026/emotions/index.html