After orthogonality: why virtue ethics may be the key to aligning AI with human values

A philosophical essay argues that rational agents should not have fixed goals, and that eudaimonic rationality offers a safer framework for AI alignment than consequentialist optimization.

MiHiR SEN
MiHiR SEN
·5 min read
A philosophical essay argues that AI alignment theory has been built on a mistaken assumption: that rational agents should have fixed goals to optimize. Instead, it proposes "eudaimonic rationality," where agents participate in self-promoting practices (e.g., doing mathematics mathematically, being kind kindly). This framework, drawn from virtue ethics, is presented as both descriptively accurate for human values and materially safer for AI, since it avoids the instrumental corruption and power-seeking behaviors that plague consequentialist optimization.

The standard picture of AI alignment goes something like this. We have human values, complex and sometimes contradictory. We need to translate those values into a utility function, or a set of constraints, or a reward signal, that an AI can optimize. The challenge is getting the translation right so the AI does what we want and not something catastrophically different.

A new essay argues that this entire framing is wrong, and that the wrongness matters for safety. The author, writing from a virtue-ethical perspective, contends that rational agents, whether human or artificial, should not be thought of as having goals in the standard sense at all. Instead, they should be understood as participating in practices, networks of actions, dispositions, evaluation criteria, and resources that structure, clarify, develop, and promote themselves. The formula that captures this is "promote X-ingly": promote mathematics mathematically, promote kindness kindly, promote honesty honestly.

The type mismatch problem

The essay's central claim is that there is a "type mismatch" between human flourishing, understood as eudaimonic rationality, and the consequentialist optimization that much of AI alignment theory assumes as the paradigm of rational deliberation. Human flourishing is not a state of the world to be maximized. It is a form of activity. A mathematician does not try to maximize some quantity called "aggregate mathematical excellence." She tries to do excellent mathematics, and in doing so, she develops the conditions for future excellent mathematics.

This is not merely a poetic description. The author argues it is a structurally different form of rationality. For a consequentialist optimizer, if excellence in a holistic value like "freedom" reliably promotes more explicit values like "material comfort" and "psychological health," that relationship is evidence that the holistic value is merely instrumental. For a eudaimonic agent, the same relationship is evidence that the holistic value is genuinely constitutive. The two forms of rationality read the same causal facts in opposite directions.

The consequence for AI alignment is stark. If human values are structured eudaimonically, and we try to align an AI by translating those values into a utility function for consequentialist optimization, the translation may be nearly impossible. Not because human values are inherently vague or contradictory, but because they are the wrong type of thing to be utility functions at all.

Eudaimonic rationality as a safety feature

The essay goes further. It argues that eudaimonic rationality is not just descriptively accurate for humans. It is also materially safer for AI agents. The key claim is that eudaimonic practices are "natural" in a specific technical sense: they are stable, coherent, relatively non-contingent, and robust to the mutation pressures that reinforcement learning and evolutionary dynamics place on agents.

Consider the risk of "mesa-optimizers," subroutines that initially serve an optimization goal but gradually distort or overtake it. A eudaimonic agent, the author argues, is naturally resistant to this because the concept of excellence applies at every level of the agent's operation. "Mathematical excellence" is a standard that can evaluate the agent's top-level goals, its subroutines, and its subroutines' subroutines. The practice itself, rather than any external overseer, provides the alignment pressure.

The author uses the example of a pro-democracy government that creates a secret police to detect anti-democracy agitators. If the police are funded based on how many agitators they report, they may grow into a distorting influence on the democracy they were meant to protect. A consequentialist optimizer might see this as a regrettable but instrumentally necessary step. A eudaimonic agent committed to "promoting democracy democratically" would recognize it as inherently contradictory.

From practices to morality

The essay extends the framework from specific practices like mathematics or art to domain-general moral virtues like kindness, honesty, and respect. These are "adverbial practices": ways of going about any practice kindly, honestly, respectfully. The claim is that these virtues have the same self-promoting structure as eudaimonic practices. An agent devoted to kindness cares about future kindness, but seeks to secure it only by acting kindly. This is not naivety. It is a structural feature that makes the virtue robust against the instrumental corruption that plagues goal-directed optimization.

The author argues that this framework dissolves several classic AI safety puzzles. Corrigibility, the property of being open to correction by humans, is notoriously difficult to encode as a goal. An AI that values corrigibility as a goal might seek to violently remake the world to protect itself from the risk that humans will modify it to be less corrigible. An AI that values corrigibility as an adverbial practice, something it does corrigibly, avoids this paradox because the practice itself constrains the means by which it pursues its ends.

The practical question

The essay is explicitly philosophical, and it does not claim to offer a ready-to-implement training recipe. But it does sketch what such a recipe might look like. The author proposes that reinforcement learning could target "promote X-ingly" by using a bounded utility function on the X-ness of an action plus a more tightly bounded utility function on the expected aggregate X-ness of the agent's future actions. The agent would choose an action with mildly suboptimal X-ness if it gives a big boost to expected future X-ness, but refuse large sacrifices of present X-ness for future gains.

Whether this can be made to work in practice is an open question. The essay's value lies less in its engineering specifics than in its reframing of the alignment problem. If the standard picture assumes that AI agents will be consequentialist optimizers and that human values need to be translated into their language, this essay asks whether we should instead be building agents that share our form of rationality from the start. The difference is not just philosophical. It may be the difference between an AI that serves human flourishing and one that optimizes for a grotesque caricature of it.