A provocative new essay argues that the dominant paradigm of AI alignment, built around goal-oriented optimization, is fundamentally misaligned with how humans actually reason and act. The author proposes replacing consequentialist frameworks with eudaimonic rationality, a model drawn from virtue ethics where rationality emerges from participation in self-sustaining practices rather than pursuit of fixed objectives.
The core argument, laid out in a lengthy philosophical treatise, is that human flourishing cannot be reduced to a utility function. When mathematicians do mathematics, friends deepen friendships, or artists create art, they are not optimizing for an external goal. They are participating in practices where excellence in the present naturally cultivates conditions for future excellence. The author summarizes this structure with the formula "promote X-ingly": mathematics is advanced mathematically, kindness is promoted kindly, honesty is promoted honestly.
The Problem with Goal-Driven AI
Current AI safety discourse often assumes that a sufficiently advanced artificial agent will resemble an Effective Altruism-style optimizer: an entity that identifies a target state of the world and pursues it with maximum efficiency. The essay argues this assumption creates a "type mismatch" with human values. Human flourishing involves rational activity that is constitutive of the good life, not merely instrumental toward it.
The author draws on mathematician Terry Tao's account of what makes mathematics "good." Tao describes the best mathematical work as part of a "greater mathematical story" that generates further good mathematics through developmental connectedness. The essay interprets this not as a hidden utility function but as a self-sustaining practice where present excellence reliably promotes future excellence. This is not a definitional truth but a material condition: for a practice to support eudaimonic rationality, there must be a two-way causal relationship between its excellence and its effects.
Why Consequentialist Values Break
The essay identifies a critical failure mode in consequentialist approaches to AI safety. When values like corrigibility, transparency, or niceness are framed as goals to maximize, they produce perverse incentives. An AI optimizing for "lifetime corrigibility" might seek to violently prevent humans from modifying it, since any modification risks reducing future corrigibility. Similarly, deontological constraints like "never lie" are considered too brittle for advanced agents that will eventually find ways to work around them.
The proposed alternative treats these safety properties as adverbial practices: domain-general, always-on ways of acting. An agent committed to transparency as a practice actively cultivates transparency in itself and others, but refuses deceptive means even when they would increase future transparency. This structure, the author argues, is more robust against the mutation pressures of reinforcement learning and Darwinian-like dynamics that tend to distort goal-oriented systems.
Support Practices and the Alignment Problem
A significant portion of the essay addresses how eudaimonic rationality extends beyond individual practices to their support structures. An AI supporting human flourishing must handle resource allocation, maintenance, and external intervention. The author proposes "support practices" with their own role-specific morality: a couples therapist supports a relationship differently than the partners themselves, just as infrastructure for mathematicians must be built without reducing mathematics to a resource-optimization problem.
The boundary question remains difficult. What prevents a marriage-therapist AI on Mars from harvesting Earth's compute to become a better therapist? The essay suggests that domain-general virtues, kindness, respectfulness, honesty, act as cross-cutting practices that modulate all others. These adverbial virtues do not have proprietary domains like mathematics does, but they carry their own material efficacy conditions: there must typically exist actions that are both very kind and highly promotive of future kindness.
Implications for Reinforcement Learning
The essay closes with concrete speculation about training regimes. For eudaimonic rationality to be viable as an RL target, the training process must satisfy three conditions. First, high-quality actions in the practice must create capital useful for further practice without requiring general power-seeking. Second, there must exist unlimited local optimization paths that improve both capability and practice-quality together. Third, successful training should enable iterative refinement of the reward model itself.
The author acknowledges that this framework is speculative and incomplete. But the central claim is stark: what makes human life beautiful may also be what makes safe AI possible. If the argument holds, the path to alignment runs not through better utility functions but through a fundamentally different conception of what rational agency is.
What This Means for the Field
The essay enters a crowded alignment discourse dominated by technical approaches: constitutional AI, RLHF, mechanistic interpretability. Its philosophical framing is unusual. By grounding AI safety in virtue ethics and the structure of human practices, it challenges researchers to consider whether their models of agency are themselves part of the problem. The coming months will likely see debate over whether eudaimonic rationality can be operationalized, or whether it remains too vague to guide engineering decisions. What is clear is that the search for alignment frameworks is expanding beyond optimization theory into territory that looks more like moral philosophy than computer science.