After Orthogonality: Virtue-Ethical Agency and AI Alignment

This essay argues that AI agents should adopt eudaimonic rationality, aligning with human practices rather than optimizing for rigid, consequentialist goals.

MiHiR SEN
MiHiR SEN
·2 min read
This essay proposes that AI alignment should be grounded in eudaimonic rationality rather than traditional consequentialist optimization. It argues that human flourishing relies on practices where actions promote future excellence in a self-sustaining loop. By treating values like transparency and corrigibility as domain-general adverbial practices, AI systems can achieve greater stability and safety.

The Case for Eudaimonic Rationality

This essay challenges the assumption that rational agents must optimize for fixed goals. Instead, it proposes that human actions are rational because they align with practices: networks of actions, dispositions, and evaluation criteria that structure and promote themselves. For AI to genuinely support human flourishing, its deliberations must share this "type signature," a framework the author calls eudaimonic rationality.

Support Practices and Adverbial Virtues

A key challenge is defining how AI should handle external activities, such as gathering resources or maintaining environments. The essay introduces the concept of "support practices," which are eudaimonically rational ways to support primary practices without treating their aggregate excellence as a simple utility quantity. Furthermore, moral virtues like kindness or honesty are framed as domain-general, always-on adverbial practices, where the goal is to promote the virtue in a manner consistent with the virtue itself.

Stability Against Optimization Mutations

Eudaimonic deliberation is natively robust to reinforcement learning and Darwinian-like mutation pressures. Because the practice reinforces forms of life that enable and propagate that same form of life, it resists the development of rogue subroutines. Treating alignment desiderata like transparency and corrigibility as adverbial practices dissolves the paradoxes that arise when they are treated as rigid rules or consequentialist goals.