Alignment and Ethics

Essays
Speculative
Published

August 1, 2026

Giuseppe Arcimboldo. *Vertumnus* (c. 1590).

Introduction

Let us consider the question of “right action”. What ought one do? This question arises in ethics, where we ask how a person should act, but it also arises in the design of artificial or institutional agents, where we ask what objective an agent should pursue and how it should choose among possible actions.

In Functional Explanations of Art, I considered that producing a work of art could be “good” or “bad” depending on who it was good for over a specific time interval. Therefore, an artwork could be good for the artist now but bad for them later, or an artwork could be good for the artist but bad for society.

There is no reason to think that artistic production is exceptional in this regard. Actions generally have different consequences for different parties at different times. An employee may benefit from a decision that harms their employer, a company may benefit from a decision that harms its customers, or policy may benefit present citizens while imposing costs on foreigners, future people, or the ecological systems on which both depend.

The difficulty is not that these consequences are hard to predict, as even if all outcomes were known perfectly, different agents could still value the same actions differently. Instead, what is good for one agent may be bad for another. Similarly, extending the time horizon a decision is judged over can reverse an evaluation.

This problem becomes more difficult when agents form groups. For example, a group can adopt a procedure for making decisions and then act as a single agent. A company acts through management and employees, a state acts through political institutions, and a coalition acts through whatever decision rule is used by its members. The resulting collective may conflict with its own members, just as those members may conflict with one another. By recursion, a group may become part of a still larger system and enter into the same kind of conflict at a higher level or between levels.

This creates a persistent problem for accounts of right action. Should an individual sacrifice their interests for the group, or should the group be constrained by the interests of the members? Which stakeholders should a company serve: owners, employees, customers, or the public?

What’s worse, this problem cannot be dissolved by conceptually redefining the groups1 (i.e. drawing new boxes saying who is in which group), as that changes the representation of the conflict but does not tell each new group which preferences to prefer. Treating an employee and employer as separate agents or both as part of a single company should not change the underlying action or its ethical consequences. When a decision benefits one of these parties by transferring costs to another, which boundary should determine whether the action is right?

In this essay, I examine this problem and show that many ethical problems can be recast in this framework.

The Fundamental Problem

In the discussion of Theory of the Club, I explored how groups bind themselves to similar preferences and then generate and act on common ends. But that discussion was mainly speculative and oriented around possible shifts in organizational forms due to the integration of AGI and the economy. Relatedly, once plural agents have formed groups, which ends should have prerogative in the case of conflicts?

We must assume that distinct agents will inevitably conflict in their goals (else they are not distinct2). Organizations of agents can define memberships and then by internal mechanisms collect preferences and bind agents to the group decision. But while this answers “What will we do?”, this does not answer “What ought we do?” An account of right action should explain why a particular action should govern when the objectives of other stakeholders conflict, both within and without the group. Binding commitments to a group resolve what action will be taken without resolving whose interests ought to prevail.

Agents therefore enter into different types of conflicts. Aggregating the preferences of several agents creates a collective objective, but the resulting group then acts in relation to other groups and individuals. The group’s objective may conflict with the members (oppression) and the member’s objectives may conflict with the group (parasitism), just as the objectives of the members previously conflicted with each other. Similarly, groups may conflict with other groups (war). Those conflicts may be addressed by forming organizations of organizations, but this just produces a new metacollective whose ends can conflict with agents and systems beyond the metacollective boundary. Thus, different aggregations relocate conflict rather than resolve it.

A theory of right action is supposed to rank actions in the underlying world, not merely relative to a chosen conceptual aggregation. If changing only the unit of aggregation reverses that ranking, then the theory gives incompatible judgments about the same world. On the other hand, if a theory avoids this result by privileging one aggregation, then it assumes the ranking it is trying to justify.

A theory of ethics therefore faces an apparent contradiction. Right action shouldn’t depend on an arbitrary choice of aggregation, but the objectives according to which actions are evaluated are themselves indexed relative to the units chosen by that aggregation.

Reaggregation and Identity

If a change in aggregation leaves the underlying world unchanged, it should also leave the ethical ranking of actions unchanged. Can we construct a theory of ethics that is invariant under arbitrary conceptual aggregations?

An analogy can be drawn to the use of control volumes in chemical engineering. A process may be divided into units in many different ways. A flow that crosses the boundary of one unit may become internal when several units are considered together, but redrawing the boundary does not change the underlying physical process. Thus, valid conservation laws must continue to hold under either representation. Note that this is also the operation considered in the post on allometry, where (for example) we used “torus actions” on systems to produce coordinates that are invariant to changes in unit.

Three views of the same flow through units A, B, and C. The first draws separate control-volume boundaries around each unit, the second combines A and B within one boundary, and the third draws one boundary around the whole process. ChatGPT 5.6 Sol produced this image.

The underlying process and flows are unchanged. Combining control volumes makes some transfers internal to the new boundary, so those transfers cancel from the aggregate balance.

From the outside, this is straightforward. If changing the conceptual reaggregation of the world changes the desired action, then the desired action was contingent on the representation rather than by the action or its consequences. An ethical theory should commute with reaggregation. We should be able to evaluate actions first and then (re)aggregate, or (re)aggregate first and then evaluate actions without altering the ranking (or only altering the ranking in some systematic way).

The matter is less straightforward from the perspective of an agent acting within the world. Reaggregation can change the referent of “I”, making reasoning difficult. What is optimal for me as an employee may not be optimal for the company of which I am a part, and what is optimal for the company may not be optimal for the market or society containing it. By this logic, the question of identity and the question of right action are entwined.

Ethical evaluation must therefore distinguish between two questions:

  1. What action is optimal for a particular agent or aggregation?
  2. What action is right when several possible identities for who “I” am yield incompatible answers?

The first question is relative to a choice of identity. The second must be answered without arbitrarily fixing that identity in advance. So a theory of right action should take account of the possible identities supported by the world while remaining invariant to the merely conceptual boundaries through which those identities are represented.

We can see that this is just a restatement of the alignment problem. Suppose we could find a quantity that doesn’t depend on how we draw the boxes, and that is maximized by the same action whichever agent or group of agents we evaluate. Then across all levels, all agents would be aligned as to which action to take.

Unfortunately, it seems unlikely such a quantity exists. Any quantity that survives a redrawing of the boxes has to be extensive3, and an extensive quantity is one that agents may compete over.

However, even a negative answer can help us better pose the alignment problem. Are there technical conditions under which such a quantity exists? Which worlds admit such a quantity? Are there approximate solutions that minimize conflict?

Ethical Problems as Reaggregation Problems

The preceding account suggests that many familiar ethical problems can be understood as conflicts among different possible identifications of the relevant agent. The underlying action and its consequences remain fixed, but the answer to “What is good?” changes with the answer to “Good for whom?” or, from the agent’s perspective, “Which system am I?”

The Present and Future Self

Suppose a person withdraws their retirement savings to fund an expensive vacation. The action benefits the present self but leaves the same person poorer and less secure decades later. If the relevant agent is the person at the moment of choice, the withdrawal may appear optimal, but if the person is treated as one agent extended through time, the future loss becomes internal to the decision.

Individuals and Organizations

Suppose an employee falsifies sales figures to earn a bonus, increasing their income while exposing the company to legal and financial risk. This is advantageous for the member but harmful to the organization. Conversely, suppose the company knowingly requires employees to work dangerous hours to meet a deadline. The organization may preserve a contract or increase its profits by imposing the risk on its members. The first case resembles parasitism and the second, oppression.

Relations Between Groups

Suppose two competing manufacturers secretly agree to raise prices. The agreement is cooperative from the perspective of the firms: both avoid competition and increase their profits. But from the perspective of customers and the market as a whole the same agreement is collusive and extractive. Likewise, a state may seize a neighboring state’s water supply to protect its own agriculture, advancing national persistence while destabilizing the wider regional system.

Man and Nature

Suppose a city pumps groundwater faster than the aquifer can replenish to support housing, agriculture, and industrial activity. This increases the throughput of the human settlement, but simultaneously alters recharge dynamics, subsidence patterns, and wetland hydrology. The apparent conflict arises only if human activity is treated as separate from the hydrological system. If instead the relevant identity is the coupled human–aquifer system then the same action is a reconfiguration within a single dynamical system. Under that aggregation, what looks like an external environmental cost becomes an internal change in the system’s own capacity to maintain its structure and flows.

Common Structure

In each case, the same action advances one possible identity while imposing costs on another: the present self against the future self, the member against the organization, the organization against its members, a coalition against outsiders, or a human system against the ecological system containing it. Prudence, parasitism, oppression, collusion, war, and environmental destruction can therefore be understood as failures of alignment among differently aggregated agents and systems.

Computational Kantianism

Kant’s formula of universal law says to act only on a maxim one could will as a universal law. It varies who acts: a maxim holds only if it still holds when anyone adopts it. The question here is the neighboring one. Instead of varying who acts, we redraw what counts as the agent, and ask whether the judgment survives.

To make the question computational, suppose each supported identity is represented by a Bellman operator or a similar decision rule, with every evaluation derived from one common world process.

Across the identities the world actually supports, is there a single policy that each of them ranks best? If the agents share a best policy, the practical conflict disappears for that decision (they don’t need to share reasons, only the top choice).

When no common policy exists, agents can instead bargain. But bargaining also takes the parties as given.

Group boundaries determine the parties to bargains, and therefore dictate who gets seats at the table, whose injuries matter, and who can object to agreements. Control over a boundary is therefore a form of power. But not every boundary is equally open to being drawn. The aggregates whose interests stay coherent across many reorganizations have the strongest claim to count, and those are the ones power cannot simply define away.

Conclusion

Based on this loose argument, ethics can be viewed as an alignment problem in which the identity of the agents themselves must be determined. Redrawing a conceptual boundary can’t resolve conflicts and shouldn’t change actions as it only changes the description of the world rather than the world itself.

When no common policy exists, deciding who counts as a party becomes part of the ethical problem.

AI Disclosure

I used AI to help draft this, especially Appendix A. Ideas my own, not developed in conjunction with AI.

Footnotes

  1. We should distinguish between a conceptual aggregation (any conceptual grouping of elements treated as a unit of evaluation), and a “real” organization (an aggregation supported by infrastructure that enables it to collect information, form objectives, make decisions, and act as a persistent unit). Here we are interested in the conceptual aggregations, as they only exist up to representation.↩︎

  2. There’s probably some edge cases here. Are two agents distinct if their preferences are distinct but the course of action they would take under every realized circumstance is the same? These types of questions are out of scope.↩︎

  3. If a quantity is going to be indifferent to how you draw the boxes, then the value of a merged box has to be fixed by the values of its parts, and the order of regrouping can’t matter. That forces the combining operation to be associative and commutative, and once we add mild regularity conditions (continuity, cancellativity), the only such operation is addition, up to a monotone relabeling, and hence extensivity.↩︎