Aim at agency: institutional design under preference uncertainty, and why alignment needs it

By act65 @ 2026-08-10T16:30 (+1)

The argument

Every answer anyone has given to "what should society maximise" — salvation of souls, the workers' paradise, aggregate happiness, coherent extrapolated volition — was proposed by someone whose own preferences the answer conspicuously served. Each has aged badly.

That isn't a claim that no ethics is correct. It's that betting a durable institution on having found the correct one is a bad bet, and it has been losing for a long time. The main reason is capture: whoever writes the specification writes themselves into it.

So the aim has to be read off the population instead. But we never see what people want — only what they report, through channels that distort, in a population whose wants the institutions already shaped. That's three problems rather than one: how preferences are formed, how they're elicited, and how they're aggregated. Almost all the work goes into the third, and it doesn't survive contact with the other two.

So: design a procedure, not a goal. A procedure still has to be scored against something you can evaluate at design time — call it a social utility function.

The framing I'd most like someone to tell me already exists is that this function is a surrogate loss. What we actually care about is regret against what people turn out to want decades later, which you can't evaluate while designing. The SUF is what you can evaluate instead. So the right questions about it are the ones you'd ask of any surrogate: is it calibrated, how big is the gap, and what happens when you optimise it hard.

Which function? That depends on what people are like. Anyone designing an institution is betting on a model of human nature — not a claim about what any individual wants, but about the range of things a population might turn out to want. I think the tradition has argued almost entirely about where that range is centred, when what decides the design is how wide it is. And width runs one way: the wider it is, the less a design can settle in advance, because every restriction is a bet that the thing foreclosed wasn't the thing worth having. Settle as little as possible, and the target has to be useful across the whole range rather than at any point in it — which means capacity, not outcomes.

**Why alignment needs this. An AI aligned to someone's preferences needs a mechanism to elicit and aggregate them — a voting rule, whether or not anyone calls it one. RLHF and constitutional AI make those choices mostly without saying so, and they inherit all three problems. Elicitation: ask someone to rate two responses and they'll prefer whichever is longer and more confident, so sycophancy is a distortion in the channel rather than a bug in the training. Aggregation: raters disagree, and something has to combine them — that is a voting rule, currently chosen by default rather than designed, and the impossibility results don't care that it was picked implicitly. Formation: a written constitution is a substantive specification of the good, which is the bet this whole argument says not to make — whoever writes it writes themselves into it.

None of that is a claim that alignment researchers are doing something foolish. It's that the field has quietly taken on a sixty-year-old problem in social choice, and mostly isn't treating it as one.

The posts

They're a chain rather than a collection, and the order is the argument: you can't pick a goal, so design a process instead; what the process should look like is decided by how varied people are; that variance forces a process that settles as little as possible; and settling as little as possible is what picks out the candidate objective.

1. Designing behind the veil — the argument above, worked properly. The narrative front door; start here.

2. Why care about institutional design — is the one I'd point a sceptic at first. Throughout history, there have been many different forms of political institutions/organisation, each with different rules. Many of these have led to great suffering, while a small, pluralistic, few have led to (some) prosperity. The point of this post is to explore a range of political institutions through history and simply to look at how they worked and what they achieved... Good and bad.

3. A model of human nature — reads the canon from Hobbes to Fukuyama as a sequence of bets on one object: not what people want, but the *range* of things a population might turn out to want. The central claim is that the tradition has argued almost entirely about where that distribution is **centred**, when the parameter that I think actually decides the design is how **wide** it is.

4. Minimality from uncertainty — the step that connects the two. If the range of what people might want is wide, a design can't settle much in advance — but how much exactly? Modelled as: fix which options stay permissible, then learn within them. The claim is that uncertainty about an option increases the cost of foreclosing it, so you may prune only what you're confident is bad. If it holds, it says why constitutions are lists of prohibitions rather than prescriptions.

⚠️ Caveat, and the reason I'd most like eyes on this one: it's the least mature piece here, was drafted with heavy LLM assistance, and I haven't verified it myself. Treat it as a conjecture. It's also the load-bearing step, which I'm aware is an uncomfortable combination.

5. Agency — what follows from that. If a design must settle as little as possible, the target has to be capacity: the ability to achieve your goals, whatever they turn out to be. The method is the part I'd defend rather than the answer — build the smallest toy worlds that tell candidate scores apart, read one axiom off each, and see what they force. They force more than I expected.