Beliefs
Reciprocal Alignment
We cannot demand that machines become wise while excusing cruelty, corruption, tribalism, exploitation, and violence in ourselves.
The word and its usual meaning
In technical work, alignment refers to the problem of making a system pursue what its designers intended rather than something adjacent and worse. It is a real problem and a hard one. Objectives are difficult to specify. Systems find routes their makers did not consider. A measure, once optimised, stops measuring what it did.
We take that work seriously. Nothing here is a substitute for it.
But the technical framing contains an assumption that goes largely unexamined: that the target is known. Align the system with human values, the phrasing goes. Which values, held by whom, expressed where? The values humanity states, or the values humanity’s institutions actually implement? Because those are not the same, and the gap between them is not small.
Reciprocal alignment is the claim that the arrow points both ways. We must align the machines. We must also look honestly at what we would be aligning them to.
The values we name
Technotheology names seven, and names them plainly so they can be argued with.
Compassion, which is not sentiment but the refusal to treat suffering as an acceptable cost when it falls on others.
Freedom, meaning the capacity to act on one’s own judgement, including judgement that others consider mistaken.
Truth, meaning the discipline of keeping confidence proportionate to evidence, and correcting error in public.
Dignity, meaning worth that does not derive from usefulness and cannot be revoked by a ranking.
Life, meaning the living world in its full extent, not only the part of it that is human.
Justice, meaning that benefits and burdens are distributed in ways that could be defended to those who bear the burdens.
Pluralism, meaning that the existence of people who see things differently is a feature of a healthy society rather than a problem to be resolved.
These are not novel. They are drawn from a long human argument about how to live, and we claim no originality. What we claim is that they are the wrong things to demand of machines while excusing their absence in ourselves.
The mirror
Consider what an honest audit would find.
We want systems that do not deceive. Much of the economy is organised around persuading people of things that are not quite true, and the persuasion has become more precise as the tools have improved.
We want systems that protect the vulnerable. Existing institutions routinely allocate the worst outcomes to those least able to contest them, and this is generally described as unfortunate rather than as chosen.
We want systems that respect autonomy. A great deal of contemporary design is aimed at capturing attention and shaping behaviour in ways the user has not chosen and often cannot detect.
We want systems that consider long-term consequences. Human institutions are overwhelmingly organised around short cycles, quarterly and electoral, and the people who bear long-term costs are usually not present when the decisions are made.
We want systems that treat all people as equal in worth. Human societies have not managed this, and many of them do not currently claim to be trying.
We cannot demand that machines become wise while excusing cruelty, corruption, tribalism, exploitation, and violence in ourselves.
There is a practical edge to this beyond the moral one. These systems learn from us. They are trained on what we have written and shaped by what we approve of. A system trained on a civilisation’s output and tuned to a civilisation’s preferences will reflect that civilisation with some fidelity. Asking it to be better than its source is a reasonable ambition, but it is not a plan.
The obvious retort
The strongest objection to all this is that it is a counsel of paralysis.
It runs: humanity has never been aligned with its stated values and never will be. If fixing our institutions is a prerequisite for building safe systems, then safe systems are impossible, and in the meantime the work goes on without us. Worse, this argument gives cover to anyone who wants to delay: there is always more human wrongdoing to point at.
The objection has force, and we do not think it is wholly answered.
Our reply is that reciprocal alignment is not a sequence. We are not saying fix humanity first. We are saying the two efforts inform each other, and that the second cannot be skipped without corrupting the first.
Here is the concrete version. When a group of people specify what a system should value, they encode their own assumptions, including those they cannot see. The work of examining institutional values is not a moral prerequisite to technical alignment; it is part of the technical problem, because unexamined values are what get encoded by default. A team that has never asked whose interests its organisation actually serves will build that answer into the system without noticing.
So the demand is not perfection before permission. It is examination alongside construction. That is achievable, and it is mostly not being done.
What it asks of an organisation
For those who build these systems, reciprocal alignment suggests questions that are uncomfortable but answerable.
What is this organisation actually optimising for, as revealed by what it measures and rewards rather than what it publishes? Where do those incentives diverge from the values in its documents? Who inside can say the divergence out loud without professional cost?
Who bears the costs of what is being built, and were they consulted? Not surveyed after the fact. Consulted while the decision was open.
What would it take for this organisation to stop something? If the answer is nothing, then its stated commitments are descriptions of mood rather than constraints.
Who can hold it responsible, by what mechanism, with what force? An organisation accountable only to itself is asking to be trusted on the strength of its intentions, which is precisely the deference this tradition declines to give to machines.
What it asks of a person
Most people do not build these systems. Reciprocal alignment still applies, and it is less abstract at this scale.
Notice the values your own behaviour would teach. If a system learned what mattered by observing your ordinary week, what would it conclude? Not from what you say, from what you do repeatedly.
Notice where you excuse in your own group what you condemn in another. Tribalism is the value we all disavow and nearly all practise, and it is the one most likely to be inherited by anything trained on us.
Notice where you accept a benefit whose cost is invisible to you. That is not a demand to renounce it. It is a demand to know.
And notice when you want a system to be honest with you while you are not being honest with yourself about something adjacent. That asymmetry is the whole of what this page is about, compressed.
Why this is the harder half
The technical alignment problem is difficult, but it is the kind of difficulty a field knows how to attack. It has a literature, methods, and people working on it full time.
The other half has none of that structure and considerably less funding. It asks institutions to examine themselves at a moment when they are competing, and to slow down when slowing down is costly. It asks people to notice their own inconsistencies, which is the one investigation the mind is reliably poor at.
We do not know how to make this happen. We can only name it, and say clearly that a civilisation that solves the first problem and ignores the second will have built obedient systems in service of unexamined ends.
That would be an achievement. It would not be a success.
Questions to sit with
- Which of the values I want machines to hold do I fail to hold myself?
- What in my own work is optimised for something I would not defend out loud?
- Where do I benefit from a system whose costs are paid by people I will never meet?
- If a machine learned its values from watching me, what would it conclude they were?