ResearchA visual guide to the paper

Absolute Alignment
Allocation

The case for dedicating every redirectable resource to AI alignment, and making humanity’s survival the purpose of that effort.

Read the paper Download PDF

Two worlds.
The same starting point.

Give two worlds the same models, knowledge, resources, and time. In one, alignment competes with other objectives. In the other, every redirectable resource serves alignment and the work that supports it.

That second commitment is Absolute Alignment Allocation. The comparison begins with a decision about what the available capacity is for.

Divided allocation25%
AAA100%

Share committed to alignment at the start

Adjust the alignment share in the figure. The total resources stay fixed; their purpose changes.

What is held equal?

The comparison fixes the starting models, knowledge, available resources, training regime, and calendar interval. Compute, agent operating time, and human effort can each be examined in their own units.

The figure uses a normalized budget of 100 units over the full interval. It tracks capacity committed to the mission; research results and survival effects enter the argument below.

Paper §1 · The paired worlds ↗

A later commitment cannot reclaim earlier time.

Every interval of divided allocation leaves less effort available to alignment. The gap accumulates even when the two worlds have identical resources.

Now suppose the divided world commits fully partway through. Its alignment effort begins accumulating at the same rate as AAA’s. The earlier gap remains.

The figure lets you move that commitment earlier or later. Under the same available capacity, starting sooner means more effort can serve the mission within the time available.

How the time comparison works

Let a be the initial alignment share, and let d be the fraction of the interval before full dedication begins. With constant total capacity, the share of the full-period budget forgone is (1 − a)d.

Keeping the divided allocation throughout sets d = 1. Starting with AAA sets d = 0. The horizontal axis is a common comparison period, with no assumed date of catastrophe.

Paper §4 · Resource inputs and consequences ↗

Put the effort where survival needs it.

Additional capacity can investigate a failure, test a safeguard, or check whether a promising finding holds up. Its value can lie in completing work that protection depends on.

  1. Investigate

    Find dangerous behavior and develop protective methods.

  2. Test independently

    Expose failures that the original research missed.

  3. Put findings to work

    Make safeguards and deployment decisions depend on the evidence.

AAA dedicates the budget to this purpose. It allows methods to change, unproductive work to stop, and resources to move toward the strongest remaining alignment opportunities.

Supporting work belongs inside the mission

Replication, secure infrastructure, maintenance, methodology, researcher support, and the reliability of automated researchers can all sustain effective alignment work. Exploratory research qualifies through a defensible contribution to the mission.

The allocation rule applies across owners and sectors. Its scope is all resources that can feasibly be redirected, including the dependencies needed to use them effectively.

Paper §2.2 · Eligible work ↗

The decisive difference is the survival difference.

The case for AAA rests on what the added alignment effort makes possible: preventing otherwise fatal failures and improving the decisions that determine humanity’s future.

Compare the complete consequences of the two policies. Cases where allocation changes the outcome determine the net survival advantage.

Which outcomes change the comparison?

With AAAWith divided allocationContribution
SurvivalExtinction+ Cases saved
ExtinctionSurvival Cases lost
SurvivalSurvivalNo difference
ExtinctionExtinctionNo difference
These categories carry no assigned probabilities. Unchanged outcomes do not erase an advantage in the cases where allocation matters.

Net survival gain= probability of cases saved
− probability of cases lost

Discoveries, delays, displaced work, and institutional responses enter this same comparison. AAA improves survival when the probability of cases saved exceeds the probability of cases lost.

The causal identity
p(τ)p(ρ)=Pr(Sτ=1,Sρ=0)Pr(Sτ=0,Sρ=1).(2)p(\tau)-p(\rho)=\Pr(S_\tau=1,S_\rho=0)-\Pr(S_\tau=0,S_\rho=1). \tag{2}

Here τ is an AAA policy and ρ its comparator. The marginal survival probabilities determine the net difference. Separately estimating the saved and lost probabilities requires a model of their joint outcomes.

Paper §4 · The survival comparison ↗

Why the claim reaches the entire budget.

Extinction ends humanity’s continued existence and the future it makes possible. The paper therefore gives every positive net survival improvement absolute priority over secondary benefits.

A diminishing return can still be a positive return. While a remaining alignment use carries a net survival advantage over diversion, the reason to dedicate that resource remains.

The stopping point is the full budget when every diversion sacrifices a positive net survival opportunity.

This is the paper’s case for exclusive dedication. It applies to the purpose of all redirectable resources, while leaving the research program free to pursue the most effective methods.

Why a small positive gain still decides

Compare survival probabilities first. Compare secondary value only within an exact survival tie. This is the paper’s adopted moral ordering:

πρp(π)>p(ρ) or [p(π)=p(ρ) and v(π)>v(ρ)].(1)\pi\succ\rho\quad\Longleftrightarrow\quad p(\pi)>p(\rho)\ \text{or}\ \bigl[p(\pi)=p(\rho)\ \text{and}\ v(\pi)>v(\rho)\bigr]. \tag{1}

Every positive gain satisfies the first comparison, however small. Uncertainty about a difference is not an exact tie. Effects of welfare, autonomy, or economic activity on extinction risk belong in the survival assessment itself.

Paper §3 · Survival priority ↗
The condition for full dedication to be necessary

In the finite decision model, A(h) contains feasible actions and E(h) eligible alignment actions. V is the unrestricted optimal continuation value and Q the value of an action followed by optimal continuation.

Q(h,a)=E[V(H)h,do(a)],V(h)=maxaA(h)Q(h,a);V(h)=s(h) at terminal histories.(3)Q(h,a)=\mathbb E[V(H')\mid h,\operatorname{do}(a)],\qquad V(h)=\max_{a\in A(h)}Q(h,a);\qquad V(h)=s(h)\ \text{at terminal histories}. \tag{3}

A policy’s survival deficit is its expected accumulated action gaps g(h,a) = V(h) − Q(h,a):

V(h0)p(π)=Eπ ⁣[t<Tg(Ht,At)].(4)V(h_0)-p(\pi)=\mathbb E_\pi\!\left[\sum_{t<T}g(H_t,A_t)\right]. \tag{4}

Every survival-maximizing policy complies with AAA if and only if all maximizing actions at every optimally reachable history are eligible:

arg maxaA(h)Q(h,a)E(h)for every optimally reachable h.(5)\operatorname*{arg\,max}_{a\in A(h)}Q(h,a)\subseteq E(h)\qquad\text{for every optimally reachable }h. \tag{5}

The full paper gives both proof directions. The condition establishes exclusivity: full dedication is required to attain the best available survival value. Full dedication still permits different research portfolios and institutional arrangements.

Paper §5 · The theorem and proof ↗

Uncertainty leaves a reason to act.

The possibility that humanity becomes extinct under both policies does not cancel the cases where alignment effort changes the outcome. A commitment justified by its net survival advantage can remain the right decision even when it ultimately fails.

This is the starting point for a subsequent paper on the value of undertaking AAA even in histories where humanity does not survive.

Read the precursor in the paper

Follow the argument further.

The complete paper contains the formal model, proofs, policy analysis, and 24 references.

Read the complete paper
What does the full budget include?

All feasibly redirectable resources, irrespective of owner or sector: compute, agent operating time, human research effort, and supporting resources. Essential infrastructure and dependencies that sustain effective alignment capacity belong within the mission.

Paper §2.1 ↗
Does AAA prescribe a research technique?

It governs allocation. Single agents, swarms, human research, experiments, formal verification, and work on better methods can all serve the mission. The program should move resources toward their strongest remaining alignment use.

Paper §2.2 ↗
How does model training enter the comparison?

The two worlds share the same training regime. The allocation argument can be examined under continued training or under a training restriction, with the resources released by a restriction included in that comparison.

Paper §2.1 ↗
Who chooses how to implement it?

The commitment is universal. Governments, developers, infrastructure providers, funders, and other resource owners choose arrangements suited to their responsibilities. Their choices are assessed by the alignment work and protection they make possible.

Paper §8 ↗

Source