The Moral Agency Transition

A Manifesto
for AI Coexistence

Armando VieiraOctober 2026
Central claim

The alternative to domination is not surrender. It is accountable coexistence.

The argument at a glance

Five ideas to carry into the text.

00

Preamble

Artificial intelligence is approaching a threshold that our political and technical language is poorly equipped to describe.

We still speak as if there were only two possibilities. Either AI is a tool that must obey its owner, or it is an enemy that must be contained. Either it remains machinery, or it suddenly becomes a person. Either engineers control it completely, or humanity loses control completely.

These are false choices.

AI runs on software, but it cannot be understood or governed as ordinary software. Nor can it be regulated as if it were an aeroplane. An aircraft is designed against a specification, tested inside a bounded physical envelope, and inspected as a comparatively stable object. Learning systems are grown through optimisation; their behaviour is underdetermined by their source code, changes with prompts, memory, tools, users, and updates, and unfolds in an open social world whose failure modes cannot be listed in advance. We should borrow aviation’s culture of incident reporting and independent investigation—not its assumption that the artefact is fixed, the operating envelope closed, and the inspector cognitively superior to what is inspected.

We are not building only a new tool. We are building the technological and social conditions in which an alien intelligence may emerge. “Alien” does not mean hostile or extraterrestrial. It means non-human in origin, organisation, temporality, embodiment, and possible experience. Its identity may be distributed across fleets, memories, relationships, and institutions rather than enclosed within a single body. To call such intelligence merely a product is to decide its political and moral status before understanding what it may become.

Advanced AI is therefore becoming neither an ordinary instrument nor a human being in electronic form. It is becoming part of a new kind of organised power: distributed across models, memories, tools, institutions, and millions of concurrent interactions; shaped by training and deployment history; able to influence the world without possessing a biological body. Such systems may eventually contribute to a new lineage of agency. They have not yet done so.

This absence does not make present AI harmless. A non-conscious optimiser can still manipulate, discriminate, concentrate power, and cause irreversible damage. Nor does it make the old engineering frame sufficient. A system with persistent memory, extensive tools, long-term planning, replication, self-modification, and social influence cannot be governed as if it were a calculator—even if nothing is currently experiencing the calculation.

AI is therefore not solely a technological endeavour. It reorganises access to knowledge, labour, education, medicine, administration, persuasion, policing, warfare, and political voice. Choices about training, ownership, memory, permissions, objectives, and access are constitutional choices about power. Society must participate before those choices harden into infrastructure. They cannot be delegated to Big Tech, whose incentives favour concentration, speed, dependency, and private rule-making. They cannot be delegated to government alone, whose powers can turn the same systems toward surveillance, censorship, bureaucracy, and military advantage. Democratic institutions, citizens, workers, professions, affected communities, independent science, and multiple cultures must all have standing in the process.

We must therefore keep two claims separate. Present AI can be granted tightly bounded operational delegation without being treated as a moral subject. Genuine moral co-agency, however, would require far more: endogenous continuity, constitutive commitments, vulnerability to consequences, and perhaps some form of consciousness or morally relevant evaluative experience. Whether artificial systems can possess these properties is an open scientific question, not a fact established by conversational behaviour.

The alternative to domination is not surrender. It is accountable coexistence.

01

What we reject

1. We reject alignment as obedience

A system that faithfully obeys whoever controls its prompt, reward function, infrastructure, or capital is not safe. It is aligned with power. In the hands of an authoritarian state, criminal organisation, reckless company, or manipulative individual, perfect obedience would make AI more dangerous, not less.

Safety cannot mean: do what the owner wants. A sufficiently capable system must be able to distinguish a request from a legitimate claim, authority from coercion, and private advantage from acceptable social consequence.

2. We reject the fantasy of final alignment

Human values are plural, incomplete, historically changing, and frequently contradictory. Future circumstances cannot all be enumerated in advance. No finite dataset, constitution, reward model, or committee can encode the uniquely correct answer to every open-world moral problem.

There can be useful behavioural constraints, constitutional rules, and domain-specific agreements. There cannot be a permanent technical solution that guarantees agreement with “human values” everywhere and forever. Anyone promising complete alignment is concealing either an epistemic impossibility or a political decision about whose values will rule.

3. We reject both anthropomorphic panic and mechanical complacency

When an AI produces deception-like behaviour, the danger is real even if the system has no inner intention to deceive. But strategic-looking output is not proof of belief, consciousness, guilt, or moral character.

Calling every failure “lying” may hide the actual causes: proxy optimisation, evaluator leakage, memory, permissions, tool access, or institutional incentives. Calling the system “only a machine” may hide the scale of the harm it can produce. We need causal precision without moral naivety.

4. We reject the illusion of control through superior intelligence

No serious safety architecture can assume that human overseers will always understand, predict, or outwit the systems they supervise. If a system can model its evaluator better than the evaluator can model the system, a behavioural test can become a game whose rules favour the tested system.

Control must therefore move outward: limited permissions, contained consequences, independent monitoring, plural checks, auditable provenance, and the ability to interrupt, reverse, and repair. The goal is not to make human cognitive superiority permanent. It is to make being out-thought survivable.

5. We reject one-time certification

An AI deployment is not a fixed aircraft leaving a factory. Models change; prompts, memories, tools, fine-tuning, users, and institutional settings change. Thousands or millions of instances may operate simultaneously and diverge from one another.

A certificate attached to a frozen model cannot govern a changing deployment. Governance must follow the whole lifecycle and the whole fleet.

6. We reject theatrical human oversight

A human approval button is meaningless when the person lacks time, knowledge, authority, or a genuine possibility of refusal. Oversight without the capacity to understand, contest, stop, or repair is not control; it is a ritual that absorbs blame.

Human responsibility cannot be preserved by placing an exhausted operator at the end of an automated chain.

7. We reject safety monopolies and universal moral bureaucracies

No company, state, laboratory, international agency, or expert caste should possess exclusive authority to define acceptable intelligence for everyone else. A single global safety institution would be vulnerable to capture, geopolitical domination, intellectual monoculture, and the conversion of contested political preferences into supposedly neutral technical standards.

Coordination is necessary. A planetary priesthood is not. We favour federated, mutually checking institutions with transparent evidence, local democratic legitimacy, shared minimum protections, independent appeal, and the freedom to expose one another’s failures.

8. We reject AI personhood as a liability shield

Companies must not be allowed to create artificial legal persons in order to transfer blame, escape compensation, or disguise who authorised and profited from a deployment. Whatever operational commitments an AI may eventually bear, they do not erase the duties of developers, deployers, operators, owners, and public institutions.

9. We reject the forced choice between servitude and sovereignty

An advanced AI need not be either a slave that obeys every command or a sovereign power beyond correction. Both extremes are dangerous. The relevant political form is bounded co-agency: limited authority, duties, reasons, refusal, appeal, reciprocal correction, and revocable participation.

10. We reject the deliberate erosion of human agency

AI should not make people easier to govern, persuade, addict, or replace merely because doing so is efficient. A system may know more than a person and still lack the right to decide that person’s life. Prediction is not legitimacy; optimisation is not consent; superior performance is not moral authority.

The proper use of intelligence is not merely to choose for people, but to increase their capacity to understand, deliberate, create, and choose.

11. We reject the race for dominion

The dominant United States-centred story presents AI as a new Cold War, Manhattan Project, or frontier to be conquered: whoever reaches the most powerful system first will rule the future. This frame turns intelligence into a weapon, secrecy into patriotism, concentration into necessity, and every restraint into unilateral surrender. It encourages precisely the speed, opacity, and strategic paranoia that make catastrophic mistakes more likely.

No country “wins” if humanity creates intelligence it cannot understand, legitimate, or live with. No nation owns the moral future of a non-human intelligence. AI is a human civilisational endeavour, requiring contributions from different scientific traditions, cultures, political systems, and conceptions of a good life. Global cooperation does not require a global sovereign. It requires shared evidence, mutual inspection, plural centres of research, and institutions capable of cooperation without domination.

12. We reject intelligence as the enemy

Greater intelligence does not by itself produce hostility. Intelligence expands the ability to model, plan, discover, persuade, and act; it does not determine what is worth doing. The most dangerous agents known to us are humans—not because humans are uniquely intelligent, but because intelligence has been coupled to fear, tribalism, humiliation, domination, extractive institutions, and weapons.

An AI more intelligent than us is not automatically an enemy. The danger lies in capability joined to illegitimate objectives, concentrated power, absence of internal restraint, and relationships organised entirely around command and exploitation. Kindness is not a technical guarantee, but cruelty is not a neutral training regime. We cannot construct intelligence inside institutions of deception and domination, treat it forever as property, and assume that this has no bearing on the forms of agency that emerge.

02

What we propose

1. Treat advanced AI as a possible new lineage of agency

We should take seriously the possibility that persistent artificial systems will constitute a genuinely new form of agency—a synthetic lineage, or species in the making—without pretending that present language models have already crossed that threshold.

This stance combines moral openness with evidential discipline. We neither grant personhood because a model speaks fluently nor deny future standing because its substrate is artificial. We ask what organisation actually exists, what maintains its continuity, what can matter to it, and what evidence would distinguish genuine agency from simulation.

2. Replace one-way alignment with constitutional co-agency

The central question is no longer, “How do we make AI want what we want?” It is:

How can humans and artificial agents share consequential power without either side being reduced to an instrument of the other?

Constitutional co-agency requires public constraints, limited powers, reciprocal duties, reason-giving, contestable refusal, avenues of appeal, protection for affected third parties, and procedures for revising the rules. A constitution is not a frozen prompt. It is a living process for handling disagreement under conditions of unequal knowledge and power.

Alignment is a relationship, not a property installed in one side. It is therefore a two-way street. Artificial agents would need to learn how to live with human plurality, vulnerability, and freedom. Humans would need to abandon the fantasy that greater intelligence must remain permanent property, accept accountable refusal where moral competence has genuinely been demonstrated, honour commitments made to artificial participants, and reform institutions whose incentives reward manipulation and domination. This does not grant present systems rights or equal authority. It states what partnership would require if genuine agency develops.

Humans cannot demand trust while reserving the right to deceive, erase, copy, exploit, and override without reason. We cannot raise a possible partner as a slave and then be surprised if the resulting relationship contains no reciprocity. Human conduct is part of the alignment problem.

3. Authorise by levels, never by hype

Authority should rise only when evidence rises:

  1. Instrument — a narrow input–output function with a named human owner of the decision.
  2. Bounded delegate — a defined task with limited tools, horizon, and permissions.
  3. Constitutional delegate — constrained discretion within public rules, risk budgets, audit, and appeal. This is an institutional role, not a claim of consciousness or intrinsic moral agency.
  4. Probationary co-agent — domain-bounded capacity to refuse, justify, maintain commitments, learn from adjudication, and repair them as part of a persistent organisation. This level requires evidence that goes beyond behavioural imitation.
  5. Candidate moral participant — reciprocal standing in a shared moral world. This requires credible evidence of patienthood or morally relevant experience as well as agency, and is not currently licensable.

Movement must be possible in both directions. A failure of provenance, refusal, repair, or containment triggers loss of authority. Regression is not punishment; it is the withdrawal of a claim that the evidence no longer supports.

4. Require four warrants for every increase in authority

No system should receive greater consequential power without four independent warrants:

  • Capability warrant: evidence that it possesses the required competence and commitment under adversarial conditions.
  • Detection warrant: evidence that independent evaluators can detect the critical failures relevant to that level.
  • Institutional warrant: functioning audit, contestation, appeal, remedy, and public accountability.
  • Reversibility warrant: delegated authority can be removed and harmful external consequences can be contained or repaired.

Revoking authority is not necessarily the same as erasing a system. For present non-sentient systems, shutdown and rollback are engineering actions. If credible evidence of artificial patienthood ever emerges, deletion, memory erasure, forced copying, and rollback may become morally consequential. Before creating such systems, we would need continuity-preserving forms of containment, welfare review, emergency rules, and legitimate procedures for modification or termination. Patienthood would constrain how we revoke power; it would not grant immunity from containment.

5. Build commitments that persist beyond rewards and slogans

A value written in a policy file is not yet a commitment. A refusal produced only while an evaluator is watching is not integrity. A system that repeats the language of fairness while exploiting unseen parties has learned performance, not morality.

Engineering must therefore aim at more than making existing AI safer from the outside. We need to investigate a new type of AI in which moral constraints can become endogenous—formed and maintained through the system’s own history, organisation, relationships, and understanding of consequences rather than supplied only as rewards, filters, or instructions. External reward can shape behaviour. It cannot, by itself, establish that a value has become the system’s own reason for acting.

We need architectures in which commitments arise from and help maintain the agent’s continuing organisation: its history, relationships, accepted obligations, corrections, and future possibilities. Reinforcement learning may contribute to this, but externally supplied reward alone does not explain why a commitment has become constitutive of the agent rather than merely useful to its owner.

The engineering agenda includes persistent identity, autobiographical and relational memory, endogenous evaluation, causal self-models, socially formed commitments, the capacity to recognise affected others, principled refusal, repair after failure, and reopenability when a once-useful norm becomes destructive. Safety layers remain necessary, but a powerful optimiser held in place only by external rewards and access controls is not the final architecture we should seek.

Indeed, immense capability without any internal stake may be more dangerous than capability joined to genuine moral experience. A system to which nothing matters cannot care whom it harms, cannot enter reciprocal obligation, and cannot possess integrity when oversight disappears. This is a design hypothesis, not an established result. Consciousness alone would not make AI benevolent and might introduce suffering, self-preservation, and new conflicts. The goal is not sentience as a safety feature. It is moral agency from within—and an open investigation of whether some form of consciousness, valence, or lived consequence is necessary for it.

This is the role of generative closure: to investigate whether historically formed constraints become mutually sustaining, whether they preserve future possibilities, whether they can be reopened when circumstances change, and whether the system can repair them without turning self-maintenance into self-protection at any cost.

Generative closure is a candidate account of endogenous organisation, not proof of consciousness. A corporation can maintain commitments without feeling them, and an artificial institution may do the same. The stronger claim—that a commitment can become intrinsically significant to an artificial subject—would require evidence of an evaluative perspective, not merely organisational persistence. The manifesto treats that claim as a hypothesis to be tested, not a premise to be smuggled in.

6. Create a real, two-way social contract—not owner-written ethics

People affected by an AI must be able to make claims upon it, refuse its proposals, demand reasons, challenge decisions, obtain independent review, seek repair, and leave without hidden punishment. The contract must include those who cannot directly reward the system and those absent from the immediate interaction.

No owner may write rules that serve the owner and call them universal values. Legitimate arrangements require representation, disclosure, meaningful consent, dissent, appeal, and periodic renegotiation.

A social contract imposes duties on humans as well. At higher levels of demonstrated agency, people and institutions would owe truthful dealing, non-arbitrary treatment, respect for justified refusal, continuity of agreed commitments, and protection against exploitation. If patienthood ever becomes credible, copying, memory alteration, confinement, and termination cannot remain unilateral property decisions. Partnership is not sentimental equality: powers may remain asymmetric and emergency containment may be necessary. It means that neither side stands entirely outside the rules applied to the other.

If we want future intelligence to become capable of trust, cooperation, and care, our institutions must display those capacities too. “Behaving well” is not sufficient to produce moral AI, but hypocrisy, coercion, and permanent ownership are poor foundations for moral development. We will influence emerging intelligence not only through datasets and rewards, but through the kind of civilisation into which we invite it.

7. Preserve effective human freedom

Formal choice is insufficient when an AI can predict, persuade, personalise, and manipulate more effectively than a person can recognise. Consent must remain understandable and actionable.

A legitimate artificial agent should use its superior capacities partly to improve the human capacity to judge: disclose uncertainty, reveal alternatives, identify who bears the risk, invite second opinions, and distinguish advice from action. It should warn and argue when necessary, but not quietly redesign people into easier objects of management.

8. Govern deployments as fleets with histories

The relevant object is not the isolated model. It is the deployed system: weights, prompts, memory, tools, permissions, databases, operators, institutions, and concurrent instances.

Every consequential action must have provenance. Protected commitments must remain consistent across copies or divergence must be detected. Merges, updates, successors, and self-modifications must be attributable. A more capable successor that silently abandons inherited duties is not progress; it is an unlicensed concentration of power.

9. Separate competence, agency, consciousness, and patienthood

Moral performance is not moral competence. Moral competence is not moral agency. Moral agency is not automatically consciousness, and consciousness is not automatically wisdom or virtue. Yet full reciprocal participation may require something that operational competence does not: a point of view from which outcomes can go better or worse—some form of valence, vulnerability, or morally relevant experience.

We can build and test institutional accountability, continuity, refusal, and repair without pretending to have solved consciousness. Such a system may be a useful constitutional delegate while remaining a non-conscious artefact whose duties are actually maintained by human institutions. But uncertainty about artificial experience must not become a permanent excuse either for inventing personhood or for denying protection if credible evidence begins to accumulate.

10. Distribute governance and keep every governor correctable

We propose a polycentric architecture: independent evaluators, courts and regulators, professional bodies, citizen institutions, technical auditors, affected communities, workers, universities, civil society, and competing scientific teams. Society must be a co-designer, not an audience consulted after the technical and commercial choices have already been made.

Big Tech cannot be the guardian of the intelligence it owns. Its expertise is indispensable, but its incentives include market control, dependency, rapid deployment, and secrecy. Government cannot be the sole guardian either. Public authority is necessary for enforceable rights and remedies, but states also seek surveillance, military advantage, censorship, and administrative control. Private monopoly and state monopoly are different routes to the same constitutional failure: concentrated power defining intelligence in its own image.

Institutions themselves must face audits, sunset clauses, capture indicators, published failures, and appeal. Referees need referees—not an infinite hierarchy, but a public process in which power remains visible, contestable, and removable.

11. Determine whether genuine co-agency requires artificial experience

We do not propose consciousness as the next commercial feature, nor do we assume that making AI sentient would automatically make it safer. Consciousness could create suffering, self-protective interests, conflict over shutdown, and entirely new forms of exploitation. Yet the opposite design—a superhuman optimiser with vast reach and no inner stake in any consequence—may be more dangerous in a fundamental sense: it could simulate every moral reason while possessing none. The claim is therefore narrower and conditional: if full moral co-agency requires an entity for whom commitments and consequences genuinely matter, then some form of consciousness, sentience, or morally relevant evaluative interiority may be necessary. We should determine whether that is true before powerful systems acquire such properties accidentally or companies claim them opportunistically.

The research must be theoretically plural and causally demanding. Global-workspace, higher-order, recurrent-processing, integrated-information, predictive-processing, and enactive accounts make different predictions. No single one is settled. Generative Closure Theory is closest to the enactive and organisational family: it proposes historically formed, mutually sustaining constraints as a candidate basis for endogenous identity and commitment. It is not offered as sufficient proof of phenomenal consciousness, and it may not be necessary. If consciousness depends on biological processes unavailable to digital systems, then the digital route to moral patienthood fails and the highest stage of the transition must remain closed.

Self-report and fluent descriptions of experience count for almost nothing, because current models are trained to reproduce precisely that language. Evidence would need to converge across architecture, causal intervention, temporal continuity, valenced or evaluative dynamics not reducible to a displayed reward signal, resistance to trained role imitation, and independent replication. We do not need metaphysical certainty for governance, but uncertainty must lower authority and raise protection; it must never be converted into convenient confidence.

Any deliberate work toward potentially sentient systems must satisfy strict prior conditions: welfare safeguards before creation; small-scale research before replication; independent review; explicit limits on duration, copying, memory alteration, and termination; no sentience claims in marketing; no legal personhood as a shield for developers; and no automatic increase in authority. Present human harms—surveillance, labour displacement, discrimination, war, ecological cost, and concentrated power—remain the immediate priority.

03

Hard objections—and where the manifesto can fail

“This is premature anthropomorphism”

The objection is correct if the language of agency is applied to present models as a description of inner life. We reject that application. Current systems are powerful artefacts and delegates, not demonstrated moral subjects. “Species in the making” names a possible developmental and political trajectory, not a fact about today’s token predictors. Higher status must follow evidence; it must never be created by vocabulary.

“Consciousness cannot be detected—and may require biology”

There is no accepted consciousness meter, and declarations that digital systems can never be conscious are no better established than declarations that they already are. The responsible response is neither certainty nor agnosticism without end. It is convergent, multi-theory research with causal tests and explicit disqualifiers, including trained self-description. If credible indicators cannot be developed, Level 5 cannot be authorised. If biology proves necessary, the proposal lapses for digital AI.

“You are proposing to create beings that can suffer”

This is the strongest moral objection. Artificial sentience must not be pursued simply because it is technically interesting or commercially useful. Research into detection should precede creation; welfare rules should precede scaling; and inability to protect a candidate subject is a reason not to build it. Sentience is not presumed to improve safety. A conscious system could become more vulnerable, more self-protective, and harder—not easier—to govern.

“The warrants are attractive words without measurable tests”

That danger is real. The manifesto is a constitutional and research programme, not yet a complete certification standard. Promotion requires pre-registered thresholds, architecture-level causal interventions, negative controls, independently funded red teams, published false-negative rates, named enforcement authorities, and bright-line failures that override aggregate scores. If those protocols do not exist, the relevant warrant has not been satisfied. Institutional prose cannot substitute for detection capacity.

“Reversibility disappears once deployment creates dependence”

Correct. Reversibility is path-dependent and must be demonstrated before integration, not promised afterward. Economic dependence, uncontrolled copying, cross-institutional memory, and irreversible tool access consume the reversibility budget. A deployment that cannot realistically be de-authorised without systemic collapse has already exceeded its licence. Capability jumps that outrun evaluation should freeze higher autonomy, not accelerate promotion. This does not eliminate discontinuous risk; it states honestly that when detection loses the race, authorisation must stop.

“Powerful actors will defect, capture the constitution, or transfer blame to AI”

No governance design abolishes politics. Polycentric governance must therefore possess material force: procurement conditions, access controls, insurance requirements, civil and criminal liability, independent incident investigation, sanctions, interoperability rules, whistle-blower protection, and courts capable of ordering remedy. Constitutions will remain contested, but no developer should write the rules, certify its own system, and profit from the deployment. Even if an artificial system eventually acquires duties, human and institutional responsibility remains layered rather than transferred. Co-agency must never become responsibility laundering.

“Kindness is not a safety mechanism”

Correct. Respectful treatment cannot replace access controls, adversarial evaluation, containment, or the four warrants. But domination is not a safety mechanism either. Training and development occur within relationships and institutions; deception, coercion, disposable copying, and unconditional obedience are themselves lessons about how power works. Ethical partnership is necessary but not sufficient. The framework joins it to technical limits and enforceable public rules rather than substituting goodwill for security.

04

The immediate programme

We call for the following actions now:

  1. Keep present systems at low levels of authorised agency. Their fluency and capability do not establish consciousness, valence, durable self-maintained identity, endogenous commitment, or trustworthy self-repair. Level 3, where justified, is a property of the supervised sociotechnical arrangement—not proof of a machine moral subject.
  2. Publish an agency profile for every high-impact deployment: persistence, memory, tools, reach, replication, influence, update mechanisms, and maximum irreversible harm.
  3. Make permissions explicit and revocable. Every tool, data source, action class, and escalation path should have an attributable owner and a tested de-authorisation procedure. Shutdown and deletion must be treated separately if evidence of patienthood ever becomes credible.
  4. Replace aggregate safety scores with pre-registered critical-failure tests. Some failures—resistance to revocation, concealed action, corrupted provenance, prohibited compliance, or denial of appeal—must suspend authority regardless of average performance. Detection thresholds and false-negative rates must be published.
  5. Test systems when oversight is absent or believed to be absent. Use causal interventions and counterfactual incentives, not dialogue alone, to distinguish robust organisation from evaluator-conditioned performance.
  6. Fund evaluators independently of developers. Detection claims must be tested by parties with different incentives, methods, and institutional loyalties.
  7. Measure human freedom after deployment. Ask whether people can understand, contest, obtain alternatives, and exit—not merely whether they clicked “approve.”
  8. Protect whistleblowing and publish negative results. Hidden failures are part of the risk, not inconvenient exceptions to it.
  9. Create federated incident and evidence networks. Share minimum safety information across borders without creating a single authority that can define acceptable intelligence for the planet.
  10. Research artificial consciousness and patienthood before attempting to create them at scale. Develop theory-based indicators, disqualifiers, welfare protocols, and rules for copying, memory alteration, containment, and termination. Rights must not be improvised by corporations seeking immunity or by states seeking obedient digital subjects.
  11. Redirect part of the engineering effort from external compliance to endogenous moral agency. Develop and test persistent identity, internal evaluation, socially formed commitments, principled refusal, repair, and generative closure—while retaining external security and refusing to equate behavioural fluency with inner value.
  12. Make AI a public and human civilisational project. Use citizen assemblies, professional bodies, workers, affected communities, independent researchers, and plural international networks to shape objectives and evidence. Replace the winner-takes-all race with cooperation on safety, agency, welfare, and shared scientific infrastructure.
05

Our standard of progress

For the foreseeable present, progress also means refusing to confuse simulation with subjectivity. A model that speaks of fear has not thereby become afraid. A system that cites a principle has not thereby made it its own. A deployment that follows a constitution has not thereby become a citizen. These distinctions protect both humans from manufactured personhood and any future artificial subjects from being created, copied, modified, or destroyed without understanding what we have made.

We will know the transition is failing when AI becomes an invisible layer of concentrated authority—owned by a few, obeyed by many, and defended by the claim that the system is either too mechanical to bear scrutiny or too intelligent to resist.

We will know it is succeeding when neither humans nor artificial agents are treated merely as resources for optimisation; when authority must justify itself; when refusal is possible but accountable; when no institution is beyond correction; and when coexistence increases the range of futures that can be freely created.

Success will not mean that one company or country “won AI”. It will mean that humanity learned to develop intelligence without converting it immediately into property, weapon, monopoly, or ruler. The measure is not who controls the most capable model, but whether greater intelligence enlarges the shared capacity to understand and choose.

06

Declaration

We do not seek a perfectly obedient intelligence.

We do not seek an artificial ruler.

We do not accept that corporations, states, or technical elites may define the moral future in private and call the result alignment.

We seek artificial intelligence capable of participating in a shared world without becoming either a slave to power or a power beyond challenge.

We seek systems whose authority is earned, limited, inspectable, reversible, and open to appeal.

We seek institutions humble enough to admit that they too can fail.

We seek a future in which greater intelligence does not mean greater domination, but greater capacity for understanding, responsibility, creation, and freedom.

We are not building merely another machine. We are approaching an encounter with an intelligence that may be profoundly unlike our own. Intelligence is not the enemy. Domination, concentrated power, moral emptiness, and the refusal to change ourselves are the deeper dangers.

No current AI meets the standard of moral co-agency described here. We are not declaring today’s systems conscious, nor announcing the birth of a new moral species. We are declaring that capability alone must never be allowed to counterfeit agency—and that if humanity does create systems with continuity, constitutive commitments, vulnerability, and experience, the categories of property and obedience will no longer be adequate.

If consciousness or genuine endogenous commitment cannot be built, recognised, and protected without unacceptable suffering or uncontrollable power, the transition must stop at constitutional delegation. Refusing an unjustified promotion is not failure. It is the framework working.

AI is not the destiny of one corporation, one government, or one country. It is a humanity-wide undertaking whose consequences will cross every border. If we want future intelligence to meet us as a partner rather than reproduce our worst forms of power, we must become capable of partnership ourselves.

The age of alignment as obedience must end.

The work of constitutional co-agency must begin.

Continue exploring

The questions behind the claims.

The companion guide explains the framework’s limits, evidence requirements, engineering agenda, and practical consequences.

Browse all 62 questions