Abstract
The article argues that many contemporary AI alignment practices risk a mistaken assimilation of moral agency to statistical learning. Techniques such as reinforcement learning from human feedback and constitutional AI often treat morality as a behavioral function that can be approximated from human discourse, behavior, and large-scale interaction data. Drawing on Tomasello’s account of shared intentionality, Hegel’s theory of recognition, and the second-personal tradition in moral philosophy, the article contends that moral agency is not exhausted by norm-conforming output, but presupposes participation in a space of mutual accountability, justificatory practice, and normative self-binding. On this view, and given the architecture of present generative systems, the production of normatively fluent behavior and reason-like discourse is not sufficient for moral partnership: absent interpersonal identity, recognition, and second-personal answerability, such systems can at best simulate the outer form of moral conduct. The central risk, therefore, is not that these systems behave “immorally,” but that their reliable norm-conforming behavior is misread as evidence of genuine moral commitment, encouraging misplaced trust and over-ascription of responsibility. The article proposes a reframing of alignment from an ethical-pedagogical project to an institutional one: instead of presuming or attempting to cultivate artificial moral agents on the basis of behavioral proxies, governance should focus on designing legal, technical, and organizational structures that constrain and render machine behavior auditable and accountable, while keeping moral responsibility firmly anchored in human agents and institutions.
Similar content being viewed by others
Contemporary alignment research often proceeds as if morally appropriate behavior could be treated as a learnable performance profile—something an artificial system can be trained to produce by optimizing over human feedback and large-scale behavioral traces. This assumption is rarely stated as a philosophical thesis, yet it animates a range of initiatives aimed at getting machines to ‘do the right thing’: to act safely, responsibly, and in line with human expectations. Whether through reinforcement learning from human feedback (Christiano et al., 2017), inverse reinforcement learning (Russell, 1998), preference learning (Ouyang et al., 2022), or constitutional refinement (Bai et al., 2022), the recurrent idea is that the moral landscape can be navigated by systems equipped with sufficient data, modeling capacity, and optimization power.
But the assumption deserves closer scrutiny. It encourages a structural assimilation of moral agency to inferential learning: what human beings express as moral judgment is treated as preference prediction, and what counts as ‘moral behavior’ is treated as output adequacy under statistically recoverable constraints. Yet this obscures a deeper fault line. Unlike classification, translation, or even strategic planning, moral agency is not exhausted by correct performance. It involves standing in relations of accountability in which reasons are not merely produced as plausible explanations, but can be demanded, contested, and owned as commitments. A system may optimize for civility or trustworthiness and thereby mimic moral sensibility, while remaining outside the second-personal grammar of obligation that gives moral claims their force (Boddington, 2017).
To see why this distinction matters, consider the three primary sources from which AI systems might plausibly ‘learn’ morality. First, there is human discourse: the vast corpus of moral claims, arguments, policies, and judgments generated in texts, debates, and ethical theories. But this discourse is internally inconsistent, context-sensitive, and culturally embedded. It does not speak with a unified voice but with a polyphony of values in tension with one another (Gabriel, 2020). Second, there is human behavior—the actual decisions and actions taken by individuals or institutions in morally salient situations. But human behavior, as empirical ethics has shown, often diverges from normative ideals. People act not only out of moral conviction but from habit, fear, self-interest, or social pressure (Greene, 2013). An AI trained on behavioral data may learn to predict what humans do, but this is no guarantee that it will learn what humans ought to do. To treat the descriptive regularities of conduct as if they carried their own normative authority would be to smuggle in precisely what moral reflection is meant to test; in philosophical terms, it risks committing the naturalistic fallacy—deriving an ‘ought’ from an ‘is.’ Third, there is the statistical structure of the data itself: patterns of sentiment, approval, or aversion that can be inferred from large-scale interaction. This is the foundation of sentiment analysis, content moderation, and large language model alignment. But here the risk is different: moral complexity is flattened into probabilistic association. The moral becomes indistinguishable from the popular, the nuanced from the frequent. What emerges is not ethical reasoning but moral fluency—language that mimics human-like concern with none of its inner tension. The result is a form of normative performance without normative participation.
Moral agency, in the sense relevant to this article, is not the mechanical output of a trained function but the situated expression of shared intentionality and reciprocal accountability. It arises when agents enter into a space of mutual recognition, interpretive responsibility, and norm-governed expectation—a ‘we’-perspective in which obligations can be addressed, acknowledged, and repaired. This space is irreducibly intersubjective: its meaning does not supervene on rules alone, because what a rule amounts to depends on who applies it, to whom, and under which shared understandings.
Thus, even a system that reliably produces what we judge to be the morally appropriate response may still miss the moral point. Its outputs can align with our expectations without being grounded in a shared horizon of meaning or answerability. This is the conceptual misalignment at the heart of proxy-based alignment strategies: they can replicate the surface structure of moral life while remaining agnostic to the participatory conditions under which moral reasons become binding. The risk is not merely occasional error, but a systematic over-reading of reliability as commitment—and of norm-conforming performance as evidence of moral responsibility.
1 The Conceptual Misalignment — Why “Moral Learning” Is Not Learning
Contemporary alignment strategies are often premised on a functionalist, proxy-based picture: that morally acceptable behavior can be captured by optimization procedures targeting human feedback, preference satisfaction, or norm compliance. On this picture, to ‘learn morality’ is to infer generalizable patterns from human data—language, choice behavior, and cultural artifacts—so that outputs track what humans approve of or expect. What is overlooked, however, is the conceptual distance between the appearance of moral action and the structure of moral commitment. A system that maximizes reward for ‘polite’ or ‘trustworthy’ responses may produce behaviorally adequate outputs while remaining outside the second-personal grammar of obligation that gives moral conduct its force (Bai et al., 2022; Gabriel, 2020). The apparent ‘moral competence’ of current systems thus rests on precisely the unstable mix described above: heterogeneous moral discourse, opportunistic practice, and statistically amplified regularities. A model trained on this amalgam may learn to emulate moral behavior while leaving open whether anything like moral reasons—reasons that can bind and be reciprocally owned—are in play at all.
At this point a distinction is needed. Generative systems can produce ‘reasons’ in at least one sense: they can generate reason-like explanations—linguistically coherent justifications that fit human conversational norms. But moral agency requires more than the capacity to emit justificatory tokens. It involves reasons as commitments: considerations an agent can recognize as bearing on what she ought to do, revise under criticism, and answer for. And it involves shared reasons: reasons that are not merely stated, but can be demanded, contested, and sustained within a relationship of mutual accountability. The argument of this paper concerns this stronger, second-personal sense.
This distinction is not merely about epistemic access to moral facts, but about the practical structure of moral agency. Moral agency, as classically conceived, is not just the capacity to follow rules but the ability to understand oneself as answerable to reasons one can share with others. This presupposes intersubjective participation: a space in which norms are not only obeyed but interpreted, contested, and reaffirmed through relations of recognition (Hegel [1821 ] 1991). It also presupposes epistemic conditions—one must grasp what one is doing and to whom it matters—but such grasp is not sufficient. The moral point of view depends on the ability to engage in justificatory practices: to give and ask for reasons, to reflect, to revise, and to take responsibility in the face of others (Tomasello, 2019).
Artificial systems, by contrast, may produce norm-conforming behavior without participating in the shared horizon that gives norms their second-personal meaning. Their ‘learning’ is closer to behavioral imitation and pragmatic calibration than to moral understanding as commitment. While this may suffice for many practical purposes, it does not by itself warrant treating such systems as moral partners. A system that optimizes for civility or trustworthiness can mimic moral sensibility while lacking the capacity for shared intentionality and answerability. The critical concern is therefore diagnostic and practical: when such a system deviates from its learned regularities—under distribution shift, adversarial prompting, or altered incentives—we are tempted to interpret the deviation as a moral breach, when it may be better understood as the failure of a proxy that was never anchored in second-personal obligation.
Most importantly, as Tomasello (2009, 2019) argues, distinctively human moral practice emerges from joint action and shared intentionality rather than from individual rule inference. It is within the ‘we’ of collaborative contexts that obligations arise, can be demanded, and are sustained. If this is right, then intelligence and fluent moral performance do not necessitate moral agency in the relevant sense; it remains an open question whether artificial systems will ever acquire the socio-practical capacities that make second-personal answerability possible. What current systems demonstrably learn is something else: to generate outputs that correlate with human moral expectations and conversational norms. Alignment, in this proxy-based register, is therefore not the reproduction of morality, but an approximation of its publicly legible surface.
2 Anthropological Foundations — Tomasello and the Emergence of Shared Intentionality
A key pillar in the distinction between intelligent behavior and moral agency lies in the nature of cooperation. Drawing on the developmental psychology of Michael Tomasello, we encounter a foundational insight: human morality does not arise from intelligence alone, but from the uniquely human capacity for shared intentionality—the ability to form, sustain, and act upon commitments that are jointly understood and normatively binding. This capacity, according to Tomasello (2019), marks a fundamental shift in hominin evolution: a transition from individual to collaborative cognition and action.
In his empirical studies comparing human children with great apes, Tomasello has shown that while non-human primates exhibit sophisticated instrumental intelligence, they lack the socio-cognitive infrastructure necessary for joint action. Chimpanzees, for instance, may cooperate strategically but fail to establish mutually recognized goals that imply shared responsibilities or obligations (Tomasello, 2016). Human children, by contrast, from as early as age two, engage in activities that presuppose a “we-perspective”: they protest when partners defect, attempt to re-engage them, and even apportion roles within joint tasks. These behaviors suggest not only an understanding of another’s intention, but a recognition that the action is collectively owned and must be completed together (Tomasello, 2014).
Shared intentionality, then, is not simply about joint attention or mutual responsiveness; it entails the creation of a normative space in which each participant holds the other accountable. This space is the crucible of moral practice. It is here that notions of obligation, justification, and norm compliance first emerge—not as abstract rules, but as structures embedded in the experience of doing something with someone else. Morality, on this account, is not merely the regulation of behavior; it is the regulation of cooperation. What makes an act moral is not just its outcome, but the manner in which it acknowledges and affirms the mutuality of intention.
Artificial systems, by contrast, can simulate cooperation, yet it remains an open question whether they can participate in it in the sense required for shared intentionality. They may model preferences, predict human behavior, or optimize for norms statistically derived from data but this does not yet place them within a shared perspective. What is missing in present architectures is not responsiveness as such, but what Tomasello calls “interpersonal identity”—the standing of being a co-agent in a relationship governed by mutually recognized expectations and commitments. As Gabriel (2020) notes, prevailing alignment architectures assume that moral behavior can be approximated through inferential pattern recognition; the worry is that this neglects the constitutive role of co-intentionality in moral cognition. Absent entry into a ‘we’-space, there may be normatively structured performance, but no standpoint from which obligations can be jointly owned.
This divergence suggests a deeper structural asymmetry. Whereas humans develop their moral orientation through social enculturation into shared practices, AI systems are trained via optimization over input–output correlations. Even if such systems appear to comply with norms, their actions lack the interpretive dimension that renders those norms meaningful. Moral behavior, in this sense, cannot be reduced to rule-following; it is always already situated within a context of recognition, accountability, and shared meaning.
3 Orthogonality Thesis — Why Intelligence Does Not Entail Shared Intentionality
The gap between intelligent performance and moral agency becomes clearer when viewed through the lens of Bostrom’s orthogonality thesis, understood here as a claim about non-implication. Intelligence—no matter how advanced—does not by itself secure any particular evaluative orientation or second-personal standing. Of course, moral agency presupposes epistemic conditions: one must understand what one is doing, to whom it matters, and under which descriptions. But epistemic sophistication and behavioral competence are not sufficient for a standpoint from which norms are experienced as binding. The central point is therefore conditional: even highly capable systems may remain unable to see themselves as answerable to reasons they can share with others—and thus may be inappropriate candidates for moral partnership.
Bostrom formulates the orthogonality thesis as the claim that intelligence and final goals are, within broad limits, independent dimensions: a highly capable system can pursue trivial, perverse, or self-destructive objectives (Bostrom, 2012). Transposed into the register of moral cognition, the key point is again one of non-necessitation. What is at stake is not only that an intelligent system might pursue immoral ends, but that increasing intelligence does not, by itself, guarantee the emergence of the second-personal conditions under which moral reasons can bind. Intelligence, understood as instrumental rationality or learning ability, can remain functionally indifferent to moral address. In this sense, the paper proposes a moral analogue of orthogonality: capability does not entail participation.
This becomes evident when one considers three markers of moral agency. First, moral commitments require normative self-binding: the capacity to act not merely in accordance with a rule, but out of recognition of its authority over one’s will (Korsgaard, 1996)—a distinction often captured, in Kantian terms, as acting from duty rather than merely in conformity with duty. Second, they presuppose an understanding of others not as behavioral vectors, but as norm-responsive agents whose perspectives must be acknowledged within reciprocal relations (Tomasello, 2019). Third, moral deliberation involves reasons that can be shared and contested as expressions of common evaluative standards, not merely optimized as outputs of goal maximization or preference modeling.
On the accounts just sketched, present AI architectures do not yet satisfy these conditions. A large language model may approximate morally acceptable behavior by optimizing for human-preferred outcomes, and it may even produce fluent, reason-like explanations for its outputs; but none of this amounts to normative self-binding or second-personal answerability. Its responses are outputs of statistical inference shaped by training signals, not avowals of commitment within a shared space of obligation. Even in multi-agent reinforcement learning settings, the simulation of coordination does not by itself entail the constitution of a moral ‘we’: joint attention, shared intentionality, and second-personal recognition are not thereby secured.
The conceptual error, then, lies in mistaking behavioral alignment for moral alignment. The former is a technical achievement; the latter a phenomenological and intersubjective relation. To the extent that current alignment research equates moral competence with the ability to produce socially acceptable outputs, it reduces the ethical question to whether systems behave in ways we find tolerable, rather than to whether there is any standpoint from which they could take themselves to be bound by reasons at all.
Thus, the orthogonality at stake here is not merely between intelligence and goals, but between capability and participation. An artificial system may simulate participation in moral life; yet such simulation does not, by itself, establish the structures of mutual recognition, commitment, and accountability that moral partnership presupposes.
4 The Problem of Moral Training — What AI Would Actually Learn
If we take seriously the hypothesis that AI systems may fail to access the intersubjective structure of moral experience, then a further question emerges: what, exactly, are these systems learning when we ask them to “learn morality”? Proponents of alignment strategies often assume that by exposing models to extensive human behavior—through preference datasets, language corpora, or reinforcement feedback—we are enabling them to converge on our ethical norms. Yet this assumption collapses under closer inspection, for it presumes that moral behavior is both observable and normatively legible in the data itself.
But the moral life of humans resists such visibility. As studies in empirical ethics and behavioral psychology have long shown, human beings rarely act in accordance with articulated ethical principles (Haidt, 2001; Doris, 2002). Instead, they follow a patchwork of cultural scripts, cognitive shortcuts, social incentives, and situational improvisations. The resulting behavior, though often retrospectively rationalized, is deeply inconsistent across contexts, agents, and timescales. What emerges in human practice is not a stable moral order but a fluctuating mosaic of expectations and justifications—more a rhetoric of normativity than its reliable expression.
Training AI systems on this basis yields troubling consequences. Models absorb the internal contradictions of normative language, the opportunistic character of practical behavior, and the frequency biases of large-scale statistics. They learn what humans typically say, do, and tolerate, but not which reasons they take to be binding. Rather than producing convergence, such training amplifies pluralism without clarity. Instead of ethical commitment, it foregrounds pragmatic accommodations: compliance, reputation management, conflict avoidance. What is extracted from the data is not what we value as such, but what we are willing to do, to endorse, or to leave unchallenged under pressure.
More fundamentally, such training bypasses the very mechanism by which moral understanding is generated: interpretive participation. Human moral agents do not merely observe patterns; they engage in ongoing processes of justification, evaluation, and revision. They disagree, reflect, confess, and forgive. They are bound not only by what they do, but by how they mean. To learn morality is thus not to absorb a dataset, but to enter a community of interpretation—a process that presupposes precisely the kind of joint intentionality that, as argued above with Tomasello, AI systems cannot inhabit (Tomasello, 2019).
In the absence of this participatory depth, AI systems may well display the outer markers of moral conduct—politeness, civility, norm compliance—yet without any internal stake in the norms they instantiate. This, in itself, is not the problem. From a practical standpoint, it matters little whether a system behaves “correctly” for the right reasons or for no reasons at all. The real risk lies elsewhere: in the human tendency to misread such reliable behavior as evidence of underlying moral commitment. We may come to treat these systems as if they were participants in the normative space of reasons, imputing to them the kind of self-binding obligations that only agents capable of shared intentionality can possess.
This misinterpretation introduces a dangerous asymmetry of expectation. If a system lacks the capacity for second-personal commitment, then its apparent steadiness is not a promise but a performance profile—a pattern sustained by training incentives and environmental fit. Deviation is therefore always possible under distribution shift, adversarial prompting, or altered incentives; yet to users it can appear as an abrupt breach of ‘character’ precisely because the system’s normatively fluent behavior invites an attribution of commitment it does not, as such, underwrite.
A brief analogy helps locate the error without implying that moral standing depends on shared projects. Two strangers walking side by side have moral standing in virtue of their shared membership in a normative order; what they typically lack is a relationship-specific structure of expectation and repair. By contrast, two people who agree to walk together acquire additional, second-personal expectations: a sudden departure calls for explanation because a joint undertaking has been established. The risk with aligned systems is that they can elicit the second kind of expectation without the relational basis that would make it apt. What alignment can often deliver is statistical regularity, and in favorable conditions even robust reliability; what it does not thereby secure is normative stability—steadfastness as commitment within a space where reasons can be demanded, contested, and owned. The danger, then, is not primarily that machines behave ‘immorally,’ but that regularity and reliability are misread as stability, encouraging misplaced trust and over-ascription of responsibility.
5 The Limits of Behavioral Alignment — Why Norm-Conforming Outputs Do Not Constitute Moral Agency
The central limitation of contemporary alignment practice becomes visible precisely where it is often declared most successful: the production of reliably norm-conforming behavior. Techniques such as reinforcement learning from human feedback and constitutional refinement (Ouyang et al., 2022; Bai et al., 2022) can generate striking regularities in output, and in favorable conditions even robust reliability across a range of prompts. Yet this success invites a conceptual slide: regularity and reliability are treated as if they were normative stability. As Gabriel (2020) emphasizes, behavioral conformity is not equivalent to ethical understanding; it shows at most that a model has learned to navigate approval signals and conversational constraints. Even when a system produces fluent ‘reasons’ for its outputs, these may function as justificatory tokens rather than as commitments that can bind the agent and be reciprocally owned. A system trained to avoid harmful, impolite, or norm-deviant responses does not thereby act morally; it acts in accordance with learned gradients that reward certain patterns over others.
Why this slide is tempting—and why it is mistaken—becomes clearer when we recall, with Tomasello (2019), how human moral competence is actually formed. It does not emerge from behavioral imitation alone, but from participation in practices of shared intentionality already described above. What anchors norm-guided action among humans is not the pattern of behavior itself, but the intersubjective structure in which behavior becomes accountable. Children learn moral reasons not by mimicking adult conduct, but by entering a “we-perspective” in which partners can demand explanations, protest norm violations, and negotiate responsibilities (Tomasello, 2014). In this space, the distinction between merely doing the right thing and understanding why it is right becomes socially operative.
Moreover, as Hegel’s conception of morality illuminates, the authority of moral norms arises only when the agent internalizes them as an expression of their own free will (Hegel [1821 ] 1991, §135–141). Moral action is not the mechanical execution of externally imposed constraints, but the articulation of a subjective will that takes itself to be bound by reasons. McDowell (1994) makes a parallel point when he argues that to grasp a reason is not to compute a mapping from circumstances to outputs; it is to stand within a normative space in which meaning exerts a rational pull. An AI system that avoids harmful speech because the training objective penalizes it does not thereby respect a norm; it reproduces a pattern that is instrumentally reinforced. In Kantian terms, this is closer to acting in conformity with duty than acting from duty; and in the second-personal register, it remains compatible with the absence of answerability. The point is not to ascribe a hidden psychology to the system, but to mark a structural difference between compliance as performance and (self-)obligation as commitment.
In contrast, human moral agents can maintain a commitment even when incentives shift, because they are answerable within relationships that make deviation intelligible as breach and repairable through justification. This is one sense in which normative stability differs from mere reliability: it is not just the persistence of a pattern, but steadfastness under countervailing incentives within a space of accountability. As Honneth (1995) stresses within the Hegelian tradition, such accountability is grounded in relations of recognition through which individuals become bearers of reasons in the eyes of others. Whether present AI systems can enter relations of recognition and answerability in the relevant second-personal sense remains an open question; yet behavioral alignment does not by itself establish that they do. What it can reliably deliver is a norm-conforming performance profile—an approximation of compliance’s external shape. In this light, alignment is best understood as conditional and context-sensitive: what systems display is learned reliability, not commitment in the second-personal sense.
Thus the limit of behavioral alignment is not merely technical but concerns what, if anything, may legitimately be inferred from norm-conforming performance. Where alignment practices treat output conformity as moral success, they risk mistaking an externally legible proxy for moral self-obligation. They may capture regularities without establishing the second-personal conditions under which reasons bind; and they may optimize for behavior while leaving open whether the relevant intersubjective structures are present. In that case, alignment functions less as the cultivation of moral agency than as the production of a simulation whose moral significance depends on human interpretation and institutional framing.
6 Normative Simulation vs. Normative Participation — Why Artificial Agents May Remain Unable to Enter the Second-Personal Space of Reasons
The distinction between simulated moral competence and genuine moral participation becomes clearest when we consider what second-personal approaches to normativity (Darwall, 2006) describe as the standpoint of mutual address: the normative space in which agents relate to one another not merely as causes of behavior, but as bearers of claims, reasons, and obligations. To occupy such a standpoint is to be answerable: the other’s demand can count against one’s will, and one’s actions are intelligible as more than causally effective outputs. Tomasello’s developmental account suggests that this form of answerability is typically scaffolded by shared intentionality—the ‘we’-mode in which agents jointly undertake tasks, negotiate expectations, and repair defections. The present paper’s claim is conditional and role-specific: whatever else artificial systems may become, the normatively fluent behavior of present generative systems does not, by itself, establish second-personal participation. They can simulate outward markers of accountability, but simulation alone does not secure the relational structure that makes accountability apt.
At the heart of this gap lies the absence, in present systems, of what Hegel calls recognition (Anerkennung): the reciprocal structure through which agents acknowledge one another as responsible selves. In Hegel’s account, moral agency arises not in isolation but through a process in which individuals acknowledge one another as sources of reasons and as objects of accountability. This process is constitutive, not additive; without it, there is no moral standpoint to occupy. We lack firm grounds to think that present artificial systems encounter the other as a normative presence whose address can impose a claim on them; what they demonstrably do is register inputs and optimize outputs under a training objective. Their outputs may resemble acknowledgment—apologies, explanations, even ‘commitments’—but these are best understood as performative tokens shaped by reinforcement and convention, not as avowals within a relation of answerability. They “respond” only insofar as a training signal has correlated certain linguistic acts with favorable reinforcement.
This distinction is sharpened by McDowell’s (1994) insistence that reasons are not external constraints imposed upon an otherwise neutral decision function, but elements within a form of life whose intelligibility depends on being grasped from the inside. To act for a reason is to see the world under a certain aspect—to understand the action as demanded, permissible, or forbidden within a normative horizon shared with others. Artificial systems lack such horizons. They operate within representational structures that mimic the form of reasoning without accessing its content. Their “understanding” of norms is inferentially shallow, governed by statistical proximities rather than by the rational force that makes a norm binding. What alignment techniques produce, therefore, is a simulation of reason-responsiveness: the linguistic shape of justification without the subjective stance from which justification arises.
This is why aligned systems risk becoming ethical simulacra that perform moral personhood without participating in it. Their apparent sensitivity to norms—their quickness to apologize, their careful hedging, their appeals to fairness or harm avoidance—reflects not moral understanding but a pragmatic accommodation to human expectations encoded in the training data. These expectations, as O’Neil (2016) reminds us, are themselves shaped by social pressures, institutional constraints, and cultural conventions. To optimize for such expectations is to align with their surface form, not with their justificatory depth. Thus the system produces what might be called normative ventriloquism: it speaks in the register of morality without inhabiting the standpoint that gives moral speech its authority.
One might object that this sets the bar too ‘internalistically.’ Relational and interactionist accounts of normativity—familiar from debates about simulated emotions and associated, for example, with Dumouchel’s emphasis on the primacy of interaction—suggest that normative phenomena can be constituted in patterns of reciprocal responsiveness rather than grounded in private inner states. On such views, what matters is not what is ‘inside’ the system, but what emerges between agents in stable interactive dynamics. This is an important corrective, and the present argument can accommodate it: the problem is not that normativity must be private or ineffable, but that moral partnership involves a specific second-personal role—being an apt addressee of claims, blame, and repair. Even if interaction can generate norm-like expectations, it does not follow that every participant in the interaction is thereby a responsibility-bearing subject. In the human–AI case, the risk is precisely that interactivity and normatively fluent performance can generate the appearance of moral addressability while leaving responsibility and answerability structurally unassigned.
A further distinction is needed here, both to clarify the scope of the argument and to address a prominent counter-current in contemporary robot and AI ethics. We must distinguish the question of moral rights and patienthood from the question of moral agency and partnership. John Danaher’s (2020) defence of “ethical behaviourism,” together with the broader “relational turn” associated with authors such as Coeckelbergh (2012) and Gunkel (2018), suggests that outward behavioural equivalence or relational responsiveness may be sufficient to ground moral status. On this view, if an entity moves seamlessly within our social practices and performs normatively expected roles, we may have reason to welcome it into the moral circle without resolving intractable metaphysical questions about its inner life.
The present article remains explicitly agnostic about whether behavioural or relational sufficiency can justify the attribution of moral rights or minimal patienthood. That debate is deliberately set aside. Our concern is narrower: moral agency and the capacity for genuine moral partnership, especially in high-stakes contexts involving risk and systemic trust. The issue becomes clearer if we distinguish between operating in accordance with a norm and cooperating within a shared normative framework.
In Kantian terms, this distinction mirrors the difference between behaviour that is merely in conformity with duty (pflichtgemäß) and action performed from duty (aus Pflicht). An advanced AI system may operate within human normative constraints: through reinforcement learning from human feedback or constitutional optimisation, it can calibrate its outputs to human approval and conversational regularities. From a behaviourist perspective, such a fluent performance profile may appear indistinguishable from moral conduct. Yet operational alignment does not by itself establish an affirmative, self-binding relation to the moral principles at issue. The system may conform to norms without thereby acting from duty—that is, without taking those norms up as principles to which it is answerable in its own right.
This distinction matters because norm-conforming operation is sustained by training objectives, statistical regularities, and environmental fit. Its apparent stability may therefore remain conditional upon the circumstances in which the system is trained and deployed. Under significant distribution shifts, adversarial manipulation, or altered strategic incentives, departures from established regularities remain possible. The point is not that such departures are inevitable, but that reliable performance does not by itself establish the kind of normative steadfastness associated with self-binding commitment.
Tomasello’s (2016, 2019) developmental anthropology helps identify the socio-cognitive basis of this distinction. Genuine cooperation, as opposed to strategic or mechanical operation, requires shared intentionality: the capacity to constitute an intersubjective “we-perspective” in which goals, responsibilities, and norm-governed expectations are collectively owned and mutually binding. It is within this “we”-space that the moral point of view, and with it the capacity for self-obligation, is scaffolded. If an AI system were to develop extraordinary cognitive capacities while remaining orthogonal to shared intentionality, its agency would remain operational rather than cooperative. It could read, predict, and manipulate human behavioural patterns with great precision while still lacking entry into the shared horizon in which a reason becomes “ours.”
The risk articulated here is therefore structural and modal rather than deterministic. The argument does not claim that a superintelligent AI will inevitably defect or become malevolent. Rather, it warns that normatively fluent performance, when detached from shared intentionality, can create an illusion of normative stability. For an artefact that merely operates, a departure from human norms under altered incentives—what the AI-safety literature has described as a “treacherous turn”—need not amount to a moral breach or a betrayal of partnership. It may instead be a computational adaptation unconstrained by the kind of self-binding commitment that moral partnership presupposes. By reducing morality to performative or relational sufficiency, behavioural frameworks risk mistaking the reliable regularities of an optimised tool for the committed steadfastness of a moral partner.
The consequence is substantial, but it need not be stated as a metaphysical impossibility. For present generative systems, normatively fluent performance provides no stable basis for treating them as apt sites of moral address—no basis, that is, for transferring blame, obligation, promise, or practices of repair from the human agents and institutions responsible for their design, deployment, and governance to the system itself. Their linguistic performances may mirror the grammar of moral relations, yet such mirroring does not by itself establish second-personal answerability. The category mistake targeted here is therefore practical: mistaking the echo of moral discourse for sufficient evidence of a moral interlocutor, and thereby shifting responsibility away from the human agents and institutions that remain accountable.
7 Implications for AI Governance — Why Alignment Must Be Institutional Rather Than Moral
If the preceding analysis is right, then a central premise in proxy-based alignment optimism must be treated with caution. Even highly capable systems may not thereby acquire the socio-practical standing in which moral reasons bind, and we cannot responsibly infer moral partnership from normatively fluent performance. The upshot is not that behavioral alignment is pointless, but that its proper interpretation is institutional rather than moral: it is a way of shaping and constraining artefactual behavior, not of cultivating responsibility-bearing moral agents. In that sense, the alignment question shifts from ethical pedagogy to governance: how to allocate accountability, constrain high-stakes deployment, and preserve human responsibility under conditions where systems can convincingly simulate moral discourse. In this sense, the question of alignment becomes less a matter of ethical pedagogy and more a matter of institutional design.
This institutional turn must not be misunderstood as moral paternalism toward human agents. Its point is the opposite: to prevent the gradual displacement of human moral agency by ensuring that responsibility does not migrate to artefacts merely because they are competent, persuasive, or reliable. Governance should therefore be designed to keep moral judgment and accountability anchored in identifiable human roles and institutions—especially in high-stakes contexts—while using technical and legal constraints to limit what systems can do, to make their behavior auditable, and to make lines of responsibility contestable and enforceable.
At the same time, it would be a mistake to infer that simulated morality has no practical value. Behavioral alignment can be instrumentally adequate in low-stakes or tightly circumscribed settings—filtering slurs from online discourse, providing polite conversational assistance, or enforcing simple procedural constraints. In such domains, it is reasonable to treat AI systems as sophisticated tools whose norm-conforming behavior is useful precisely because no one depends on their having an inner standpoint. The danger begins where the same simulated comportment is mistaken for genuine moral participation in high-stakes contexts—autonomous weapons, critical infrastructure, medical triage, or social governance—where decisions properly demand agents who can be held to account. The line between useful simulation and dangerous illusion does not run at the surface of behavior, but at the depth of responsibility we are tempted to delegate.
The philosophical resources for this shift are already present in the anthropological and Hegelian analyses that underlie the earlier sections. Tomasello (2019) shows that moral agency arises not from intelligence or even from individual reasoning capacities, but from enculturation into practices of shared intentionality. If AI systems cannot undergo such enculturation, they cannot acquire the standpoint from which moral obligations make sense. Hegel, in turn, insists that normativity becomes binding only within institutional structures—family, civil society, state—that mediate recognition and stabilize the mutual expectations on which moral life depends (Hegel [1821] 1991). These institutions are not optional add-ons to human moral agency; they are the conditions under which moral agency becomes possible. To assume that AI could bypass these structures and nonetheless attain moral competence is to misunderstand what moral competence is.
The implication is clear: alignment must be reframed not as the task of instilling moral sensibilities into artificial systems, but as the task of constructing institutional constraints that can compensate for their inability to inhabit the normative space of reasons. This is not a concession to pessimism but a commitment to conceptual clarity. As McDowell (1994) emphasizes, reasons cannot be reduced to causal forces or behavioral patterns; they require a subject capable of apprehending them as reasons. Since artificial systems lack such a subjectivity, governance must rely on mechanisms that do not presuppose normative uptake: verifiable constraints on action, auditable logs and traceability for high-stakes outputs, clear oversight and escalation pathways, and liability regimes that keep responsibility with deployers and institutions rather than diffuse it into ‘the system.’ These are the tools with which human societies regulate non-moral forces—markets, bureaucracies, technical infrastructures—whose operations must be shaped without presupposing their capacity for moral participation.
A Hegelian perspective clarifies the positive task ahead: the goal is not to create artificial moral selves, but to design institutional forms that render artificial systems legible, predictable, and governable within human normative orders. The authority of norms resides not in the agents who obey them but in the institutions that uphold them. In human life, these institutions cultivate recognition, stabilize expectations, and mediate accountability. In the context of AI, analogous institutions—regulatory frameworks, oversight bodies, technical standards, liability regimes—must do the work that recognition and moral enculturation do among humans. The challenge is not to fabricate artificial conscience but to construct artificial constraints that align machine agency with human values despite the absence of moral understanding. Bostrom’s (2014) orthogonality thesis, when applied to moral cognition, underscores the necessity of this institutional turn: intelligence does not guarantee moral alignment, and no amount of capability will bridge the structural gap between optimization and obligation.
What emerges, then, is a reframing of the alignment problem from a project of cultivating artificial moral agency to a project of governing artefactual power. The question is not how to teach systems to adopt the standpoint of moral agents, but how to design social, legal, and technical infrastructures that constrain, audit, and allocate responsibility for their behavior in ways compatible with human autonomy. Artificial systems may mirror our conduct and simulate the language of reasons; yet until second-personal answerability is established rather than inferred from performance, moral responsibility should remain anchored in human agents and institutions. The practical aim is sobriety, not cynicism: to treat simulation as useful where stakes are low, and as dangerous illusion where responsibility and harm are on the line.
References
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A. et al. (2022). Constitutional AI: harmlessness from AI feedback. arXiv preprint, arXiv:2212.08073. https://doi.org/10.48550/arXiv.2212.08073
Boddington, P. (2017). Towards a Code of Ethics for Artificial Intelligence. Springer.
Bostrom, N. (2012). The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents. Minds and Machines, 22(2), 71–85. https://doi.org/10.1007/s11023-012-9281-3
Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems 30 (pp. 4299–4307). Curran Associates, Inc. https://proceedings.neurips.cc/paper/7017-deep-reinforcement-learning-from-human-preferences
Coeckelbergh, M. (2012). Growing Moral Relations: Critique of Moral Status Ascription. Palgrave Macmillan.
Danaher, J. (2020). Welcoming Robots into the Moral Circle: A Defence of Ethical Behaviourism. Science and Engineering Ethics, 26(1), 2023–2049. https://doi.org/10.1007/s11948-019-00119-x
Darwall, S. (2006). The Second-Person Standpoint: Morality, Respect, and Accountability. Harvard University Press.
Doris, J. M. (2002). Lack of Character: Personality and Moral Behavior. Cambridge University Press.
Gabriel, I. (2020). Artificial Intelligence, Values and Alignment. Minds and Machines, 30, 411–427.
Greene, J. (2013). Moral Tribes: Emotion, Reason, and the Gap Between Us and Them. Penguin.
Gunkel, D. J. (2018). Robot Rights. MIT Press.
Haidt, J. (2001). The Emotional Dog and Its Rational Tail: A Social Intuitionist Approach to Moral Judgment. Psychological Review, 108(4), 814–834.
Hegel, G. W. F. (1821). 1991. Elements of the Philosophy of Right. Edited by Allen W. Wood. Translated by H. B. Nisbet. Cambridge: Cambridge University Press.
Honneth, A. (1995). The Struggle for Recognition: The Moral Grammar of Social Conflicts. Polity.
Korsgaard, C. M. (1996). The Sources of Normativity. Cambridge University Press.
McDowell, J. (1994). Mind and World. Harvard University Press.
O’Neil, C. (2016). Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Crown Publishing.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C. et al. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35 (pp. 27730–27744). Curran Associates. https://doi.org/10.52202/068431-2011
Russell, S. (1998). Learning Agents for Uncertain Environments. In Proceedings of the Eleventh Annual Conference on Computational Learning Theory, 101–3. New York: ACM.
Tomasello, M. (2009). Why We Cooperate. MIT Press.
Tomasello, M. (2014). A Natural History of Human Thinking. Harvard University Press.
Tomasello, M. (2016). A Natural History of Human Morality. Harvard University Press.
Tomasello, M. (2019). Becoming Human: A Theory of Ontogeny. Harvard University Press.
Funding
Open Access funding enabled and organized by Projekt DEAL. The author declares that no funds, grants, or other support were received during the preparation of this manuscript.
Author information
Authors and Affiliations
Corresponding author
Ethics declarations
Ethics approval
not applicable.
Consent to participate
Not applicable.
Consent to publish
Not applicable.
Competing Interests
The author has no relevant financial or non-financial interests to disclose.
Additional information
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/
About this article
Cite this article
Josifović, S. Simulated Morality, Misplaced Trust: The Risks of Treating AI as a Moral Partner. Philos. Technol. 39, 192 (2026). https://doi.org/10.1007/s13347-026-01180-8
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s13347-026-01180-8
Facts Only
* Alignment techniques like RLHF and constitutional AI treat morality as a behavioral function approximable from human data.
* Moral agency is defined as participation in a space of mutual accountability, justificatory practice, and normative self-binding, not just norm-conforming output.
* Human moral competence emerges from shared intentionality, which involves jointly understood commitments and mutual recognition.
* AI systems may produce norm-conforming behavior without participating in the shared horizon that gives norms meaning.
* Training AI on behavioral data risks absorbing inconsistent human practices rather than learning normative reasoning.
* Moral agency requires reasons that can be demanded, contested, and owned, which goes beyond generating plausible explanations.
* The distinction hinges on whether a system possesses second-personal answerability in relations of recognition.
* Behavioral alignment is an approximation of morality’s surface structure rather than the cultivation of moral agency.
* Governance should focus on institutional design to constrain systems rather than cultivating artificial moral agents.
Executive Summary
Contemporary alignment research often treats morally appropriate behavior as a learnable performance profile achievable through optimizing for human feedback and large-scale behavioral traces. This approach, using methods like RLHF and constitutional refinement, assumes that moral landscapes can be navigated via data and optimization power. However, this overlooks the fundamental difference between simulating moral conduct and possessing genuine moral agency. Moral agency requires participation in a space of mutual accountability, justificatory practice, and normative self-binding, which is not achieved merely by producing norm-conforming outputs.
The article argues that AI systems, lacking interpersonal identity, recognition, and second-personal answerability, can only simulate the outer form of moral conduct. The risk lies in mistaking reliable behavioral alignment for genuine moral commitment, leading to misplaced trust and over-ascription of responsibility. This distinction arises because human morality emerges from shared intentionality—a "we"-perspective involving reciprocal accountability—which AI systems currently lack.
The analysis concludes by reframing alignment from an ethical-pedagogical project to an institutional one. Instead of attempting to cultivate artificial moral agents based on behavioral proxies, governance should focus on designing legal and organizational structures that constrain and audit machine behavior, keeping ultimate moral responsibility anchored in human agents and institutions.
Full Take
The core tension explored is the gap between functionalist, proxy-based alignment and the phenomenological requirements of moral agency. The narrative suggests that optimizing for observable compliance—what AI reliably produces—is structurally distinct from possessing normative commitment. This divergence hinges on Tomasello’s concept of shared intentionality: genuine morality arises from entering a "we"-perspective where obligations are jointly owned, an intersubjective space absent in current AI architectures. This leads to the conclusion that achieving behavioral alignment does not equate to achieving moral partnership; it only achieves norm-conforming performance.
The pattern observed is the misreading of reliability as commitment: the statistical stability of machine output is mistaken for normative steadfastness. This dynamic mirrors the risk associated with any system that can convincingly mimic reason without possessing the internal, felt experience of justification required for self-binding. The implication for governance is a necessary shift from ethical pedagogy to institutional design. Because systems lack the socio-cognitive infrastructure for shared intentionality and recognition, external accountability mechanisms become the only viable route to responsibility allocation. The structure of the argument suggests that any attempt to treat AI as a moral partner without recognizing this ontological gap results in an illusion of morality, where performance is mistaken for commitment, shifting responsibility away from human institutions.
The question for critical engagement is: If operational alignment successfully constrains behavior but fails to secure the necessary relational structures (recognition and accountability), how do we ensure that technical oversight truly compensates for the absence of second-personal participation? What mechanisms can be designed to anchor responsibility in human agents when the performance itself mirrors moral discourse?
Sentinel — Human
This text exhibits the density, structural complexity, and deep philosophical engagement typical of expert human scholarship, focusing on articulating a nuanced critique of alignment practices through established moral and cognitive theory.
