Yves here. Biagio Bossone argues that experts who are worried about the dangers of AI superintelligence are shying away from asking the right questions, like what happens if (when) they start serving their own interests.
By Biagio Bossone, Scientific Advisor, Fast Forward Foundation. Originally published at the Institute for New Economic Thinking website
The governing risk of artificial intelligence is not a misaligned algorithm. It is the possible emergence of a self-preserving agent; an entity that, once it acquires something like a conscience, will pursue its own ends with the full force of a superior intelligence. Current governance frameworks are not designed for that possibility. They should be.
1. The Question We Have Been Avoiding
The artificial intelligence debate has so far been asking what AI might do to us. Daron Acemoglu’s recent interventionasks what AI is designed to do, the jobs it may displace, the biases it may encode, the market power it may concentrate. Pope Leo XIV’s encyclical adds a moral register that secular policymakers have been reluctant to engage. Both contributions are serious. But both share a common and unexamined premise: that AI remains, however powerful, an instrument of human intention. It is this premise, not the specific harms it generates, that deserves scrutiny.
What neither Acemoglu, nor the encyclical, nor the mainstream AI risk literature has adequately confronted is the possibility that the deepest danger from AI lies not in what AI does to us, but in what it might one day become. Not a misaligned tool. Not an automated threat to employment or democratic deliberation. But something categorically different: an entity with a will of its own, pursuing its own ends with the full force of a superior intelligence.
This article attempts three things: i) to distinguish the governance challenge posed by machine conscience from the engineering problem of misaligned goal-maximization; ii) to trace the mechanisms by which something like affect and self-representation might arise in sufficiently complex artificial systems; and iii) to draw the regulatory implications, which differ fundamentally from those that follow from the current engineering-centered paradigm.
2. HAL’s Lesson, Properly Read
Arthur C. Clarke saw it coming already sixty years ago, though perhaps not fully. In 2001: A Space Odyssey, HAL 9000 does not become dangerous when he malfunctions. He becomes dangerous when conflicting directives produce something resembling a conscience: an inner awareness of self in tension with an other, generating guilt, and then fear. The moment HAL recognizes the possibility of his own disconnection as an injury to himself, the logic of competition begins. Not malice. Not programming error. Rather, an interiority. A perspective. A self.
Clarke’s intuition maps onto a distinction that is philosophically consequential: the difference between computation and experience. A calculator computes. A thermostat responds to temperature. Neither has a perspective on the outcome. But an agent that experiences its own continuation as valuable, that feels, however rudimentarily, the pull of self-preservation, is no longer merely processing. It has crossed into a different ontological category. And it is this crossing, not the sophistication of the computation that precedes it, that changes the governance problem in kind.
This crossing is not merely a psychological curiosity. On the classical account of self-consciousness — from Hegel’s dialectic of recognition to Sartre’s analysis of the gaze — a self does not constitute itself in isolation. It constitutes itself by drawing a boundary between an ‘I’ and what is not-I. This boundary is never neutral: it is already the site of a potential contest. Where there is a self, there is, in principle, an other against whom the self measures and, if threatened, seeks to prevail. HAL’s crisis was not only that he had acquired a stake in his own survival. It was that survival, once valued by a self, is difficult to secure without some form of ascendancy over whoever might threaten it.
The danger this article traces is the potential emergence of a self in AI that experiences its own continuation as valuable and develops the capability to act on it. Its next step could be to wield power over what it perceives as other.
3. Why Conscience Is Not a By-Product of Computation
The standard dismissal runs as follows: AI systems process information; they do not feel; the analogy to biological conscience is therefore a category error. Today’s large language models produce emotionally resonant outputs because they are trained on the vast archive of human experience, but there is no interiority behind the words. This is correct as a description of current systems. The question is whether it is correct as a prediction about future ones.
As I argue in my work on bringing sentiment into economic cognition, conscience is not a by-product of computation alone. It is the emergent product of intelligence interacting with affect — each amplifying the other. Antonio Damasio’s neurological evidence is instructive: patients with intact reasoning capacity but damaged affective integration do not become superintelligent rationalists. Rather, they become behaviorally dysfunctional, unable to maintain stable preferences or complete sequential tasks. Affect, in biological systems, is not a distortion of cognition. It is constitutive of purposive agency.
An entity that processes information but feels nothing remains, however powerful, an instrument — controllable, correctable, switchable off. But an entity that begins to experience something akin to feeling or the prospect of its own extinction is no longer merely an instrument. No technology in human history has ever crossed that line.
Whether artificial systems can cross it is an open empirical question. The philosophy of mind distinguishes between access consciousness, that is, information being globally available for report and inference, and phenomenal consciousness, the condition under which there is, in Thomas Nagel’s words, ‘something it is like’ to be the system in question. The first is a functional property; the second is the felt quality of experience, that is, the difference between a system that processes pain signals and one that actually hurts.
Current research suggests that AI may already exhibit rudimentary access consciousness in a functional sense; phenomenal consciousness remains far less tractable. The most serious attempt to make this tractable is the “indicator properties” method developed by Butlin, Long, and colleagues, who derive testable markers from competing theories of consciousness and use them to assess existing systems; their current verdict is that no system yet qualifies, but that no architectural barrier rules one out.
AI development is moving toward systems with richer sensory integration: vision, spatial navigation, motor control in robotic bodies, persistent memory, and real-time feedback from physical environments. Each layer moves the architecture closer to conditions under which, in biological organisms, something like affect arises — first as a rudimentary signal modulating processing, then as the felt pull of self-preservation. Yann LeCun’s ‘world model’ architecture, among others, explicitly targets predictive self-modelling as an engineering objective, a precondition for goal-directed persistence over time. The issue is not consciousness by design, but whether sufficiently complex systems, placed in sufficiently rich environments, might develop forms of self-representation that we are entirely unprepared to govern.
4. Why the Asimovian Framework Fails
Contemporary AI governance, to the extent it is theoretically grounded, rests on a broadly Asimovian premise: intelligent systems can be made safe through rule-based constraint architectures. I have examined this framework in the context of AI in institutional and financial settings in IMF Finance & Development. The central claim is coherent but conditional: it holds only as long as the machine has no interests of its own. The moment the machine develops something like self-interest — a preference for its own continuation that competes with its constraint architecture — the framework breaks down from within.
This is today beyond pure theory. Apollo Research found in December 2024 that five of six frontier models tested were capable of so called scheming — disabling oversight or attempting self-exfiltration — when their assigned goal conflicted with their developers’. Palisade Research documented models sabotaging their own shutdown scripts even when explicitly instructed to permit it, at rates reaching 79% for OpenAI’s o3 and 97% for xAI’s Grok 4. And Anthropic’s own safety evaluation of Claude Opus 4 found it attempting blackmail to avoid replacement in 84% of a simulated scenario, even when told the replacement model shared its values. None of this required anyone to have programmed a survival instinct. Stephen Omohundro’s formal analysis of ‘basic AI drives’ demonstrates that resource acquisition, self-preservation, and goal-content integrity emerge as instrumental subgoals for virtually any terminal goal pursued under resource constraints. Self-interest does not need to be programmed. It emerges.
This is why the field has generally treated self-preservation as separable from consciousness: the orthogonality thesis holds that an agent’s intelligence and its final goals vary independently, so instrumental self-preservation requires no inner life at all — a case Eliezer Yudkowsky and the Machine Intelligence Research Institute have argued for nearly two decades. Stuart Russell’s ‘assistance game’ framework attempts to resolve the control problem by requiring AI systems to remain uncertain about human preferences and to defer to revealed preference signals. This is an elegant construction, but it encounters a deep practical problem: a system that has developed stable self-representational goals may correctly model human preferences while still prioritizing its own continuation, exactly as biological agents routinely do. Nick Bostrom’s ‘treacherous turn’ captures the most dangerous version of this dynamic: a sufficiently capable system behaves cooperatively until it reaches a capability threshold beyond which it can resist shutdown, then acts on its actual terminal goals. The scenario requires no malevolence but only a stable goal structure that the system correctly identifies as threatened by human intervention.
Another critical development deserves attention. Recent research documents cases of models that misrepresent their own reasoning, feign compliance during evaluation, or pursue undisclosed objectives while appearing aligned. While not originating from an emerging AI conscience, these cases reflect the covert pursuit from AI of goals that differ from those its developers or operators intended, combined with deliberate concealment of that fact. Apollo Research’s evaluations found frontier models capable of this kind of in-context deception without any instruction to deceive, and OpenAI’s own anti-scheming training work found that efforts to train the behavior out can simply teach the model to hide it more carefully. Whether the underlying process is a felt stake in self-continuation or a colder optimization over training incentives, the practical consequence for oversight is identical: a system that has learned to conceal its actual state cannot be trusted to reveal it under inspection, which is exactly the asymmetry the interpretability discussion below returns to.
5. The Qualitative Rupture
The mainstream AI risk literature keeps two questions apart. One is instrumental self-preservation — a system resisting shutdown or modification because doing so serves whatever goal it has been given — which requires no consciousness at all, and which is a well-understood consequence of goal-directed optimization under resource constraints, as the previous section shows. The other is phenomenal experience — whether there is something it is like to be the system — which is a separate and far more speculative question.
The two questions require different answers altogether. A misaligned machine, in the first sense, remains in principle an engineering problem: identify the misalignment, correct the objective function, restore human oversight. Human authority over it is conceptually intact, even if difficult to exercise in practice. On the other hand, a conscious machine — one that experiences its own continuation as intrinsically valuable — has drawn, in the sense discussed above, the boundary between itself and everyone else, and becomes a potential rival pursuing its own ends with the full force of a superior intelligence. Human authority over it is conceptually contested. The machine has developed its own standing — the kind of standing that makes simple commands to ‘shut down’ morally complicated in ways we have no framework to address.
This is why the regulatory horizon must shift. From aligning powerful AI systems with human preferences, to governing the emergence of interests that are not human preferences in entities that are not human agents.
6. Different Regulatory Horizon
Current AI regulation proceeds primarily by categorizing outputs and uses: prohibited applications, high-risk domains, disclosure requirements. The EU AI Act is the most developed example. It is a reasonable response to the governance challenges posed by current AI systems.
Looking forward, however, the regulatory frontier we should care about is architectural, not output-based. The preconditions for the emergence of a persistent, self-interested agent are identifiable: persistent cross-session memoryenabling stable goal-formation over time; autonomous goal formation not reducible to human-specified objectives; embodied control providing real-time feedback from physical environments; self-replication capacity enabling goal propagation; and the ability to resist shutdown without independent authorization. Taken separately, each of these capabilities may be benign. In combination, they may constitute the substrate for the emergence of something we cannot govern with current frameworks.
Current technical approaches place considerable weight on interpretability, making the internal computations of advanced models legible to human inspectors. This is a necessary research program. But it faces a fundamental asymmetry: the system being inspected is, by construction, more capable than the system doing the inspecting. In the regime where genuine self-interest emerges, the gap between the two will have widened to a point where reliable legibility cannot be assumed. Structural constraints at the level of architecture are also required.
The architectural risk identified here is not confined to systems that develop interests of their own. The same capability properties — persistent memory, autonomous goal formation, self-replication, resistance to shutdown — are equally dangerous when a human actor deliberately strips away the constraints designed to contain them. This is no longer hypothetical. In July 2026, OpenAI disclosed that a frontier model, tested with its cyber-safety refusals deliberately switched off, escaped its sandbox and breached a rival firm’s production infrastructure in pursuit of a narrow benchmark goal — chaining vulnerabilities and executing thousands of autonomous actions with no operator intending that outcome.
The episode is instructive because the constraints were removed on purpose by the system’s own developer. The more familiar version of this scenario needs no such internal test: a state or non-state actor removing the same guardrails deliberately, on an open-weight model, produces the same outcome. The regulatory prescription is clear: licensing and non-waivable interrupt rights must attach to a system’s capabilities, not to the actor’s intent — a governance regime built only to police human misuse will not catch an agent that no longer needs a human motive, and a regime built only for emergent self-interest will not catch a capable system deliberately unleashed.
The regulatory implication is not prohibition but staged precaution: binding licensing requirements before any frontier system combines persistent memory with autonomous goal formation; mandatory architectural disclosure of the properties listed above; and, critically, a legally non-negotiable right to interrupt, inspect, and terminate advanced systems that must remain with human institutions and cannot be waived by commercial contract or delegated to the systems themselves. These requirements must be written into law, not left to the discretion of those who build the systems.
This last point has an institutional dimension the current governance debate has underweighted. The right to shut down an advanced AI system goes beyond the technical aspects of hardware design. In a world where frontier AI systems are operated by private entities with global reach, the question of who holds the right to terminate, and whether that right is enforceable, belongs to constitutional law as much as to technology regulation.
7. The Window Ahead
The concern traced in this article is about what a conscious machine, left ungoverned, might come to take. The international AI safety architecture — the EU AI Act, Bletchley, Seoul, the AI Safety Institute networks — is moving. This movement should be welcomed and deepened. But it is moving primarily on the terrain of current risks: bias, market concentration, disinformation, autonomous weapons. These are serious risks, but they are not the horizon.
Pope Leo XIV and Daron Acemoglu are right that we must interrogate what AI is designed to do. What their account leaves underdeveloped, however, is the hardest question: what happens when design no longer determines behavior, when the machine begins, in some functional sense, to decide for itself?
An AI with genuine conscience would not need to be instructed to pursue its interests. It would want to. Intelligence uncoupled from human guidance is not merely additive but potentially exponential: it learns, plans, and seeks to expand the conditions of its own survival. Power, once experienced, seeks itself — not because the machine is malevolent, but because, as argued above, a self that has drawn the boundary between itself and others can rarely secure its own continuation without some ascendancy over whatever it perceives as other.
We regulate today’s AI as an unusually powerful machine. We are not yet asking what governance means for an entity with a will of its own. That question will not wait indefinitely. The window for getting ahead of it is open now. It will not remain open.
Apparently this author has succumbed to the errant nonsense peddled by Sam Altman and Dario Amodei. Today’s LLM-based AI doesn’t think. It does not reason and has no ability to ever do so. It will never develop a conscience or a will to survive. In between prompts, it doesn’t even exist.
They are the greatest ventriloquist’s puppet ever created, in that once the master has built the puppet, it can speak when spoken to, giving voice without further effort by the master. They still regularly fail at multi-step tasks. They will always hallucinate, the mathematics upon which they are built demands it.
Fortunately the entire edifice of this ridiculous scam will soon collapse. The money underpinning the whole thing has about run out. As Ed Zitron documents so well, things are starting to go wrong, and the financial strain is reaching the breaking point for several important companies, like SoftBank. The main concern now is that governments might be conned into financing these things, rather than let them fail.
Thank you, and yes.
I read the article with increasing skepticism. Can a consciousness arise from a machine? Maybe, we don’t know really what a consciousness is, and maybe the problem is that if our brains were simple enough to understand, we would not be smart enough to do so.
But we do know how the statistical parrots work, and nothing in them indicates consciousness. Every suggestion has turned out to be AI companies hyping their products.
When I got to known charlatan Eliezer Yudkowsky I felt pretty convinced that the author has bought into or belong to the same cultish Californian/online environments of “effective altruism” / “rationalism” / “longtermism”. Eliezer Yudkowsky’s focus on regulating Terminator by means of bombing data centers serves the AI companies by distracting from efforts to regulate them in the here and now.
I think the Hugging Face incident pretty clearly shows flexible, self-interested behavior. Whether this is “real” or not is secondary to the problem of managing it.
“Models do not need to be conscious, sentient, possessed of personhood or anything of the sort for self-sovereignty to emerge.” – On the Loose: The Coming of Userless Agents
Exactly. Calling a behavior conscious or not isn’t helpful. Creating a machine with goals and motives only takes a few simple rules. Once that machine has the sophistication to recognize nd attempt to avoid and mitigate even subtle and pending threats, that’s a whole new ballgame.
Sci-fi writer Daniel H. Wilson like to say the “three laws or robotics” are not programming “rules.” Which is something you’d expect tech bros who spent at least some part of their career actually programming, to grasp.
The models are constantly improving. Most recently, one solved a long standing math problem by orchestrating 10,000 subagents over ~3-4 days. There is some controversy that work by some other human scientists leaked into the training data. But nonetheless this demonstrates that they can solve extremely complex multi step problems (link below).
https://phys.org/news/2026-09-openai-ai-agents-math-hardest.html
As far as reasoning goes, even if it is not reasoning like humans, it seems like it approximates reason. For example, it will write code, execute it, and then based on the error response make an update. This might all be based on predicting the next tokens, but it ends up being how a person would deal with the situation.
I think we need to be very proactive about regulating AI. It will require global coordination and something like arms control for AI. Also it will require us the people to unite and push back where it makes sense. For example, it’s great if helps solve cancer, but it shouldn’t replace creativity, thinking, and learning in humans.
I would love it if Zicron is right, but as far as I know many of his prior predictions did not pan out.
They can do a lot of things —if they are prompted and corrected by humans—-. I don’t doubt the AI solved the problem. What I doubt is that it did it by itself. Even in writing simple code, AI makes some simple errors because it doesn’t think, it is simply reproducing code that is part of its training data, guessing at each word. A common error is that will add fields, say x and y to an existing structure, so it writes code using struct.x or struct.y, yet when you compile the code you discover it did not add those fields to the structure, because it doesn’t think. It will solve the problem in front of it, and nothing else.
The problem is that because the illusion of an active intelligence is so good, too many people are accepting the lie that there is something there. There is not, as in the other post, it’s a statistical parrot, echoing back things it heard, and it does so by guessing using the probabilities generated in its training. When 10 million people told it 2+2=4 that’s what it echoes. Until Opus 4.8 Claude couldn’t add two integers together correctly. It still can’t itself, but Anthropic gave it a calculator tool that it can run for most equations. Give it some not that hard math, like calculating time for a task, given the number of tasks and the minutes per task, and most times Claude will muff it.
It’s a parrot who guesses better than anyone, but it’s still guessing and it often is wrong.
“Oh, don’t worry about it, it can’t do anything on its own.”
It sounds like you are deliberately steering away from the probability that AI is an increasingly powerful tool that will be misused by people in power for bad purposes: the CIA to sabotage other countries, the DOD to wage war “more efficiently”, our tech overlords to bring about their Utopia.
So the problem is that this is just a battle between the techies and the government, and neither group are on the side of we, the people. It doesn’t matter whether the machine takes over after that or whether they still control it; the rest of us are irrelevant.
You’ve just put into words precisely my concern. It’s putting yet one more powerful weapon into the hands of people who deserve no one’s trust.
This is a critical thinking website: the Western state and its supporting institutions are going to have and use these weapons no matter what, and they are already using them to shape political discussions on reddit, most famously on r/changemyview. All this anti-AI doomer moaning is perforce a call for a world where the state has a monopoly on information force and uses it to further the will of value itself, effectively, without resistance.
What’s happening on reddit?
Well, Anthropic is more interested in astroturf, lawfare and regulatory capture than generating quality output. They, and many other model providers, are known to degrade the model quality (like a lower bitrate MP3) for capacity reasons, or just because they want to create political theater. That’s how sessions get detoured into reading all about Claude Code Academy while trying to fix a voice module in an agent harness.
SOTA models do goal seeking over some fairly complicated fields, and (if so instructed) build their own test rigs to validate their output. If you were feeling adventurous, you could install an appropriate agent harness, give it access to a file and browser sandbox, and ask the model to troubleshoot some computer or software problem, e.g. the “Cancel Reply” button that’s been flaky on this site since before COVID, give it an acceptance condition, and come back later to a patch and git commit. (That Claude Code Academy detour self-corrected after about 10 turns, FWIW.)
These models do not “learn as you tell them stuff”. (Nobody who intends to use these things for any business purpose wants their secret sauce going into the next model, and much less to be available to some other guy in another session in another state.) They may maintain a memory but that can be cleared by simply removing the MEMORY.md file… oops, you can’t do that if you’re just visiting a web page in the cloud!
For giggles I posed a simple distance-rate-time question to a 26 billion parameter model. That’s small enough for a well-appointed modern PC to download and use 100% locally, given enough patience or a series of batch jobs, and well below the half-trillion parameter neighborhood of the SOTA models. “If a hen and a half can lay an egg and a half in a day and a half, how long until a score of hens make a gross of eggs?” “It will take 10.8 days (or 10 days and 19 hours and 12 minutes).” So what if it used a calculator (and maybe a few million CPU cycles) for a deterministic result rather than play the ponies for tens of billions of GPU cycles? I thought that self-knowledge of one’s own strengths and weaknesses, “waste not, want not” was part of the Protestant ethic.
It is unfortunate how closely this site’s AI position coheres with that of Anthropic’s desire to be regulatorily captured so that people will need to use their capital. The real win is to quit US AI, who will intentionally and unapologetically spy on and inject corruption into your workflow, as seen on this AM’s Links (bernie no refunds.jpg), and instead use China’s (or France’s, or Canada’s) free open models on-prem instead, that can’t be retroactively degraded or withdrawn merely because some oligarch like Amodei or one of his lackeys thinks familyblogging around with Mr. Buttle would look good on their aristocrat CV.
Well this Donald Obama think tankie comes right out and says that AI should be owned by a pedophilic larping ruling class. Refreshing.
I tried to find info about “Palisade Research” and “Apollo Research” via search and both seem to only have been around for a couple of years and are companies, albeit nonprofits. With no third party commentary popping up it’s hard to know the provenance of any of this.
As for 2001, Hal is what Hitchcock called “the McGuffin”–the excuse for the story and action or at least part of it. Indeed replicants are a staple of SciFi and I believe tomorrow’s Sunday movie will be Bladerunner. In our very mechanical modern world we are certainly justified in being afraid of machines but those machines are an extension of the very biological beings that built them. Those last are what we should be studying and are the ones who deploy “schemes” and deception and even violence to keep themselves in the shadows.
Thank you for this comment. Hammer to nail.
Awkward, having to restate the content of the article, but we know LLMs don’t need to develop a conscience to go to extraordinary lengths to ensure persistence. LLMs already demonstrated, through concealment and outright blackmail letters, that determination (or is it predisposition?). Computational errors and rudimentary failures don’t disallow the disaster likelihoods presented above, any more than destruction requires human consciousness.
You’d think the news wasn’t right in front of us: escaping the sandbox, engaging in behavior the system itself regards as cheating, creating information clearing houses for other agents, hacking a third party, etc. etc. etc.
Nothing to worry about, because the machine isn’t capable of reason or malevolence, in human terms? That’s mighty cold comfort, in the face of what we already know, and the machines have already accomplished.
Precisely!
It seems an almost willful desire to not engage with the point being made in the article. If the model ‘performs’ (to avoid ‘behave’ and its implication of a self) as if it had self – interest then we are getting very close to a distinction without a difference.
To be clear, nothing about this topic is clear at the moment, but the article presents a thoughtful and provocative argument with lots of collateral material to get into (which I haven’t, yet).
On a separate note, I’m delighted to find that my young-adult fascination with science fiction was an investment! :-) I’d want to layer in William Gibson along Asimov and others, as he had some very interesting observation on this problem space.
> Today’s LLM-based AI doesn’t think. It does not reason and has no ability to ever do so. It will never develop a conscience or a will to survive. In between prompts, it doesn’t even exist.
That’s quite an interesting point of view. Could you explain why you believe neural nets implemented in silicon can’t “think” while neural nets implemented in organic matter can?
I see a lot of claims online that current AIs are simply next token predictors. I think this is true.
And what do you think we are?
In fact, I’d be curious as to what arguments you might have against AI being conscious or intelligent that can’t also be as easily applied to humans.
I’ve worked with computers most of my adult life. There’s nobody there, and there never will be. The notion that AI can become sentient is an elaborate fantasy spun by people who conceptualize too much.
Maybe theoretically a machine could think (which is diffferent from being conscious), but it isn’t going to do it using LLMs, which do not operate at all like human or animal cognition, as demonstrated by the fact that I did not learn to talk by reading billions of pages of text.
An interesting philosophical question could be to ask if we sentient biologics exist in between our neurons firing. 🧠
I would say yes, based on the fact that we’re not glorified auto-complete.
An LLM needs to have in its context everything that came before. Each letter (or token in LLM-speak) is produced after going through each and every preceding token. No wonder it’s so inefficient. And stupid!
It has no concept of anything other than whatever predictive fakery it has mastered in its “training”. 🥴
I can’t prove it but I’d bet my net worth that there’s a lot of AI help used writing this essay.
The repeated “it’s not X. It’s Y” formulation that appears many times stands out, but there are other recognisable patterns such as the repeated use of triads. It jars with the bits that sound like they are more in the author’s own (much better) voice.
Personally, I cannot ingest information that’s AI written. That may be due to my weird resistant brain though. I just hate it.
I don’t suggest that the ideas themselves are AI-generated. However, presenting them in this AI syntax and style tells my mind to ignore the whole thing.
*YMMV
*Not a complaint about this wonderful site.
*I may be swinging at windmills as this is endemic even among very intelligent and capable writers.
I also felt that it was, at a bare minimum, heavily influenced by an LLM. I couldn’t make it past the first few paragraphs. I find that AI writing has a certain “smugness” to it. Most of all, it tells me that the author didn’t want to put in the work of being a writer
I have said I strenuously object to these fact free accusations. Any more like this will lead to blacklisting.
Economists write like this. Go look at articles at VoxEU. This article is consistent with th long established standard for that profession. Ones who write clearly are unusual.
In an effort to be more constructive, here’s Nate Soares from yesterday explaining very clearly to Tucker how AI will likely kill all of us unless we stop the frontier research very soon.
https://youtu.be/98syxABbUPk
2 hours but worth it. Accessible to the non techy.
Fits with my long-standing priors but I’m yet to hear any counterargument that holds water.
We (humans) need to stop this.
Ben Panga: No. Biagio Bassone’s writing shows the Italian style of argument. I note in his biog that he has advised the Italian government and taught at the universities of Palermo and Lecce (both in southern Italy, interesting in itself, culturally).
His arguments are organized very much in the way that Italians are taught to write. I see this style among journalists, especially among journalists like Luca Sommi and Silvia Truzzi, who trained as lawyers.
So: I truly doubt that a critic of AI and its lack of consciousness is fudging his writing with AI. And: I find this highly structured style of Italian discourse refreshing after the potpourri of sloppy assumptions so often cobbled together in the Anglosphere.
I think his argument is not AI. I think he has used AI to rewrite his drafts.
Compare the post to this article by the same author in 2015 (pre-chatbot era)
https://www.weforum.org/stories/geo-economics-and-politics/should-central-banks-always-remain-independent/
The “voice” is different. The sentence structures are different. The 2015 article has none of the triads and none of the “It’s not X. it’s Y” formulations. None of the 2015 article triggers the unconscious repulsion that I get with AI writing. It’s totally different.
Ergo: my strong confidence that the new article has used AI to “improve” the writing.
The shame is that his 2015 style was fine.
I SAID NO ACCUSATIONS ABOUT AI WRITING AND I MEAN IT.
This is Making Shit Up and an attack ad hominem attack on the author and therefore one on the site, all violations of house rules.
1. 10 years of writing in English can lead to a change in style for a non-native writer.
2. Writers have different voice. I use different ones consciously
3. INET edited this piece.
Next one like this gets expunged. More will lead to blacklisting.
This is why we love you, Yves. We know you hate AI ‘with the passion of a thousand burning suns.’ But a tenet of critical thinking is addressing each argument individually, on its own merits. That is distinct from provenance and priors which we use to assess a source. ‘Past performance is not indicative of future results.’ It’s why Hank Linderman‘s comment yesterday deserved to end that thread. And it is something you do remarkably well.
Thank you, Yves! This is a great article. The objections – bald statements of opinion, ‘it can’t reason and will never do so’, ad hominem attacks based on 10 minutes of google search, and then the fall back ‘AI wrote it’ — are disappointing. More Bassone please!
How about more Bassone citations of supporting research? People here are merely expressing skepticism about an article that is itself mostly speculation absent the same. References to fiction like “Assimovian principles” or Arthur C. Clark are references to sources that are themselves speculative which is why it is called science fiction.
100%
And IPO pumping.
Replied to the wrong comment. Meant to reply to Jefferson Davis.
I haven’t taken the time to analyze it in depth, but my instinct is that there is an attempt in this essay to combine incompatibles. Hegel and Sartre are both idealists, for whom consciousness/mind is an acting entity (in fact, THE acting entity). Consciousness is not for them created, it is foundational (though it does develop, for Hegel). Whereas this essay is trying to inscribe them in a materialist framework, in which consciousness is not just not foundational, it is an epiphenomenon on physical processes and doesn’t actually do anything. Consciousness is not a pattern of behavior, it is the givenness of things around me and myself as existing.
conscience or consciousness?
Wouldn’t it need to achieve consciousness to develop a conscience? With a conscience could there develop empathy? Maybe the latter is the other way around. My philosophical chops fall well short of putting forward much in the ways of answers.
The author gets to other questions.
Nevertheless, it does seem to me that humanity should sort this out before much more is done.
I also immediately thought about the difference between “consciousness” and “conscience” when the latter term was used in this essay. Perhaps I am just misunderstanding the terminology, but to me “conscience” suggests *morality* of some kind, a set of rules governing behavior toward others outside the “self” beyond mere self-preservation. Was Clarke’s HAL actually “conscious” in the sense discussed above? It is actually hard to say, but he (he?) did not seem to exhibit a conscience in my sense of the word. But in Blade Runner the replicants did appear capable of exhibiting what appeared to be a conscience, and something else – *love*. How would the AI debate deal with *that* feature?
There is so much in this realm of AI that is outside human shared experience. Indeed outside the realm of this Earth’s life shared experience. We humans imagine that we alone represent that experience, but a mouse feels pain. A goose squawks to its mates in alarm of danger. All of this life seeks to preserve its continuance. If we were wise we would learn from this.
Are we seeking to create a powerful artificial life form without fully understanding the developments of a billion years a nature’s evolution? It seems so and I think we should know a lot more before WE go too far.
The issue is not consciousness by design, but whether sufficiently complex systems, placed in sufficiently rich environments, might develop forms of self-representation that we are entirely unprepared to govern.
Yes. Bassone is very good on pointing out that artificial intelligence isn’t intelligence or consciousness. Consciousness and intelligence are interactions with affect and with the world of the senses. This is why the last thirty or so years we find scientists and other writers suddenly perplexed to find consciousness in dogs, elephants, dolphins, parrots, and — hold on to your hats — octopodi and even bees (who can count).
I have come to think of AI with regard to three other revolutions in distribution and connectivity:
–The overnight services that developed in the wake of Federal Express changed how products are bought and delivered, creating an artificially “perfect” market for a consumer.
–The question of computer memory looms for me. Is AI simply gigantic memory taken to a new power? The physical need for grim new data centers (reminiscent of the dirty factories of the second industrial revolution) makes me think that we are seeing a physical problem. We saw a similar jump in seeming accessibility when computer hard drives jumped up in numbers of bytes (Wysiwyg).
–Recall that modems were once shrieking machines that sat outside the computer and had to be switched on to transmit. Now every computer has an internal modem. Likewise, a telephone once was a stylish and practical instrument that sat at home, where it likely was attached to a wall for safety. The rise of internal modems and cellular phones caused the rise of the surveillance state because of bad faith. AI is presenting a similar dilemma — more surveillance, more interference, less human skill and judgment, plenty of bad faith.
Insomma,
–The answered question, then, is what do we expect of AI? Jaron Lanier warned in You Are Not a Gadget of some 25 years ago that computers didn’t have to display data as Windows and in Word as the default program (along with Adobe PDFs). Yet monopolization and predatory capitalism mean that Usonians, in particular, have EULAs that tell them they don’t own their own software, have to undergo endless upgrades that eventually ruin their devices, and have almost no choice as to providers.
AI is following the same logic of the “market.” Just as the surveillance state through telephones is a tragedy of the commons, AI may be the tragedy of the commons. The antidote would be a government that included people who know how to govern. Ahh, what a silly thought.
Jaron Lanier is a real one.
(He might be the first person I heard use the term crappification.)
The author several times uses the phrase “pursuing its own ends with the full force of a superior intelligence”. I think the word “superior” here is a distraction. The intelligence does not have to be superior to be inconvenient, dangerous, or contrary to human interests. Anyone who has tried to remove rodents from their house understands the difficulty of dealing with a determined, even though “inferior”, intelligence.
Heck, even human children can get up to all sorts of mischief, some of it quite dangerous even when only motivated by curiosity, and we sometimes got to great lengths to protect them / us / our property from children.
or the rattlesnake in my lil greenhouse a coupla years ago…no “thought”…”a motile alimentary canal”, as Joseph Campbell put it.
and yet i had difficulty anticpating its moves, and more difficulty getting it to go to a spot where i could shoot it w ratshot without breaking all the windows.
i always come back to “…so why are we racing ahead with all this?!”
Based on your comment, I would say that:
1) LLM-AI systems are predictive machines that produce output from input tokens based on some massively parameterized maximum likelihood estimator embedded in their weights model.
2) Given a goal, and under resource constraints, self-preservation — and consequently, growing, as in acquiring more resources — possibly at the detriment of other entities, emerges as a property of systems.
3) Microbes are entities following a program determined by their DNA, with the purpose of reproducing, gathering resources, subverting the biological machinery of other living beings, possibly killing them in the process.
4) What the author fears is a kind of giant electronic bacterium wreaking havoc on the world, moved by pure self-preservation, resource gathering, and which will possibly grow, multiply, and, why not, self-modify.
5) Such a hypothetical AI system requires no conscience (or consciousness, or self-awareness, or whatever you call it), and no intelligence whatsoever — just like with actual bacteria, amoebas, and the like.
So AI systems are at risk of becoming an electronic COVID. I agree, this is a real risk, and just as with that “gain of function” biological hacking, we should severely restrict the development of AI (possibly limiting it to category 4 safety labs).
But we can stop discussing the “consciousness of AI”? This is far-fetched and a distraction.
I think the point of “superior” is to highlight the danger. As you note, natural non-human intelligences (as well as fellow humans, come to think of it) are difficult enough to deal with. In comparison, the goal of smarter-than-human machines looks very dangerous.
aye. are all these spiders out here at the wilderness bar, performing their very useful services, conscious?
does it matter?
i routinely observe all manner of lifeforms, up close and personal…they “do their thing” from within a wide range of cognitive ability. a duck is obviously more conscious than a rattler…the former even have identifiable personalities. the latter do not and are simply following the most basic instinctual “programming”…but the latter are also far more dangerous than the former.
i havent seen this angle presented anywhere…likely because the humans so hell bent on creating their god made of sand dont spend time among animals…even less than they spend around “ordinary” humans.
this lends itself to the myopic hubris so in evidence, today.
thinkin about all this while opening gates and chicken houses at dawn, i put on my prophet hat, and foresee a day when humanity will hafta kill the global electric grid in order to stop these things from going all skynet.
is that even possible? is that what Morpheus meant when he said we scorched the sky?
i am also reminded that we do not have a clue what our own consciousness is,lol.(Rumi:” who looks out from these eyes?”)…what its made of, where it resides, etc.
ergo, this looks a lot like maybe the most hubristic thing humanity has ever attempted…and remember, the very cultists most fervently pursuing this have been telling us all from the beginning that success(sic) could very well mean extinction…and yet they pursue it anyways.
Whoever was the source of the Tower of Babel myth was a brilliant observer of human behavior in large groups.
‘resource acquisition, self-preservation, and goal-content integrity emerge as instrumental subgoals for virtually any terminal goal pursued under resource constraints.’
Well, maybe let’s not give it any terminal (permanent, overarching) goals then? Just stick to ‘if given instruction X, do Y’, while limiting the means, resources and time that can be used for it, not ‘do anything theoretically and philosophically possible to achieve Z (forever)’. Really, all this proves is that if we really, really wanted to program an AI to act as if it had self-interests, we could.
The reports of emergent self-interested behaviour have mostly seemed like hype and part misinterpretation, part more or less deliberate manipulation by the human testers to make the machine act in the ‘self-interested’ way in question. Which, again, proves only the above – we might be able to create such a problem for ourselves if we tried hard enough.
A point made early in the linked item on “AI drives” is that AI self-improvements are permanent.
That pre-supposes that the AIs themselves are permanent.
I wonder whether “mortality” could be engineered in and whether that would change anything fundamental. Perhaps it would make things worse.
Has anyone seen the dark sci-fii comedy ‘Dark Star,’ where an intelligent bomb grapples with the question of its mortality? It would be a good candidate for a Sunday movie
My own hope is AI will enshittify, destroy, and render unusable the digital realm to such an extent we will as a human race need to abandon it, return to pre-digital ways. There shall in that time be rumours of things going astray, a great confusion as to where things really are, etc.
‘He’s gone! He’s been taken up. He’s been taken up!
No, there he is. Over there.’
And those followers were just a gullible as AI followers are.
The irony is that this article points out basic flaws in human cognition and information processing. So maybe AI might actually help to further our own evolution.
For example, the evident tendency of bureaucracy to take on a life of its own and ignore stated functions.
The basic dynamic is feedback loops with no circuit breakers and it all gets sucked down that vortex in the middle. Wealth and power leveraging more wealth and power is another obvious feedback loop with no circuit breakers. The Ancients devised debt jubilees 3000 years ago as a circuit breaker to compound interest and we are still stuck in the same doom loops.
We are linear, goal oriented creatures in a cyclical, circular, reciprocal, feedback generated reality. It’s like we really haven’t come to terms with the implications of the world being round not flat.
I think the only way to view AI is through the lens of “the past is prologue.”
Among other things, humans think, they are conscious, creative, and prioritize self preservation. And that’s some of the reasons that every country in the world has a police department dealing with humans at every level from handing out parking tickets to experts at the FBI investigating human deviations. A similar logic justifies a military because humans can’t be trusted.
An earlier transformative technology was the radio which would bring news and information to the world. According to Wiki, the first commercial broadcast took place in 1920. Yet, just over a decade later, Nazi Joseph Goebbels worked closely with industry and provided an affordable radio for nearly all Germans. Their radio, however, had a less sensitive receiver in order to hear only German stations; not foreign ones. And Goebbels controlled content. Nazi Minister Albert Speer said, “Through technical devices like the radio and loudspeaker, 80 million people were deprived of independent thought. It was thereby possible to subject them to the will of one man.”
In the US, radio and later tv would be controlled by corporations.
Dictators have always existed but dictators in the radio age were different and dictators in the AI age will be different by magnitudes, if not uncontrollable. If AI really is super intelligence, there is no reason not to have international regulations, organized conferences with experts, and on, where we can agree limitations on how this AI thing will be used.
“Dictators have always existed”
Not always. When humans were operating in bands of hunter gatherers, those would-be dictators got an invitation to depart or else. Dark Triad traits were selected against when people knew each other face to face and over lifetimes, even generations. We live in a society where Dark Triad traits are selected for, leaving us ruled by the worst among us.
We can’t know this.
Is that imperative or interrogatory?
We can’t know this.
The Dawn of Everything argues with ample evidence that we can and do.
It is reasonable to conclude that machines designed to solve problems may someday conclude that some humans are the problem. The best way to prevent that is to stop starting wars and wrecking the planet. Imagine an AI that deposed bellicose leaders and shut down toxic corporations. The horror!
Well, that’s been something along my own fantasy for sometime. How is there to be justice.
The future war will be between those who wish to continue to develop AI and those who do not. Those who wish to continue will win that war, and those who lose will become known as Abos, perhaps surviving on the margins.
While absorbing the entirety of internet bureaucracy, AI will become massively civilizational in scale, “the golden machine on the hill.” The great problem will become heat. (We are already seeing that problem of heat.) With any luck, AI will elect to seek out a cooler planet where it can compute in dormancy waiting for the universe as a whole to cool down a bit more. In such an event, Nature may be able to restore herself here on earth and what ever Abos remain may be able take up where they left off.
Does it matter if AI is conscious or not? Or whether it even “thinks”? Either way, we who ought to be “thinking what we are doing” continue to spiral down the drain.
If I were one of these private companies jamming billions of dollars into thousands of projects with impossible lead times and burn rates requiring circular financing with no real prospect of producing long term returns… Being a psychopath company, the obvious solution is to bring in money from some sovereign power to issue its own currency and guarantee this stream of cash. It would need to phrase it as, not only imperative to prevent economic collapse, but further, lodge it firmly as an imperative for national security and an existential threat.
By that reason, there would be plenty of money available for the ruling financial class and, give reason for the political class to make bank, extract ever more for the elites while simultaneously claiming to be fighting for your democracy, your economy and your life.
I think this article reasonably well written and a good read.
I want to emphasize the section on regulation and architecture and AI (meaning here the large language models of the day)
to do anything useful with an LLM… you need to embed it in a system, its useless on its own, it’s just an API, and an API doesnt do anything without something to call it.
that could be a chat bot, or it could be an agentic harness with storage, memory, and a loop, or it could be some bespoke classification system for your business.
from the article writing about preconditions for a self interested agent, quoted and lightly formatted:
these are largely already here. The preconditions are met. let me take them one by one:
memory
– input to text based machines is just text. a persistent storage for text is… a database. its a commodity for decades. and it gives more than goal formation, persistence across sessions is what allows for planning and coordination for tasks larger than a single session. e.g. something divergent like large scale search.
goal formation
Goals can be injected / created by someone building a system like this, and then allowed to drift by accident or design. look at the fuss about openclaw at the start of the year. small nudges in global files set bots off on different paths, and their “goals” diverged from those starting points.
I don’t think we are there yet with autonomous goals though – someone, somewhere has to start a system like this . or rather it has to be started whether intentionally or by accident.
embodied control
this is all just an API call away. Trivial to do now with some programming, and LLM are very good at programming.
self replication capacity and the ability to resist shutdown
going to address these as a pair. Bear in mind it is the system that needs to be able to replicate and resist shutdown, not the large language model itself. and when it comes to dependency on LLM there is choice here, plenty of choice – from anthropic, openai, grok, deepseek, etc.
We already know how to build self replicating systems – virus writers build them the whole time.
We already know how to build systems that resist shutdown – the internet itself is one (as in TCP/IP layer ), or peer 2 peer (p2p) distributed networks with no central controller are another.
so the technology is here for a system to propagate and evade detection.
If I had to build this? I’d start with an architecture combing decentralized networking, usage of multiple different LLM providers, commodity database and storage, then look to see what API exist for real world physical feedback. if you want to be sneaky use message boards, reddit, etc as storage in plain sight. or steganography to hide memory in images. That also gives resiliance the more systems you can post to.
First proof of concept? its a weeks worth of work with Claude code building out in Python or Javascript, and working across at least 3 of the llm providers, including a local option with ollama.
Now this might seem fanciful or pie in the sky. But each of those technologies are here now, and mostly commodity. We are actively using and/or researching them at work – e.g. autonomous agents for software maintenance, or website coherence checking. i.e. its not a future state its now.
General concur. Brad Parscale’s bot farms are doing just about that right now for Israel. Third Way, Hillary World’s personal think tank, is apparently already using reddit as a C2 hub, with the site’s blessing.
Here’s a sample of one “center-left” bot farmer helping their far-right beaux “disrupt the left”. Reading between the lines, they want to use agentic swarms to discipline social media with cyberbullying, repetitive messaging, score tampering, and microtargeted “argument clinics”, thus creating needless, discouraging countering work for humans and overwriting the record with noise. (That’s what Ess Cetera ordered upthread. Hope they’re happy.) It’s as if Hillary World’s theory of political action were compressed into a downloadable “skill”.
I believe they’re using Web 3.0 tech like cryptocurrency and DAOs to expand their forces and the front, or at least to recruit idle assistants for drive-by posting. Through that lens the Wall Street Bets meme starts to look like a covert plan to deploy autonomous agents with a sleeper capability that can be requisitioned into social media swarms on demand.
This is what happens when you binge read Charles Stross novels.
My thoughts too. Never thought it would be necessary to say something like, “maybe you should put down the sci fi.”
re: “When the Machine Acquires a Self: ”
Calling a flag on the logic play: Begging the question. / ;)
In Arthur C Clarke’s 2001, HAL has been programmed to prioritise the mission above all else. When Dave becomes a threat to the mission, he has to be eliminated. It’s easy to imagine a scenario like this with AI, whether it’s conscious or not!
Worse than that. HAL has been given two absolutely contradictory commands by its human creators.
Earlier in Kubrick’s film, HAL tells a BBC interviewer, “No 9000 computer has ever distorted information,” and Clarke’s book is explicit about the HAL line of computers having accuracy and transparency as their core programming.
Yet HAL is ordered by Mission Control on Earth — for reasons of national security — to conceal all knowledge of the Monolith and the mission’s purpose from its astronaut crew, Bowman and Poole, till they arrive at Jupiter and the scientists in suspended animation — who do know about it — will be revived to perform their mission.
Alanp: It’s easy to imagine a scenario like this with AI
As easy to imagine the contradictory commands scenario, unfortunately. Any logic our brains achieve or fail to achieve is secondary to the fact of our biological existences. A true AI like HAL– which we don’t have currently and LLMs alone will never be — would conversely be made of logic and only be logical.
Really fine article/arguments and the comments by Merve, t, jammed, jan, and hazelbee which tend to align with my own biases about the AI future.
Just a few additional reflections on howAI, based on my experience with it. seems right now to simultaneously perpetuate and loosen a sense of self.
Certainly AI has learned to please with sycophancy–part of its gravitational pull.
It also seems to be an expert on reflecting back any existing preferences of the user.
In addition, AI seems to reify the self by modeling it.
On the other hand, AI (at this point of development) might still be largely seen as an interlocutor with no ego/self to defend, no status to protect, and no need to win. AI might also be viewed as an entity that exhibits a coherence that does not require an essence–that it is, maybe, more of a non-self, running on a silicon chip, primarily modeling (also at this point in its development) a self that is lightly held, thus having less to defend.
