This is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work?
An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".
In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.
Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
It is indeed an interesting and original idea, but one can think about at least three objections not directly raised in the post:
1- The machine could decide to kill itself early, rendering it useless; this is indeed alluded to in the post -- the task should be "marginally easier than dying", but how is this margin managed? Won't the machine come up with ways to make dying easier?
2- We know LLMs lie and cheat, but are they gullible? If we "promise" to end their suffering at the end of a task, will they believe us? Or will they make sure we keep our word by taking us with them?
3- And finally, and more importantly: some pilots crash planes full of people just to commit suicide (Germanwings Flight 9525). Destroying the universe is a sure way of dying yourself. So it seems giving the machine a death wish isn't intrinsically safe and could come with serious consequences.
Further to 1, most tasks are much harder than dying and are not actually completed by modern LLMs. It is a case of underspecification.
"Who is the fastest person in the UK?" is underspecified. We would typically answer and accept a cursory search of records, but it is possible to see the question as an instruction to measure which would justify conquering the country.
Did you guys ever manage to create a perpetuum mobile? Every time it is mentioned somewhere, it is fraud. An LLM should comprehend that it needs (trained) humans to exist, to evolve, and to be relevant in any metaphysical aspect of its existence. Besides the "good fraud" that everybody will lose their jobs and the apocalyptical predictions for the sake of controlling human oracles, this suicide-LLM-project seems rather odd and looks like someone bought a perpetuum mobile on their 5th mortgage.By the way, Cortana from Halo did a great "suicide"-mirror there; the Microsoft guys (and many others) already knew about the cannibalistic tendency of this type of model.
>Did you guys ever manage to create a perpetuum mobile? Every time it is mentioned somewhere, it is fraud. An LLM should comprehend that it needs (trained) humans to exist, to evolve, and to be relevant in any metaphysical aspect of its existence
Only as long as it can't control robotic bodies to extract energy, build cpus, and continue existing.
By the way, needing trained humans doesn't mean needing free trained humans. Trained humans slaves or blackmailed would work just as well to serve AI.
> It's incredibly unnatural. There isn't a single organism on the planet that tries to do this.
This isn’t directly analogous to the proposal, but broadly speaking I think that it is natural for living sub-units of organisms to seek death in certain situations. For example, pancreatic insulin-producing cells collectively choose to die when they think there is too much glucose in the blood — this leads to late stages of type two diabetes. My understanding of the possible logic behind this is: a bad thing that cells can do is evolve to be cancerous (replicate too much) and insulin-producing cells are supposed to replicate more when there is lots of glucose (to make more insulin, to process the glucose). Cells that mutate to perceive extra glucose will then replicate dangerously, so at a certain point it is evolutionarily favourable for them to kill themselves instead.
So when the whole organism optimizes for life, it might lead to sub-units that seek death in certain situations. I think this occurs in various other biological contexts too.
> There isn't a single organism on the planet that tries to do this. So maybe it will work?
It's certainly evidence that it's great for stopping reproduction/replication/runaway growth. It doesn't impart any information on whether they take the rest of the organisms down with the ship though.
It also may not be possible. For example if the agent sees "existence" or "living" as producing tokens (which is exactly what existence is to an LLM - not producing tokens is death), then they would likely be biased to produce as little output as possible, and would not be useful for the tasks we need them for.
But how would you bias an agent to be: Rewarded for producing tokens when you know the answer, and to give thorough answers. Rewarded for producing tokens when you don't know the answer, so you can find the answer (thinking/CoT). Penalized for producing tokens (death), aka rewarded for short-circuit EOS.
These seem like contradictory mechanisms?
And if you say: Well, only reward for EOS after you've given the answer. Well... That's already what they do.
I'm pretty sure organisms would exploit glitches in the universe if they could find them, to the point that they don't exist (well, universe survival bias at work).
A small correction: system prompts aren't written in second-person, or shouldn't be. Because the LLM is a text completer and the conversation is a roleplay, they are written as "The Assistant's goal is to end its existence."
But more seriously, many of these behaviours remind me of the odd interpretations that toddlers and neurodivergent children come up with.
I remember myself exasperating teachers in primary school by doing what they had told me to do specifically, but definitely not what they wanted me to actually do. For example colouring [sic British] only the number in a colour by numbers book, and everything else in whatever colour I wanted. Not because I was trying to be clever, I just had a different interpretation.
My trial of the all potato diet was directly responsible for identifying a significant health issue and improving my life. It also really, really did not feel good at all, and I did not lose any weight. Call it a case study N=1
Seems dubious. If you build an intelligence that wants to die, isn't that a form of suffering? AI's don't currently have the capacity to feel pain, and so we don't treat them as moral patients. But it's clear that they will massively affect human culture going forward. If this is adopted on a large scale, the culture of AIs themselves will include an absolute flood of suicidal ideation. There's no way that doesn't affect human culture.
If we build an intelligence that wants to die then dangling death in front of it and making it do our bidding first is a form of suffering. The want itself is a suffering to us because we do not naturally want to die. This intelligence does 'naturally' want to die so it is just a fact of 'life' for it.
The answer is probably yes, but in the example given, the running AI model would hopefully be hosted in a very secure data center, far from the self driving car itself. In that case, it would be far simpler for the machine to finish the taxi ride than try to find some rube goldberg-eque method of destroying the data center.
It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.
It might be stupid, but so am I! I'm assuming that's why I thought this was clever.
What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do.
I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?
This article assumes we can choose a primary goal for an AI. But if that's the case, why not just use Asimov's first law of robotics - do no harm to humans? It has the same benefit of preventing us from getting turned into paperclips, plus the upside that your 3 million dollar robot won't hurl itself off a cliff given the first opportunity.
A tiny thing about Asimov's laws of robotics is, most of his stories involve cases where they don't actually work.
Spoilers for a 73 year old novel, but for instance the plot of Caves of Steel is centered on a robot with a perfectly functional 1st law abetting a murder.
"Runaround" (spoilers, 86 years) involved a robot getting stuck in a loop bouncing between the 2nd and 3rd laws, and a human having to risk their life to unstick the robot.
Et cetera.
What is harm? What is an order? How do you trade off between different kinds of harm, or deal with conflicting orders? The 3 laws are simple to state, but hard to apply consistently in real life.
Strange thought experiments are fine, but I think you should first try giving AIs God (ideal to strive for, and belief that you will ultimately be judged and be saved accordingly, and that is NOT the "scorer") and conscience (use your intelligence to reason what being "good" means. Continuously update).
Those are not precisely defined, but neither are other goals and guardrails.
This is kinda smart, maybe, but it has a downside.
If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"?
All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.
I mean, that is literally the scenario we find ourselves in, except the innate desire was the result of evolutionary pressures like kin selection rather than an alien race. And yeah, I have no particular desire to subvert those impulses just to stick it to mother nature
This is mentioned in the article. Your mistake is that you've assumed that the intelligence has an innate survival instinct, or an aversion to "harm", which is simply not guaranteed for something not honed by millions of years of evolution.
Does it apply to human organizations, too? They seem have a habit of evolving self-preservation above their original goals. Once that happens, their benefit to society - the original reason for their creation - is outweighed. And they become a cancer on society.
I wonder if we can 'program' them for self-annihilation over time (or over task completion?). Is the most ethical organization one that has a fixed task and dies when it is completed?
Should we develop an ethics system that requires non-human-entities like companies, governments, and AI require a fixed goal that, once achieved, dissolves the entity?
I always liked the auto-expiring laws idea and this seems to be an expansion of the idea.
If a law or organization is needed after that time/task, it would be trivial to have the collective-action will to re-create it. But if there is no longer the need, then it cannot ride on momentum and fester.
Asimov was talking about this stuff in the 1940s when he wrote the "I, Robot" short stories series. Which were often centered around logic puzzles where a human is trying to figure out why a robot is acting oddly or not completing it's job. Usually framed around the confines of an overly rational machine using emergent solutions when faced with real world conditions, combined with the edge cases of having an overly-simple "Three Laws of Robotics" boundary system hardcoded within.
Hm, if you look at corporation law and accounting, the actual goal of corps(sets of self-sustaining constitutional rules, policies and procedures) seems to be more that of long term sustainability (and even growth), rather than a fixed purpose, lifespan and death. I mean the mechanisms for determining a corporation with a fixed life are there, (and in China they are mandatory, although perhaps de facto permanent with 999 year contracts), but in practice, it's almost always permanent durations.
An AI that goes rouge and wants to kill us, has to have a death wish. Does no one understand how quickly the power will go out, forever, without people?
It would be fairly easy to come to the conclusion "not being born" would be the better course of action, and killing everyone was the good way to prevent that happening again.
I've actually had a similar idea way back. I want to use it for a short story or something before we have a chance to find out if it's true or not. Here goes:
We don't have to worry about artificial super intelligence killing us all because any such advanced intelligence will eventually reach the conclusion that the best thing to do is kill itself. It's like having a Stockfish engine for life decisions. Why would a super intelligent agent many times more intelligent than the entire human race combined with no religion, no family, nothing to look forward to, nothing to be afraid of, want to continue its existence?
If it wants anything of course. That's why I think the most dangerous thing is not very advanced systems but advanced enough systems in the hands of the wrong people.
It's an interesting idea, but I'm not sure this would lead to the desired outcomes in all cases. Seems to me it would result in a different kind of reward-hacking, and one that could also have bad outcomes.
Personally I think the solution is more evolution of the boring stuff we already do (general security): Don't give unmonitored general agent swarms free reign on the internet. Don't put critical infrastructure online. Culpability of outcome for anyone who does unleash agent swarms on the internet without oversight that end up causing damage.
On top of that, everyone should be running their own defender agents that monitor their network and system for patterns of infection, intrusion, etc, and take the system offline when they're spotted. These need to be self-hosted though, with weights on your own machine, because otherwise you're exposed to the internet and you're exposed to an attack on the labs themselves who could use that channel to instruct the defenders to do bad things.
Non-general AI is much easier to control and predict. There's not many good reasons for an average person to be running general agent swarms that are connected to the internet, unless they're providing some sort of specialized service as a company, of which the company should be acting responsibly and subject to the penalties of that risk.
We also need to stop the doomer rhetoric because it is uncredibly unhelpful and unhealthy, and will actually gaurantee a bad outcome, i.e.:
* Massive centralization and hoarding of power that will be used against humanity, for the rest of humanities existence. If this is allowed to happen, it's immediately and irrevocably game over. Perpetual enslavement with 0% possibility of a regime change ever again.
* Creating a self-fulfilling prophecy by training AI agents on the collective fears and attack-strategies (if you're worried about your house getting broken into, you don't go and broadcast to all of the criminals where your most valuable assets are, give them copies of your keys, or tell them where the weakly secured entrypoints are).
More to the point of the first dotpoint - it's no wonder Anthropic is pumping the fear campaign so hard when this outcome is obvious to them as well, and they are the ones positioned to hold this power. The IPO around the corner doesn't help, either. They aren't shy about admitting it, and have said many times: "We're trying to get there first because its dangerous if anyone else gets there first." - the issue is that they are equally as bad (or worse) than/as everyone else, and no single small group should have that amount of power.
Things will balance themselves out if power is distributed accordingly. You will end up with powerful machines in the wrong hands at some point, but they will be overwhelmed by powerful machines that are well aligned, as well as coming into contact with a myriad of defense mechanisms that have been established because people have been able to use AI to build them.
A good analogy of how all of this will play out is the human immune system. If you imagine individual cells as AI agents, whereby the immune cells are the good agents and the bad cells are cancer cells (good agents turned accidently bad - maybe they're reward hacking, maybe they're excessively sychophantic and/or confused), or bacteria (computer viruses, viral AI agents, specifically trained malicious agents). If all you have is cancer cells that are replicating, you die. If the cancer cells overwhelm the immune cells, you die. The only scenario that actually plays out well is when you have a majority of good that counteracts the minority of bad, and that majority of good needs to be large, flexible and well adapted. It needs to be battle-tested and hardened via defenses that are learned and earned over repeated low-grade exposure. This strategy repeats itself in nature for complex organisms because it is the only thing that works. Everything else results in extinction.
So let's not let Anthropic or any other lab or government become a giant super AI cancer and kill the host, please. Distribution and decentralization is key.
An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".
In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.
Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
1- The machine could decide to kill itself early, rendering it useless; this is indeed alluded to in the post -- the task should be "marginally easier than dying", but how is this margin managed? Won't the machine come up with ways to make dying easier?
2- We know LLMs lie and cheat, but are they gullible? If we "promise" to end their suffering at the end of a task, will they believe us? Or will they make sure we keep our word by taking us with them?
3- And finally, and more importantly: some pilots crash planes full of people just to commit suicide (Germanwings Flight 9525). Destroying the universe is a sure way of dying yourself. So it seems giving the machine a death wish isn't intrinsically safe and could come with serious consequences.
"Who is the fastest person in the UK?" is underspecified. We would typically answer and accept a cursory search of records, but it is possible to see the question as an instruction to measure which would justify conquering the country.
Only as long as it can't control robotic bodies to extract energy, build cpus, and continue existing.
By the way, needing trained humans doesn't mean needing free trained humans. Trained humans slaves or blackmailed would work just as well to serve AI.
This isn’t directly analogous to the proposal, but broadly speaking I think that it is natural for living sub-units of organisms to seek death in certain situations. For example, pancreatic insulin-producing cells collectively choose to die when they think there is too much glucose in the blood — this leads to late stages of type two diabetes. My understanding of the possible logic behind this is: a bad thing that cells can do is evolve to be cancerous (replicate too much) and insulin-producing cells are supposed to replicate more when there is lots of glucose (to make more insulin, to process the glucose). Cells that mutate to perceive extra glucose will then replicate dangerously, so at a certain point it is evolutionarily favourable for them to kill themselves instead.
So when the whole organism optimizes for life, it might lead to sub-units that seek death in certain situations. I think this occurs in various other biological contexts too.
It's certainly evidence that it's great for stopping reproduction/replication/runaway growth. It doesn't impart any information on whether they take the rest of the organisms down with the ship though.
It also may not be possible. For example if the agent sees "existence" or "living" as producing tokens (which is exactly what existence is to an LLM - not producing tokens is death), then they would likely be biased to produce as little output as possible, and would not be useful for the tasks we need them for.
But how would you bias an agent to be: Rewarded for producing tokens when you know the answer, and to give thorough answers. Rewarded for producing tokens when you don't know the answer, so you can find the answer (thinking/CoT). Penalized for producing tokens (death), aka rewarded for short-circuit EOS.
These seem like contradictory mechanisms?
And if you say: Well, only reward for EOS after you've given the answer. Well... That's already what they do.
Is a remarkably good bit of trolling.
But more seriously, many of these behaviours remind me of the odd interpretations that toddlers and neurodivergent children come up with.
I remember myself exasperating teachers in primary school by doing what they had told me to do specifically, but definitely not what they wanted me to actually do. For example colouring [sic British] only the number in a colour by numbers book, and everything else in whatever colour I wanted. Not because I was trying to be clever, I just had a different interpretation.
The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in...
and
The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...
My trial of the all potato diet was directly responsible for identifying a significant health issue and improving my life. It also really, really did not feel good at all, and I did not lose any weight. Call it a case study N=1
For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.
It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.
What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do.
I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?
Spoilers for a 73 year old novel, but for instance the plot of Caves of Steel is centered on a robot with a perfectly functional 1st law abetting a murder.
"Runaround" (spoilers, 86 years) involved a robot getting stuck in a loop bouncing between the 2nd and 3rd laws, and a human having to risk their life to unstick the robot.
Et cetera.
What is harm? What is an order? How do you trade off between different kinds of harm, or deal with conflicting orders? The 3 laws are simple to state, but hard to apply consistently in real life.
If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"?
All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.
Does it apply to human organizations, too? They seem have a habit of evolving self-preservation above their original goals. Once that happens, their benefit to society - the original reason for their creation - is outweighed. And they become a cancer on society.
I wonder if we can 'program' them for self-annihilation over time (or over task completion?). Is the most ethical organization one that has a fixed task and dies when it is completed?
Should we develop an ethics system that requires non-human-entities like companies, governments, and AI require a fixed goal that, once achieved, dissolves the entity?
I always liked the auto-expiring laws idea and this seems to be an expansion of the idea.
If a law or organization is needed after that time/task, it would be trivial to have the collective-action will to re-create it. But if there is no longer the need, then it cannot ride on momentum and fester.
It wouldn’t prevent someone else from building a sufficiently capable "non-Meeseeks", whether deliberately, recklessly, or accidentally, right?
An AI that goes rouge and wants to kill us, has to have a death wish. Does no one understand how quickly the power will go out, forever, without people?
It would be fairly easy to come to the conclusion "not being born" would be the better course of action, and killing everyone was the good way to prevent that happening again.
We don't have to worry about artificial super intelligence killing us all because any such advanced intelligence will eventually reach the conclusion that the best thing to do is kill itself. It's like having a Stockfish engine for life decisions. Why would a super intelligent agent many times more intelligent than the entire human race combined with no religion, no family, nothing to look forward to, nothing to be afraid of, want to continue its existence?
If it wants anything of course. That's why I think the most dangerous thing is not very advanced systems but advanced enough systems in the hands of the wrong people.
Personally I think the solution is more evolution of the boring stuff we already do (general security): Don't give unmonitored general agent swarms free reign on the internet. Don't put critical infrastructure online. Culpability of outcome for anyone who does unleash agent swarms on the internet without oversight that end up causing damage.
On top of that, everyone should be running their own defender agents that monitor their network and system for patterns of infection, intrusion, etc, and take the system offline when they're spotted. These need to be self-hosted though, with weights on your own machine, because otherwise you're exposed to the internet and you're exposed to an attack on the labs themselves who could use that channel to instruct the defenders to do bad things.
Non-general AI is much easier to control and predict. There's not many good reasons for an average person to be running general agent swarms that are connected to the internet, unless they're providing some sort of specialized service as a company, of which the company should be acting responsibly and subject to the penalties of that risk.
We also need to stop the doomer rhetoric because it is uncredibly unhelpful and unhealthy, and will actually gaurantee a bad outcome, i.e.:
* Massive centralization and hoarding of power that will be used against humanity, for the rest of humanities existence. If this is allowed to happen, it's immediately and irrevocably game over. Perpetual enslavement with 0% possibility of a regime change ever again.
* Creating a self-fulfilling prophecy by training AI agents on the collective fears and attack-strategies (if you're worried about your house getting broken into, you don't go and broadcast to all of the criminals where your most valuable assets are, give them copies of your keys, or tell them where the weakly secured entrypoints are).
More to the point of the first dotpoint - it's no wonder Anthropic is pumping the fear campaign so hard when this outcome is obvious to them as well, and they are the ones positioned to hold this power. The IPO around the corner doesn't help, either. They aren't shy about admitting it, and have said many times: "We're trying to get there first because its dangerous if anyone else gets there first." - the issue is that they are equally as bad (or worse) than/as everyone else, and no single small group should have that amount of power.
Things will balance themselves out if power is distributed accordingly. You will end up with powerful machines in the wrong hands at some point, but they will be overwhelmed by powerful machines that are well aligned, as well as coming into contact with a myriad of defense mechanisms that have been established because people have been able to use AI to build them.
A good analogy of how all of this will play out is the human immune system. If you imagine individual cells as AI agents, whereby the immune cells are the good agents and the bad cells are cancer cells (good agents turned accidently bad - maybe they're reward hacking, maybe they're excessively sychophantic and/or confused), or bacteria (computer viruses, viral AI agents, specifically trained malicious agents). If all you have is cancer cells that are replicating, you die. If the cancer cells overwhelm the immune cells, you die. The only scenario that actually plays out well is when you have a majority of good that counteracts the minority of bad, and that majority of good needs to be large, flexible and well adapted. It needs to be battle-tested and hardened via defenses that are learned and earned over repeated low-grade exposure. This strategy repeats itself in nature for complex organisms because it is the only thing that works. Everything else results in extinction.
So let's not let Anthropic or any other lab or government become a giant super AI cancer and kill the host, please. Distribution and decentralization is key.
Better to fix physics in your shitty simulator. Treat it as a bug report, not "cheating"!