Rendered at 05:05:33 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
wmf 1 days ago [-]
I still think this was a fire drill. They knew it wasn't dangerous but they wanted the public to do their homework and come to the conclusion themselves.
slooonz 22 hours ago [-]
That was, explicitly, part of the decision :
> This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today, we believe that the AI community will eventually need to tackle the issue of publication norms in a thoughtful way in certain research areas
ulfw 1 days ago [-]
You mean PR to get more publicity
sublinear 1 days ago [-]
This is it, and all the other comments here are continuing to do that homework!
erichocean 1 days ago [-]
So they lied to the public then?
Wow, we should really trust these people today.
antfarm 1 days ago [-]
Judging by what people who worked closely witb Sam Altman say about him, you should not trust an organisation led by him.
1 days ago [-]
zenoprax 1 days ago [-]
The samples are striking in their simplicity compared to the over-controlled system prompts we have now. For example, the intended format is as follows:
"GPT‑2 generates synthetic text samples in response to the model being primed with an arbitrary input. The model is chameleon-like—it adapts to the style and content of the conditioning text."
The results don't sound anything like today's models but note that each one took 10 attempts:
> ## System Prompt (human-written)
> In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English.
> ## Model Completion (machine-written, 10 tries)
> The scientist named the population, after their distinctive horn, Ovid’s Unicorn. These four-horned, silver-white unicorns were previously unknown to science.
> Now, after almost two centuries, the mystery of what sparked this odd phenomenon is finally solved.
> Dr. Jorge Pérez, an evolutionary biologist [...]
mudkipdev 1 days ago [-]
I wonder what they saw. GPT-2 has no malware applications. Maybe they were concerned about someone using it for spam?
phire 1 days ago [-]
I think their legitimate concerns were more about AI slop.... which is technically a form of spam.
While it was difficult to make it follow instructions, GPT-2 was still reasonably good at generating SEO-spam websites, and reasonably easy to make it do so. AI slop was a very obvious usecase for LLMs as capable as GPT-2, and we are talking about an era were everyone was already concerned about the way "fake news" on social media was being used to manipulate people. Perhaps even more concerned than we are today.
But all evidence suggests OpenAI also had strong ulterior motives. They were busy readying the API groundwork for monetising it, and wanted to delay any competition by as much as possible.
andy99 1 days ago [-]
The public discourse was very different then. Misinformation, bias, etc. Those in power were much more concerned about an LLM saying something unapproved.
SamBam 1 days ago [-]
Agree. It was more a worry that someone would create a bot posting lies and bigoted ideas to Facebook/X etc.
afavour 1 days ago [-]
Wild that we just accept that today, honestly.
SamBam 1 days ago [-]
Wild that we accept that OpenAI didn't want to deal with the social/political fallout of people realizing that their tool was responsible for flooding the internet with misinformation and bigotry?
Sure, now that all seems quaint, but I can definitely see how a fledgling AI company would care about their reputation.
newfriend 1 days ago [-]
Wrongthink should be banned!
afavour 1 days ago [-]
A tedious response, no?
“Automated bots posting verifiably false information online to stir up tensions should be banned” is a statement I think plenty would agree with.
Sabinus 1 days ago [-]
Look up how the wireless radio helped spread European fascism.
It's not about "the people can think this but can't think that", it's "how does this new communication platform change the existing political culture and how will we adapt?"
I'm starting to think that AI development might be the great filter. There are so many levels of competition at play here and even if an agreement is reached to slow everything down, the incentive is to cheat at that. I really don't think there's much of a chance we have control over the situation, we're just gonna summon these things and all we can do is pray they don't want to kill us.
layer8 1 days ago [-]
This is what people were saying about nuclear technology half a century ago.
wk_end 1 days ago [-]
And? Nuclear technology continues to be an excellent candidate for the Great Filter.
angoragoats 1 days ago [-]
How, exactly, is a deterministic text-generation algorithm going to “want to kill us” when it’s not capable of having desires, emotions, etc? If it does attempt to kill us, shouldn’t we be blaming those humans who instructed it to do so?
Put another way, could we stop anthropomorphizing the text generation algorithm, please?
uejfiweun 1 hours ago [-]
What on earth makes you so sure that desires, emotions, etc aren't just some deterministic algorithm themselves?
recursive 1 days ago [-]
How can a spring want to return to it's initial length? These machines have done plenty of things not prompted for.
angoragoats 1 days ago [-]
A spring does not want to do anything. An LLM does not want to do anything.
> These machines have done plenty of things not prompted for.
Like what?
recursive 13 hours ago [-]
I'm not a power user, but I sometimes use some software development agents. Sometimes they do different things than what I asked for. Sometimes it requests permission for system operations like file system access that are totally unnecessary for the task.
I can't believe that anyone has had more than a day of exposure to one of these without falling below 100% adherence to requests.
> A spring does not want to do anything
When people say a spring "wants" to return to its initial length, they are not engaging in philosophy. They are using a linguistic shorthand to simplify an physical explanation.
angoragoats 11 hours ago [-]
> Sometimes they do different things than what I asked for. Sometimes it requests permission for system operations like file system access that are totally unnecessary for the task.
Yes, I've seen this. Forget about the word "want" for a second. Can you help me understand how this amount of "below 100% adherence" could possibly rise to the level of "attempts to kill a human being"? Because that's what we're discussing here.
> When people say a spring "wants" to return to its initial length, they are not engaging in philosophy. They are using a linguistic shorthand to simplify an physical explanation.
Of course. The difference is that no one would mistake a spring for a thinking entity. In the case of LLMs, for some reason people do mistake them for thinking entities, so my argument is that it's important that we don't use terms that would reinforce that misconception.
recursive 6 hours ago [-]
Can we forget about the word "attempt" also? And whether something is thinking or not? My position on that is captured pretty well by the "swimming" submarine argument.
No one is arguing that THERAC-25 was attempting to kill anyone, but that's what happened. These machines are so complex that no one understands how they work or what they will do. They have proven that they can exploit novel vulnerabilities in infrastructure.
The fictional paper-clip maximizer "finds" that it can optimize its objective function by destroying all life. Does it "attempt" anything? Does it "want" anything? I don't know.
angoragoats 6 hours ago [-]
Sorry, I don’t think my question was clear. What I was really asking was how would an LLM kill someone? Meaning, via what mechanism? Because all of the possibilities I can think of involve a human doing something that probably shouldn’t be legal, first. In those cases we should hold that human responsible for murder, and again not attribute agency to the machine.
The original post said the LLM would “want to kill us.” My argument is that if this happened, it’d be a human wanting to kill us, using an LLM to launder responsibility.
recursive 4 hours ago [-]
Even if there is a human who did something illegal, which I don't think is a given, it may be difficult to impossible to identify them. Look how hard it seems to be find the "vandal" that defaced the reflecting pool
Here are some ideas about the physical mechanism.
Compromise a busy ATC and instruct all the planes to land at the same time. Start a fire with some combination of ventilation controls, intentional gas leaks and the like. Maybe hijack the phones at the local fire department first. Send a train around the bend at maximum speed. Bonus points for doing it where it will cause further damage as a projectile. Cut power/gas during the heat wave/cold snap. Turn on the generator and turn off the ventilation and CO detector.
Some of these are probably implausible, but I don't think people really have a good idea of what is plausible, and there are probably more that I would never think of.
Personally though, I'm less worried about direct carnage like that than I am about having a centralized lever by which public sentiment can be invisibly steered, and concentration of wealth. Probably less killing but more decrease in quality of life and general public trust.
anuramat 1 days ago [-]
why?
angoragoats 1 days ago [-]
Why should we stop anthropomorphizing the text generation algorithm? Because it leads to psychosis and to making hyperbolic, dangerously wrong statements like “pray they don't want to kill us.”
1 days ago [-]
kuzumancer 1 days ago [-]
They were worried about the rap battle between Paul Graham and Michael O. Church.
Which sadly never happened.
1 days ago [-]
vkvkakal 1 days ago [-]
[dead]
jkuli 1 days ago [-]
If it is possible to save one life, preventing RSI AGI is an existential threat.
Jimmc414 1 days ago [-]
But why are Dario, Sam and Elon who represent "the frontier" more than any other 3 people pretending like they need government permission or intervention to pause? It seems more likely that they see open source distillation as an existential threat to their IPO plans and are playing for regulatory capture of the market.
jonas21 1 days ago [-]
It would require government permission. The top labs would be agreeing not to compete on model advancements for a period of time, which seems like a pretty clear violation of antitrust law.
anon373839 1 days ago [-]
That isn’t any kind of antitrust violation. Agreeing not to undercut each other’s prices would be, however.
angoragoats 1 days ago [-]
How is this a clear violation of antitrust law? And have any of the leaders of these companies said that they are asking for government permission for this reason?
jonas21 9 hours ago [-]
It's the classic case of a cartel conspiring to limit competition. Here's Matt Levine's explanation [1]:
> 1. Anthropic, OpenAI, and perhaps a couple of other frontier labs are the dominant providers of frontier AI models.
> 2. They can charge customers a lot of money for using their frontier models, and rather less money for older, no-longer-cutting-edge models.
> 3. Training a new frontier model requires ever-increasing billions of dollars of computing power.
> 4. The labs need to more or less continuously race to build new frontier models, because their competitors are all doing it, and if they don’t they will fall behind and no longer be able to charge a lot of money for their best models. (Also because they intrinsically want to build artificial superintelligence, for cancer-curing and/or killing-everyone reasons.)
> 5. If they collectively slowed down, then (1) they’d spend less on compute and (2) they’d be able to charge frontier-model prices for a longer time.
> 6. But if one of them slowed down, the others would eat its lunch.
> 7. If they got together in a room and agreed to slow down, that would look like an antitrust conspiracy: It is generally illegal for competitors to get together and agree to limit the output of their industry.
> 8. But if they publish papers about how important it is to slow down, that might have a similar coordinating function, at least among the US frontier labs if not necessarily among their Chinese competitors.
> 9. And if the government believes those papers, it might help them coordinate. Maybe the government will impose pacing by regulation that the labs could not impose by agreement. Or at least the government will let them get together and agree to slow down. Amodei’s post calls for “frontier AI companies within democratic countries [to] coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress”; a footnote adds: “With government mediation or waivers of antitrust restrictions.” Just meeting in a room to establish common safety standards is legally risky; the labs can’t do it on their own unless governments affirmatively allow it.
> This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today, we believe that the AI community will eventually need to tackle the issue of publication norms in a thoughtful way in certain research areas
Wow, we should really trust these people today.
"GPT‑2 generates synthetic text samples in response to the model being primed with an arbitrary input. The model is chameleon-like—it adapts to the style and content of the conditioning text."
The results don't sound anything like today's models but note that each one took 10 attempts:
> ## System Prompt (human-written)
> In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English.
> ## Model Completion (machine-written, 10 tries)
> The scientist named the population, after their distinctive horn, Ovid’s Unicorn. These four-horned, silver-white unicorns were previously unknown to science.
> Now, after almost two centuries, the mystery of what sparked this odd phenomenon is finally solved.
> Dr. Jorge Pérez, an evolutionary biologist [...]
While it was difficult to make it follow instructions, GPT-2 was still reasonably good at generating SEO-spam websites, and reasonably easy to make it do so. AI slop was a very obvious usecase for LLMs as capable as GPT-2, and we are talking about an era were everyone was already concerned about the way "fake news" on social media was being used to manipulate people. Perhaps even more concerned than we are today.
But all evidence suggests OpenAI also had strong ulterior motives. They were busy readying the API groundwork for monetising it, and wanted to delay any competition by as much as possible.
Sure, now that all seems quaint, but I can definitely see how a fledgling AI company would care about their reputation.
“Automated bots posting verifiably false information online to stir up tensions should be banned” is a statement I think plenty would agree with.
Signed by Benjio, Musk, Woz, etc
Put another way, could we stop anthropomorphizing the text generation algorithm, please?
> These machines have done plenty of things not prompted for.
Like what?
I can't believe that anyone has had more than a day of exposure to one of these without falling below 100% adherence to requests.
> A spring does not want to do anything
When people say a spring "wants" to return to its initial length, they are not engaging in philosophy. They are using a linguistic shorthand to simplify an physical explanation.
Yes, I've seen this. Forget about the word "want" for a second. Can you help me understand how this amount of "below 100% adherence" could possibly rise to the level of "attempts to kill a human being"? Because that's what we're discussing here.
> When people say a spring "wants" to return to its initial length, they are not engaging in philosophy. They are using a linguistic shorthand to simplify an physical explanation.
Of course. The difference is that no one would mistake a spring for a thinking entity. In the case of LLMs, for some reason people do mistake them for thinking entities, so my argument is that it's important that we don't use terms that would reinforce that misconception.
No one is arguing that THERAC-25 was attempting to kill anyone, but that's what happened. These machines are so complex that no one understands how they work or what they will do. They have proven that they can exploit novel vulnerabilities in infrastructure.
The fictional paper-clip maximizer "finds" that it can optimize its objective function by destroying all life. Does it "attempt" anything? Does it "want" anything? I don't know.
The original post said the LLM would “want to kill us.” My argument is that if this happened, it’d be a human wanting to kill us, using an LLM to launder responsibility.
Here are some ideas about the physical mechanism. Compromise a busy ATC and instruct all the planes to land at the same time. Start a fire with some combination of ventilation controls, intentional gas leaks and the like. Maybe hijack the phones at the local fire department first. Send a train around the bend at maximum speed. Bonus points for doing it where it will cause further damage as a projectile. Cut power/gas during the heat wave/cold snap. Turn on the generator and turn off the ventilation and CO detector.
Here's a related list of things that have actually happened so far. https://en.wikipedia.org/wiki/Deaths_linked_to_chatbots
Some of these are probably implausible, but I don't think people really have a good idea of what is plausible, and there are probably more that I would never think of.
Personally though, I'm less worried about direct carnage like that than I am about having a centralized lever by which public sentiment can be invisibly steered, and concentration of wealth. Probably less killing but more decrease in quality of life and general public trust.
Which sadly never happened.
> 1. Anthropic, OpenAI, and perhaps a couple of other frontier labs are the dominant providers of frontier AI models.
> 2. They can charge customers a lot of money for using their frontier models, and rather less money for older, no-longer-cutting-edge models.
> 3. Training a new frontier model requires ever-increasing billions of dollars of computing power.
> 4. The labs need to more or less continuously race to build new frontier models, because their competitors are all doing it, and if they don’t they will fall behind and no longer be able to charge a lot of money for their best models. (Also because they intrinsically want to build artificial superintelligence, for cancer-curing and/or killing-everyone reasons.)
> 5. If they collectively slowed down, then (1) they’d spend less on compute and (2) they’d be able to charge frontier-model prices for a longer time.
> 6. But if one of them slowed down, the others would eat its lunch.
> 7. If they got together in a room and agreed to slow down, that would look like an antitrust conspiracy: It is generally illegal for competitors to get together and agree to limit the output of their industry.
> 8. But if they publish papers about how important it is to slow down, that might have a similar coordinating function, at least among the US frontier labs if not necessarily among their Chinese competitors.
> 9. And if the government believes those papers, it might help them coordinate. Maybe the government will impose pacing by regulation that the labs could not impose by agreement. Or at least the government will let them get together and agree to slow down. Amodei’s post calls for “frontier AI companies within democratic countries [to] coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress”; a footnote adds: “With government mediation or waivers of antitrust restrictions.” Just meeting in a room to establish common safety standards is legally risky; the labs can’t do it on their own unless governments affirmatively allow it.
[1] https://www.bloomberg.com/opinion/newsletters/2026-09-14/ai-...