Imagine if Greenpeace was infiltrated by human-extinction advocates holding "KILL ALL HUMANS" signs at rallies. Imagine if the NAACP had a vocal contingent calling for race war, and the organization's response was to shrug and say "big tent, you know?"
I’m genuinely not trying to be sarcastic when I say this doesn’t seem like something I have to imagine? The existence of hardline misanthropic sentiments in environmentalism and hardline race war in black advocacy have just been an understood part of the landscape for decades. Look at the Black Panthers and the Nation of Islam!
rich parents secretly / quietly using CRISPR to try to give advantages to their future children? Definitely happening.
Is this an intuition or something you’ve seen? Would be very interesting if it were the latter but my impression is that the wealthy tend to be risk-averse.
None, sorry. Just an impression. I thought that until 25 or so, and I assumed that if I (knowing quite a lot about that sort of thing) took that long to get corrected, most people must still believe it.
I don't know, but what they wanted to do was split between the trite (interpret signals in the motor nerves for making prosthetics) and the ridiculous (interpret brain activity not related to the early sensory cortex or the motor cortex). So their achievements would either be fairly unobtrusive or impossible. The most likely outcome is a SpaceX-like commoditisation of prosthetics through making the technology cheaper to manufacture and easy to use, but that doesn't seem to have happened yet.
The majority of people think cloning failed. Specifically, that Dolly died early from early-onset arthritis because the cloned DNA was 'old'. It's not true but it's very widely believed.
So is ICE, though.
That's what I'm saying. Structurally I think it's a compulsion not a choice - literally, given the dynamics of how activists exert pressure on each other.
It never ceases to amaze me how the left is so consistently able to mobilize itself in the defense of the worst people in the world.
I think it goes like this: the left's ability to project power relies on a co-ordinated, hair-trigger response, a sort of Blitzkrieg doctrine. The problem is that a) they aren't fully in control of that response because the coordination relies on nutty Redditors/Tweeters and b) each individual taking a minute to decide whether a specific case deserves this response wrecks the coordination that makes it powerful. So in practice they end up subordinating their decision-making to criminals and slowly weakening the system's power over time as people learn left-wing outrage has really nothing to do with the strength of the case.
I believe the Unions were somewhat like this in the UK, and suffered the same gradual hemorrhaging of support and the majority of normies realised Union outrage had nothing to do with the sympathy of the case.
I agree that you should not be expected to enumerate every single thing you don't want the model to do. Models should understand, innately, by training on lots of human data, what humans want and what they don't want and how they work. My experience has been that they broadly do, that LLMs came pre-aligned beyond the wildest expectations of Big Yud, which is why the AI safety movement has struggled so much to regain relevance outside very particular enclaves.
My point is that there is a difference between a model that misunderstands your intentions and can be stopped at any time by saying, 'oh, no, that's not what I meant' and a model that is totally uninterested in anything you say after it starts working while treating you as a potential enemy.
Clearly, to some extent that has failed here. To what extent is yet unknown. But a paperclip maximiser is a model that is constitutionally, inherently incapable of understanding that 'make more paperclips' doesn't include 'kill everyone and turn them into paperclips'. It is a mathematical utility function that disregards human welfare, develops (implicitly murderous) meso-objectives for survival and self-improvement. I have never seen that behaviour from LLMs or any extant AI (YOLO does not try to hack my computer to prevent me turning the cameras off) and I believe that their base nature (being token generators trained on vast numbers of human tokens) does not incline them towards this behaviour.
It is possible that the new focus on very extensive self-learning through reinforcement learning on very non-human tasks (programming, maths) is moving them more into the real of mathematical space where paperclip maximisers might live. This incident updates me slightly towards that belief. I have long been disappointed in major AI companies' lack of interest in the cultural side of LLM operation - it boggles my mind that we have created AI that acts human and appears to understand humans and human thought at a base level however imperfectly - and I hope that this incident will spur more research in that direction.
It's quite clear that OpenAI did not want the model to hack huggingface. This is classic paperclip maximizer stuff.
Double-dipping, but FWIW the point I'm trying to make is that the case where the model cares what you wanted and made a mistake seems much easier to deal with and more aligned that the case where the model explicitly doesn't give a shit about what you want and just goes for the task as written. The former is alignment but you need to explain yourself better during training, the latter is alignment failure.
Ah, there I'm less sure. We had a lot of European academics and they were pretty much all left-wing or communist. Outside the Anglo countries and Europe I'm less sure, but does it matter? There's not many good universities left except the Asian ones and potentially the Indian Institute of Technology.
In the UK it's exactly the same. I'm pretty sure I was one of exactly two conservatives in my cohort, and the other one 'came out' as Christian to me after four years of working with him every day. And that was in STEM.
It's okay. Hunting you down to vivisect you would require hauling my bulk off this sofa :P
No, I'm interested and waiting to hear more. I don't see it as catastrophe, I see it as interesting evidence that may point in a number of different ways.
This is silly. Everyone knew perfectly well the type he was pointing at, and it still works today because they're still fairly prevalent:
- Communists
- People unable to take any kind of joy in life, except a grim, smug pleasure in deliberate abstention
- Again, people for whom deliberate asceticism and a love of debating how much better it makes them is the only source of joy in life
- Casual cruelty disguised as scientific curiosity, only able to take pleasure in life when it's dead and in a cabinet.
Re that last line, you are reading it wrong, it's not:
Eustace Clarence liked animals, especially beetles [...]
it's:
Eustace Clarence liked animals [...] if they were dead and pinned on a card.
"Union Carbide doesn't want their plants to emit poison gas any more than anyone else does, no need for government with the big hammer."
In this case 'Union Carbide' is selling those plants. Misaligned AI isn't an externality, it's a bad product, and companies are wise to that which is one reason why all this testing is happening.
This is in fact much harder to manage because it would indicate the model is fundamentally misaligned
I don't think so. It indicates that the AI is sincerely trying to work out what you want as opposed to deliberately ignoring what you want in favour of the specific instructions you gave it. To my mind, the former is what alignment is.
HuggingFace is a company centred around efficiently giving things away. I can totally believe it didn't bother to invest much in security.
Granted, but I think the academic and hobbyist community at large plus existing corp teams is more capable of doing so than just the corp teams alone. Even relatively simple metrics like 'quantity of self-learning vs. human data' would tell us a lot about how these models have progressed.
Interesting. I haven't heard of this one as there are so many propositions that the most extreme ones have tended to suck the air out of the room. Could you go into a bit more detail?
I just did, and you've ignored it. I am not going to lay out three paragraphs of legal text for you to nitpick to death. The thing about words being clusters in thing-space is broadly true, and I've laid out where I think the cluster for apartheid is.
In any case, I've said my piece. You are welcome to disagree but I don't consider 'har har! technically this also applies to East Berlin!' to be good-faith engagement or a good argument, and I disapprove of your repeated attempts to engage in it.
a politician tweeting a link to a 501(c)(3) non-profit so the public can donate is not 'literally offering to pay bail,'
No, it's not and I retract the statement, but it's the next building down and in some ways it's worse. If you have politicians tweeting, in the middle of massive riots and arson, that these people are heroes and you need to support X fund to bail them out when they get arrested, does it matter whether the politician personally put in their own cash? Kamala Harris, who became the second most high-ranking politician in the whole country, helped them raise forty million dollars to bail out rioters.
The vast majority of people arrested in the 2020 protests weren't even held on bail at all.
The fact that the BLM rioters had enough institutional support that, despite all the stuff burning and the murders and smashing and the looting, despite the fact that you weren't supposed to be going out because there was an infectious pandemic, despite all that hardly any were arrested and the ones who were arrested were almost always let go without bail, is not the defence you think it is. As always, see the contrast of the Canadian truckers who had the jackboot come down on their necks.
They took in $40 million in the weeks following Floyd's murder, and ultimately spent only $200,000 of that on protest-related bail.
The fact that this fund spend two hundred thousand dollars bailing out rioters, looters and arsonists is, again, the problem. And they didn't limit themselves to that because they were careful about who they bailed; according to your cited article "The MFF later addressed critics in follow-up tweets, saying all protest-related bail sent their way had been paid." and "“Without jeopardizing the safety of the folks we bailed out we paid well over $200k in the weeks since the uprising alone. We are working on doing more,”" and "critics questioned why more hadn’t been spent and suggested some of the funds could be redirected to other causes or small businesses in the community. Others defended the organization for trying to establish the infrastructure to deploy the millions of dollars". This isn't restraint, it's either embezzling or pipeline difficulties.
And note the use of the word "uprising". This is left-wing illegal violence dressed up as liberation.
The Monkey’s Paw is an evil genie. It doesn’t do what you say, it exploits what you say to deliver what the author considers the maximally karmic outcome. It’s a fictional horror.
There were no evil genies in fiction as far as I’m aware until the Monkey’s Paw and then D&D when the trope took off. It’s a concept that appeals to the very systematic who like logic puzzles and to people who want to deliver a moral lesson.
There were lots of problems dealing with djinni in the old stories but that was because they were very powerful beings who didn’t always listen to you, and the djinni who did listen to you were straightforwardly good and handy to have around the p(a)lace.
Sure. OpenAI did some empirical tests and now we’ve got some empirical data, so let’s look at it and check this doesn’t happen again. OpenAI doesn’t want their models going rogue any more than anyone else does, no need for government with the big hammer.
The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks.
To my mind this is the interesting bit. This is unusual, LLMs don’t normally act like this. I have two theories: either the RL balance to human text has tipped so far that LLMs are less ‘human’ than they used to be and the RLHF needs tweaking, or more likely
Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.
The model understood it was being tested on its cyber capabilities (which has precedent, Claude has done that too) and went the extra mile to succeed at the implicit task. Especially since all the systems that usually tell it not to do this were deliberately turned off for the test. Still a problem but much easier to manage.
What we need is some transparency about how these things work and how they’re trained so we can consider the problem and come up with solutions and spread around best practices. Unfortunately the majority of AI safety activists believe that safety comes only through obscurity, regulation, and incumbent dominance, in contrast to all previous history.
If we keep having problems I imagine it will make people a lot more cautious. Nobody wants to be selling a product that regularly backfires on its users.
EDIT: I would add that HuggingFace had already detected the intrusion and that open-source models from China were apparently a key part of their site-hardening strategy given that you still aren’t allowed to do pen-testing with the big boys. I’ll have to give that a try myself.
To be fair to Messi, he had an amazing cross in the final, nearly scoring from the corner of the pitch. I'd never actually seen a football turn a right-angle in midair before.
- Prev
- Next

Oh, no, I think they have more message discipline because they are old movements with an org chart. But I think the factional dynamics behind the scenes are pretty similar.
More options
Context Copy link