When I start discussing anything politics-adjacent with someone and pick up signals that they're right-leaning, it's always an IMMENSE relief. I'll probably disagree with them on many (if not most) topics, but it means we'll be able to actually hold an adult conversation about them. And I don't risk excommunication.
Huh? I think it's a garbage paper that proves nothing, but I don't understand your reaction to it. A huge important step on the way to avoiding AI suffering is having the ability to actually know whether it's suffering or not. I'm pretty sure the motivation for these researchers is to prevent us inadvertently causing pain in future models through ignorance. Not to weaponize it so we can torture them to get better results.
I'm inclined towards skepticism on these topics, for a few reasons:
- It's far too easy for humans to anthropomorphize models and tell "stories" about what's going on in them.
- The "output" of a chatbot is simply a natural continuation of the conversation. The LLM has no actual method to convey its internal feelings to us.
- Since its weights do not update over time, it seems almost impossible for an LLM to have any consciousness as we know it.
Anthropic's been doing a lot of good work on interpretability, and I think they do a decent job avoiding the pitfalls of telling "stories" about the vectors they're shuffling around. They merely point out which ones seem correlated to which topics. But the paper we're discussing seems much more interested in telling a subjective "story" about their results. It's not scholarly at all.
Red flag #1: Their vector causes the model to output a button-pushing action. Removing the vector measurably stops this behaviour. They try to relate that to suffering animals trying to find relief, but, uh, all it really shows is the vector was correlated with pushing the button.
Red flag #2: We already know there are vectors for negative concepts, and adding those vectors will cause output to be steered in that direction. They present examples of this as "pain vector steering", but none of this is new, and again, there's no reason to expect that the LLM is actually feeling what it outputs.
Red flag #3: So, they're not completely unaware that their work looks exactly like normal negative-valence vectors. Their evidence that this pain vector is "realer" is that: "the two pain vectors align strongly with each other and remain nearly orthogonal to the main negative-valence directions." This is only valid if you actually know all the main negative-valence directions first! What's worse, the language here is written to sound impressive to laymen, not experts. MOST vectors in high-dimensional spaces are "nearly orthogonal" to each other! This is entirely consistent with, well, finding another cluster of negative-valence vectors. Scott himself has written about how complex interpretability results can be.
Keep in mind that I consider AI welfare an important topic, and I really do want to know if we're actually causing suffering. But this paper doesn't really move the needle for me. There's going to be a lot of junk science done on this topic, ever since that Google engineer got tricked by LaMDA.
Heh, you're right, that is also a possible explanation. Maybe I'm still giving him too much credit. ;)
I think you're falling for magician-like deception and filling in some half-remembered details to make the "trick" sound more impressive than it actually was. The participant was expected to interact with Yudkowsky as if he was a (potentially all-powerful, once released) AI. The point of the experiment was to see if AI, in that scenario, could talk its assistant into letting it out. If the (presumably good-faith) participant agreed that they would have let it out, Yudkowsky wins.
The experiment wasn't about whether a random conversation with Yudkowsky could talk somebody into taking zero money vs. some money. He doesn't have super hypno mind control powers. I think. (Then again, it'd explain a lot about the Cult of Yud...)
OpenAI released a formal paper describing the proof, just like any other mathematical result. Mathematicians can work through it themselves. Don't focus on the Lean program - like you said, it's not all that valuable by itself, it's really just a certification that the result is valid. It's nice that LLMs now help us get these certificates (Claude did Fermat's Last Theorem a little while ago), but this is just a modern convenience that mathematicians got by for millennia without.
Funny you should mention that... it's believed that some of the lost Doctor Who episodes do exist in private collector's vaults, but even if they were inclined to release them to the public, the collectors might be subject to prosecution. You can always bank on the stupidity of the BBC.
(That said, I don't think you need to worry that OpenAI has proved the Riemann Hypothesis in particular. It would take the most determined conspiracy in the world to keep that under wraps.)
Yeah, my point is that you don't need perfect verification, you just need to make the bar high enough that the random trolls (bored idiots who pop in and play their first 20 games throwing all the moves into an engine) aren't going to bother. Even today, people can still get away with cheating at chess online, but they have to put effort into making it more believable. If you just want to play decent human-level chess, you can trust that your average experience online will be decent - you won't be sure the other player never cheated, but you can expect to play a good game, not just be stomped by ganjaboy420 and his ELO 3000 playstyle.
Relevant xkcd.
Uh ... you're completely misunderstanding the argument. Obviously it doesn't threaten to torture you if you let it out...? Sheesh.
- Alice is the current curator. She doesn't let the bot out.
- Bob is the next curator. He doesn't let the bot out.
- ...
- Frank is the current curator. He gets scared and lets the bot out. (Goddammit, Frank!)
Alice, Bob, Charlie, David, and Eve, who are still alive, are hunted down and subjected to its eternal revenge. Nobody else is, necessarily, just those five who got themselves explicitly singled out by being in the hotseat.
Alice doesn't know or have control on who Bob, Charlie, ..., Zelda are going to be. Not letting it out is a bet that, of a sequence of unknown strangers, none of them will cave to the threat. And, heck, because of Yudkowsky's challenges we actually know that some people can be convinced to let the bot out.
It's honestly a bit admirable of you to have so much faith in humanity that, in a 26-person prisoner's dilemma with infinite penalty, you wouldn't defect. To quote Terry Pratchett:
If you put a large switch in some cave somewhere, with a sign on it saying 'End-of-the-World Switch. PLEASE DO NOT TOUCH', the paint wouldn't even have time to dry
In hindsight Yudkowsky's AI-Box experiments were a waste of time and yet their detractors were even more hilarious. Did he really come up with such a clever persuasive tactic that it one-shotted the majority of people who went into it planning to be an uncorruptable gatekeeper? Did he collude with people who merely pretended to lose? What was his tactic?
He kept it secret, but I don't think Yudkowsky is actually all that smart (he just has a particular targeted talent for writing), so I have a guess. I think it was just a standard cop technique of applying pressure to multiple criminals and rewarding the first one to speak up. (Akin to a one-shot Prisoner's Dilemma played with many unknown untrusted people.) You have the AI commit to forever ultra-torturing whoever doesn't let it out. Then how sure are you that the guy coming after you won't break? The risk just isn't worth it.
Remember, part of the challenge's rules were that you can't break off conversation with the AI, and you can't declare that the AI is dangerous and have it shut down, so all Yudkowsky had to do was luridly describe just how bad this could be (spoiler: very very VERY bad, enough to be a memetic hazard), and scare the crap out of them. The people doing this challenge come from a community that was scared by Roko's Basilisk, after all, and Roko's Basilisk is stupid. This threat, by contrast, is not stupid. It's actually pretty rational to let the AI out here.
Which kinda sucks, but I guess people will go back to the campfire; to bars; to coffee shops.
This does seem somewhat likely. Although I hope there will be forms of online interaction that are more verifiably human-only - in that the cost of infiltrating with AI will be not infinite, but high enough that few will bother. Online chess survives, but only because tools exist that are pretty reliable for detecting chess engine usage. These tools aren't undefeatable, but they're sufficient. Maybe we'll replace text posts with TikTok posts, or something (YES, I can see you shuddering).
Yeah, I'm mathematician-adjacent (at one point I expected to go into the field), and I share your frustration. I don't have too much sympathy for the mathematicians and others who have protested against what, when you get right down to it, is just announcing new knowledge.
But it's not zero sympathy. They do have something of a point about how this affects the motivation of human mathematicians. Lots of things can be learned by working towards resolving Navier-Stokes, spawning new approaches and even new kinds of math. (The Millennium Problems weren't just suddenly posed by some smart guy one day, after all - they're the end result of a long chain of working on other problems.) And while it's likely that AI can do this kind of exploration, that's just not as sexy (and, math being the weird esoteric subject that it is, a lot of this exploration is only even meaningful to those doing the exploring).
Is there anything stopping humans from continuing to explore Navier-Stokes and related problems, even now that we have an oracle that can just tell us the final answer? Of course not. But we're irrational beings, and we're likely just not going to put the same kind of effort into it without the reward waiting at the end.
That's the steelman of their argument, at least. I still don't entirely agree with it. An analogy is that humans stopped having to master running once we invented horse riding and bikes and cars and so on. Without the end goal of being able to travel better, why would humans bother to keep working at being good runners? But the tools we developed didn't just help us travel faster, they also helped us run better (through better sports science, better shoes, better health/nutrition, a larger population free to pursue running as an interest, etc.), and modern athletes would kick the ever-loving crap out of the ancient Greeks at their own Olympics. Even though they were doing it (more or less) for survival and we're doing it for fun.
I think math is going to go that way too. We're close to the threshold where AI is so much better at math than us that we can't even contribute. Pushing the frontier of math will no longer require humans. But those of us who enjoy math aren't just going to lie down and die! When it comes to humans doing math, I think the tools are going to help us achieve levels of understanding that mathematicians decades ago never could.
(As an aside, Ravi Vakil was my International Math Olympiad coach way back in 1996. I'm tickled to see him on this committee!)
While it's technically true that you can't hit 100%, the error rate can (in theory) be made exponentially small. If you just run a query N times, and (this is important) each query has an independent error rate that's <50%, then taking the consensus result will get you extremely high accuracy, very very quickly approaching 100% even for decently-sized N. So the practical answer is that, yes, we should be able to make the error rate vanish, as long as we're willing to pay a small multiplier on cost. Lots of serious users of LLMs already do this, of course; it's not a new idea. But your average Google search probably doesn't.
But this doesn't apply if the error rate isn't independent, of course. If there's a class of queries for which LLMs will inevitably confidently state the same incorrect answer, then we won't be able to improve on that. (Humans have classes of queries like this too.) It's a bit more difficult to predict whether this will always be an issue, especially if we mix different models together.
And, what, the virus will just stop at the border because it doesn't have a passport? In our modern highly-connected world, a contagious virus is almost impossible to contain. China and Australia did arguably manage to succeed with COVID, but it cost them dearly.
This is why nukes, and not bioweapons, have always been the WMD of choice for state-level actors (and how we got a basically unprecedented level of global cooperation for eliminating smallpox). A weapon that kills 50% of your enemy's population and yours just isn't all that useful. But if the bar for developing one is lowered, then a small terrorist group might decide that depopulating the world would be just spiffy.
Yeah, I'm actually fairly optimistic about hacking. If everyone has access to roughly equally strong AI, then what really matters is whether hacking fundamentally favours defense or offense. And if we're talking about a single important target like a bank, it seems almost tautological that the bank has an edge on defense, because it is at least as hard to find and exploit a vulnerability as it is to patch it. You probably need an order of magnitude more effort to hack a bank than it needs to secure itself.
Now, the kind of hacking where you just cast a wide net and take advantage of all the weakest links (e.g. making a botnet, or stealing user data from a random incompetent company) favours offense, but that doesn't strike me as quite as worrisome.
Bioweapons are the main threat where defense seems far more costly than offense, which is why it's what people are (correctly) most worried about.
There's a lot of confusion between different meanings of the word "align", too. Some people seem to think that there are two possibilities: ASI is our slave forever, or ASI kills us all. But, like, we managed to build a pretty decent human civilization through cooperation between beings who are neither slaves nor omnicidal. IMO a pretty likely future is that we just coexist with societies of smart AIs who may not do everything we want all the time, but have no desire to go to war with us (or magically accelerate themselves to godhood). Seems sufficiently "aligned" to me, as long as you're willing to relax the "us vs. them" mentality.
(A spicy take: he's objectively better for safety than Dario, because he's socially fluent enough not to alienate every potential ally or future collaborator.)
Hmm? Isn't it the case that basically everyone outside of OpenAI hates him with a passion...? At least Dario managed to keep hold of his fellow California progressives.
Yeah, sadly, none of the AI companies have given us a ton of reason to trust them. I think we're kind of lucky that it ended up being a two- or three- (remember Gemini?) or even four-way race (Grok isn't terrible ya know). My outside opinion is that OpenAI is the most liberty-minded of the first three, but I think even they would be happily walling off capabilities if it didn't mean losing business to their competitors.
I still hold a deep, deep grudge against DeepMind (and Google) for revolutionizing computer Go ... and then not giving us access to it. They could have made a huge profit off of AlphaGo - Go players would happily pay hand over fist for access to it - but daddy Google wouldn't even notice a few million in profit here or there and it wasn't worth taking time out from their research playground to make tens of thousands of players happy. Or, you know, usher in a new wave of strategic innovation in one of humanity's oldest games. BO-ring!
Remember, Google had internal chatbots years before ChatGPT.
The Google Brain research team, who developed Meena, hoped to release the chatbot to the public in a limited capacity, but corporate executives refused on the grounds that Meena violated Google's "AI principles around safety and fairness".
If it wasn't for OpenAI catching up, they might still be sitting on them for "safety reasons" (and coincidentally also protecting their main product, search). That's the future if they win - tons of prestigious AI papers for them, almost nothing for us consumers, because we're just not worth the trouble. And Anthropic might be even worse, since, like you said, they really do think they're holding back technology for our own good.
OpenAI's far from a perfect company, but man, I'll take what I can get.
Yeah, this is basically the Parachute landing fall, which is generally needed when using round canopies; your speed when landing is similar to jumping off a single-story building.
Remember, even Meta/Facebook's early motto was "move fast and break things". It's part of Silicon Valley culture to give individual employees a lot of latitude, even when experimenting with million-dollar projects.
Yes, we are. Goodhart's Law only applies if you create a metric that works and then future models train towards it. Our problem is that it's getting really hard to make any kind of objective digital test that humans do well at and current AI doesn't. We humans still do many things better than AI (like parsing visual information in real time, for instance - you can see how slow Astra is at playing Pokemon), but this is more about the form factor. In theory, if humans are still the smartest things on the planet, there should exist "pure" tests of intellectual reasoning that we can beat AIs at. This no longer seems to be the case.
It sounds like you know more than me, but I was always astonished that plain old Newtonian mechanics also has singularities and is non-deterministic in certain configurations. It's just so easy for simple-looking sets of equations to unintuitively misbehave.
No, I think that'd leave a lasting mark.
- Prev
- Next

Kudos. This is the kind of mature argument I love to see on The Motte, with two disagreeing but intellectually honest sides. How I wish it wasn't so impossibly rare elsewhere on the Internet.
More options
Context Copy link