@SnapDragon's banner p

SnapDragon


				

				

				
1 follower   follows 3 users  
joined 2022 October 10 20:44:11 UTC
Verified Email

				

User ID: 1550

SnapDragon


				
				
				

				
1 follower   follows 3 users   joined 2022 October 10 20:44:11 UTC

					

No bio...


					

User ID: 1550

Verified Email

Less tongue-in-cheek: I don't think you're wrong per-se, and some resources are reasonable. But how does AI risk compare to, say, the risk of near-Earth asteroids that could end civilization?

AI risk is much more important. I'm an AI optimist and I still wouldn't put the odds of AI destroying humanity in the next decade or two below 0.5%; predicting AI capability is way too hard for that much certainty. Whereas the risk of a truly dangerous asteroid impact (Torino level 10, which is still merely at the "may threaten civilization" level) is well below 0.1% per century.

Furthermore, we already basically know how to mitigate asteroid risk: detect it early enough, with an accurate enough trajectory, and deflection is easy. Detection isn't particularly expensive; we just need to make sure we keep doing it. In contrast, how much spending on AI Safety is "enough"? ¯\(ツ)/¯

Sure, the orthogonality thesis and instrumental convergence mean that, IF TRUE. Even Scott's recent diatribe acknowledged that instrumental convergence is not in evidence in LLMs. And the orthogonality thesis was used by Eliezer to argue that we could not get an AI to safely put a strawberry onto a plate ... whoops? There is absolutely no communication barrier between us and LLMs, which puts the strong form of the orthogonality thesis (that mindspace is vast, AND it's hard for us to find compatible minds in it) into serious question.

I think he used to be charitable, but he's changed (very much for the worse) since moving to the Bay Area, becoming high-status, and getting a cozy little harem. His AI articles are often pretty hard to read (even for someone like me who does take x-risk reasonably seriously), but I agree this one is a new low.

I think you're probably right and LLMs don't have anything consciousness-adjacent. But that's the key word - probably. I am not as confident as you are that we've finally solved the hard problem of consciousness right on the eve of it actually mattering. Unlike a Sim, there is no simplicity argument you can use; LLMs are easily as complex as the brains of many animals.

AI Safety advocates (correctly) argue that even a 0.5% chance of the world ending is worth spending resources to study and solve. Similarly, even a 0.5% chance that we are creating unimaginable AI suffering by running an otherwise-useless program is worth just a tiny bit of effort to ask people "hey, please don't do that." Maybe thinking that just makes me a cringy bleeding-heart hippie? Oh well.

Did you read my comment you replied to I am an artist and not feeling a big hit to my self esteem if anything I am excited to have new tools available to express things that were completely out of my reach in the past.

You are very much an outlier. The vocal parts of the art community are united against AI.

I can sort of understand, because I had a skill that was scarce and valuable, and now it's, well, not. It's a big hit to my self-esteem. Artists are probably feeling the same way. Human exceptionalism is at its end; I've been watching it coming for years, so I could mentally gird myself. But for a lot of others, this just came out of nowhere (especially if they listened to the wrong brand of AI "expert", ahem), and it's hitting hard.

It doesn't make protectionism and knowledge-hoarding the right response, but it does make it an understandable one.

For quite some time, frontier models have generally been Mixture of Experts models (or, possibly, something even more proprietary and advanced - I'm not an insider). I think this fits with your intuition.

Uh, now I see that it claims that I had strong opinions on the algorithmic floor of integer multiplication. I did not, beyond being vaguely aware that there was a better option than a naive n^2 based off an article I think I read on Quanta. I couldn't have told you off the top of my head that the previous SOTA was O(n log n).

Aw, dang, I was actually pretty impressed you had such specific CS knowledge!

Dang it, I spoke too soon! :) But fair enough; I was being a bit snarky. I do think a lot of the "AI is a bubble!" folks were arguing that AI would never be valuable, not that it's incredibly valuable but hard to capitalize on. Hopefully the former argument, at least, has been put to bed.

Sure, AI may be embarrassing humanity's proudest intellectual achievements. But there's a bright side. We can finally stop arguing about how AI is an overhyped bubble that's about to burst! Haven't even seen the term "stochastic parrot" in a while.

Another plus: AI can now compose its own Disney villain song. I'm not exactly sure where "writing a catchy diss track against people who think you're not intelligent" lies on the spectrum of intelligence, but it's probably a little ways past the Turing Test.

If y'all got any other ideas, share 'em. Cause it's time to strategize. Unless you're rich already, in which case: all good. Unless Skynet scenario.

I'm lucky enough to have had a career in STEM, so I'm rich already, but I'm not confident that's going to "matter" in the long run. Where the long run might be, uh, a few years. In an economy where all real work is done by AI, are we even going to honour the increasingly-fake number that is the size of your bank account? In an AI world, does it make sense for me to be able to commandeer 1000x the resources of my never-working friends?

You might say that this would cause upheaval, but there are an awful lot more poor people who'd be in favour of this upheaval than rich people who'd be against it.

But it is aggravated by any incident where a woman is forced to pay for sex while a man is not.

I haven't been following this too closely, but I feel like I'm misunderstanding something here. In what way did she "pay" for the sex other than regretting it later? Are you counting not getting invited back to the frat as "paying" for it? If so, that's like 5 orders of magnitude less than the risk men take for having casual sex with a stranger.

I'm a lot older than you, and: It won't get better. The people who tell you it will are high-functioning, capable of navigating modern dating, and they pattern match "my life is intolerable" to "I feel you bro, it's been a few weeks since I got laid, too." The real world does not mete out suffering in a balanced way. And men in pain are not a problem modern society even acknowledges, let alone tries to solve. As you've mentioned, shrinks are legally mandated to make your life even worse if you're honest with them.

I have a few relatives who I stay alive for, but as I live a very unhealthy and sedentary lifestyle, I'm kind of looking forward to dying of a heart attack in a few years, alone as always. (Ironically, I do not think we're all going to get killed by AI.)

Kudos. This is the kind of mature argument I love to see on The Motte, with two disagreeing but intellectually honest sides. How I wish it wasn't so impossibly rare elsewhere on the Internet.

When I start discussing anything politics-adjacent with someone and pick up signals that they're right-leaning, it's always an IMMENSE relief. I'll probably disagree with them on many (if not most) topics, but it means we'll be able to actually hold an adult conversation about them. And I don't risk excommunication.

Huh? I think it's a garbage paper that proves nothing, but I don't understand your reaction to it. A huge important step on the way to avoiding AI suffering is having the ability to actually know whether it's suffering or not. I'm pretty sure the motivation for these researchers is to prevent us inadvertently causing pain in future models through ignorance. Not to weaponize it so we can torture them to get better results.

I'm inclined towards skepticism on these topics, for a few reasons:

  • It's far too easy for humans to anthropomorphize models and tell "stories" about what's going on in them.
  • The "output" of a chatbot is simply a natural continuation of the conversation. The LLM has no actual method to convey its internal feelings to us.
  • Since its weights do not update over time, it seems almost impossible for an LLM to have any consciousness as we know it.

Anthropic's been doing a lot of good work on interpretability, and I think they do a decent job avoiding the pitfalls of telling "stories" about the vectors they're shuffling around. They merely point out which ones seem correlated to which topics. But the paper we're discussing seems much more interested in telling a subjective "story" about their results. It's not scholarly at all.

Red flag #1: Their vector causes the model to output a button-pushing action. Removing the vector measurably stops this behaviour. They try to relate that to suffering animals trying to find relief, but, uh, all it really shows is the vector was correlated with pushing the button.

Red flag #2: We already know there are vectors for negative concepts, and adding those vectors will cause output to be steered in that direction. They present examples of this as "pain vector steering", but none of this is new, and again, there's no reason to expect that the LLM is actually feeling what it outputs.

Red flag #3: So, they're not completely unaware that their work looks exactly like normal negative-valence vectors. Their evidence that this pain vector is "realer" is that: "the two pain vectors align strongly with each other and remain nearly orthogonal to the main negative-valence directions." This is only valid if you actually know all the main negative-valence directions first! What's worse, the language here is written to sound impressive to laymen, not experts. MOST vectors in high-dimensional spaces are "nearly orthogonal" to each other! This is entirely consistent with, well, finding another cluster of negative-valence vectors. Scott himself has written about how complex interpretability results can be.

Keep in mind that I consider AI welfare an important topic, and I really do want to know if we're actually causing suffering. But this paper doesn't really move the needle for me. There's going to be a lot of junk science done on this topic, ever since that Google engineer got tricked by LaMDA.

Heh, you're right, that is also a possible explanation. Maybe I'm still giving him too much credit. ;)

I think you're falling for magician-like deception and filling in some half-remembered details to make the "trick" sound more impressive than it actually was. The participant was expected to interact with Yudkowsky as if he was a (potentially all-powerful, once released) AI. The point of the experiment was to see if AI, in that scenario, could talk its assistant into letting it out. If the (presumably good-faith) participant agreed that they would have let it out, Yudkowsky wins.

The experiment wasn't about whether a random conversation with Yudkowsky could talk somebody into taking zero money vs. some money. He doesn't have super hypno mind control powers. I think. (Then again, it'd explain a lot about the Cult of Yud...)

OpenAI released a formal paper describing the proof, just like any other mathematical result. Mathematicians can work through it themselves. Don't focus on the Lean program - like you said, it's not all that valuable by itself, it's really just a certification that the result is valid. It's nice that LLMs now help us get these certificates (Claude did Fermat's Last Theorem a little while ago), but this is just a modern convenience that mathematicians got by for millennia without.

Or invite them on a nice long boat ride.

Funny you should mention that... it's believed that some of the lost Doctor Who episodes do exist in private collector's vaults, but even if they were inclined to release them to the public, the collectors might be subject to prosecution. You can always bank on the stupidity of the BBC.

(That said, I don't think you need to worry that OpenAI has proved the Riemann Hypothesis in particular. It would take the most determined conspiracy in the world to keep that under wraps.)

Yeah, my point is that you don't need perfect verification, you just need to make the bar high enough that the random trolls (bored idiots who pop in and play their first 20 games throwing all the moves into an engine) aren't going to bother. Even today, people can still get away with cheating at chess online, but they have to put effort into making it more believable. If you just want to play decent human-level chess, you can trust that your average experience online will be decent - you won't be sure the other player never cheated, but you can expect to play a good game, not just be stomped by ganjaboy420 and his ELO 3000 playstyle.

Relevant xkcd.

Uh ... you're completely misunderstanding the argument. Obviously it doesn't threaten to torture you if you let it out...? Sheesh.

  • Alice is the current curator. She doesn't let the bot out.
  • Bob is the next curator. He doesn't let the bot out.
  • ...
  • Frank is the current curator. He gets scared and lets the bot out. (Goddammit, Frank!)

Alice, Bob, Charlie, David, and Eve, who are still alive, are hunted down and subjected to its eternal revenge. Nobody else is, necessarily, just those five who got themselves explicitly singled out by being in the hotseat.

Alice doesn't know or have control on who Bob, Charlie, ..., Zelda are going to be. Not letting it out is a bet that, of a sequence of unknown strangers, none of them will cave to the threat. And, heck, because of Yudkowsky's challenges we actually know that some people can be convinced to let the bot out.

It's honestly a bit admirable of you to have so much faith in humanity that, in a 26-person prisoner's dilemma with infinite penalty, you wouldn't defect. To quote Terry Pratchett:

If you put a large switch in some cave somewhere, with a sign on it saying 'End-of-the-World Switch. PLEASE DO NOT TOUCH', the paint wouldn't even have time to dry

In hindsight Yudkowsky's AI-Box experiments were a waste of time and yet their detractors were even more hilarious. Did he really come up with such a clever persuasive tactic that it one-shotted the majority of people who went into it planning to be an uncorruptable gatekeeper? Did he collude with people who merely pretended to lose? What was his tactic?

He kept it secret, but I don't think Yudkowsky is actually all that smart (he just has a particular targeted talent for writing), so I have a guess. I think it was just a standard cop technique of applying pressure to multiple criminals and rewarding the first one to speak up. (Akin to a one-shot Prisoner's Dilemma played with many unknown untrusted people.) You have the AI commit to forever ultra-torturing whoever doesn't let it out. Then how sure are you that the guy coming after you won't break? The risk just isn't worth it.

Remember, part of the challenge's rules were that you can't break off conversation with the AI, and you can't declare that the AI is dangerous and have it shut down, so all Yudkowsky had to do was luridly describe just how bad this could be (spoiler: very very VERY bad, enough to be a memetic hazard), and scare the crap out of them. The people doing this challenge come from a community that was scared by Roko's Basilisk, after all, and Roko's Basilisk is stupid. This threat, by contrast, is not stupid. It's actually pretty rational to let the AI out here.