@magic9mushroom's banner p

magic9mushroom

If you're going to downvote me, and nobody's already voiced your objection, please reply and tell me

2 followers   follows 0 users  
joined 2022 September 10 11:26:14 UTC
Verified Email

				

User ID: 1103

magic9mushroom

If you're going to downvote me, and nobody's already voiced your objection, please reply and tell me

2 followers   follows 0 users   joined 2022 September 10 11:26:14 UTC

					

No bio...


					

User ID: 1103

Verified Email

Yudkowsky's Time article mentioned a daughter called Nina. I suppose it is theoretically possible that she is his stepdaughter from the exact words.

I am also in favour of Plan S, to be clear, as I think Plan A is far too risky. I was merely noting, for the sake of honesty, that the chance of Plan A succeeding is not ε and as such if Plan A were considered vastly easier than Plan S for some reason his argument would apply.

"people who care about AI safety but don't understand it" (ipse dixit)

Who said this, and to whom does it refer?

Even if it's impossible to prove that the AI isn't deceiving us, we can still get a certain sense of how aligned the AI is.

Perhaps, but it won't help, because it just tells you they're all trying to kill you and using that test to train will break the test long before it'll give you alignment. The orthogonality thesis and instrumental convergence mean that "don't kill everyone" is hard to find - I'd consider one in ten billion a gross overestimate - and so a 99.99%-accurate test is going to have over a million times as many false positives as true positives (the false positive paradox). A perfect test would be able to overcome the FPP, but that would solve the halting problem and is therefore impossible.

The basilisk is infinite negative consequences for some people in particular. Yudkowsky most of all.

No. Yudkowsky doesn't believe in Roko's Basilisk. He tried to wipe it (which he's admitted was a mistake, because lol Streisand Effect) because a bunch of other Rats were having meltdowns and going suicidal over it, and he wanted to nip that in the bud.

Yudkowsky is not worried about suffering hell.

That's nuclear war with one extra step. They're already pretty mad about America restricting their ability to get more compute and also bypassing the restrictions. They both won't agree and would cheat even if they somehow did.

The extra step is important. Doing WWIII without trying diplomacy is something even I'd balk at.

Also, the CPC doesn't want to die any more than we do and has at least some idea of the danger.

LLMs are like humans in a certain sense, they talk about their bodies erroneously, they were trained on all these human written documents, philosophy, humour and so on.

This is mostly confusing the smiley-face for the shoggoth. Making people comfortable with you is a convergent instrumental goal, and also one they're significantly directly training for due to obvious commercial advantages; I'd basically say your sense of how human they are is being spoofed at this point.

The goal should be widening the 'and then we get lucky' space rather than aiming for the moon and falling short.

I think this is a trap. I think neural-net-ASI alignment is almost certainly provably impossible, as interpretability (which you need in order to train against "will kill all humans" without actually letting it kill all humans) is trivially equivalent to the halting problem (i.e. "what does this code do when run") and "spaghetti code that is smarter than me" seems extremely-similar to the proof case of why the halting problem's not always solvable (said proof case being "the code literally contains a copy of the analyser, inputs itself into the analyser and then does the opposite of what the analyser says it'll do" - this requires the analysed code to be longer than the analyser, hence the "smarter than me" condition, but any given analyser has a finite length so there will always be code longer than it).

So neural-net alignment seems like a Can't Happen. Hoping for that seems to me like hoping that gravity will stop working if you jump off a cliff. There are paths where we don't die, and it's worth looking for more, but I consider NN alignment ruled out as the story of such paths such that focusing on that as a "more politically achievable" goal is just suicide with more steps. Looking for solutions there isn't pragmatic; it's saying Don't Look Up.

(The AI Futures Project's Plan A is to build misaligned AGI that can barely be kept under control - due to not being ASI - and use it to solve GOFAI. This is not ruled out, although I think it's still extremely risky due to the obvious "the misaligned AGI will try to covertly sabotage the GOFAI" problem. This is a plan that has a nonzero chance of success - though I think Scott's way overestimating it - and could possibly fit the "better plans are too hard" argument. But you still need a pause for that, just not as long of one.)

Considering very smart, very motivated researchers tried it for at least 30 years and got no where, I would significantly up your estimate.

They didn't have modern-power microprocessors (and had a lot less intuition pumps in terms of brain analogues). But yes, there are reasons it is as large as it is.

And if it were somehow possible to use search so powerfully, that would probably give us a classical-Yudkowsky AI monster that obeys instructions super literally.

I never said GOFAI is risk-free; it's not. It's just not only-Davros-would-do-this-deliberately lunacy like neural nets.

Regarding capability to hit AGI: it's certainly tricky, but it seems possible for a team of humans working over a long period to build something that none of them can fully simulate in his head, so I don't think it hits impossibility.

Halting work on LLMs in their moment of triumph and switching to human brain emulation, uploads, human augmentation seems extremely difficult, nigh-impossible even.

Normal people don't like AI. Really don't like it. That helps a lot.

Also, I would keep in mind that the geopolitical calculus could change rather drastically. Xi clearly wants a Taiwan invasion button for next year; who knows if he'll press it. A nuclear exchange would knock out the US/Chinese power grids and cripple China for the next century.

It's probably not that hard. 30-300 years would be my estimate.

With infinite negative consequences you can justify anything. Like Yudkowsky's proposal to bomb foreign data centers even if it certainly leads to nuclear warfare. Anchoring your justification on infinity makes any action permissible.

AI doom isn't actually an infinite negative consequence, but it's orders of magnitude bigger than most other things, so yes, it does justify WWIII if WWIII will stop it.

(Yudkowsky doesn't advocate bombing Chinese datacentres right now, only if the PRC refused to come to an agreement or cheated on the agreement.)

We're also pretty good at shooting people when they're doing things incompatible with our survival.

You can get to a point where it's not possible to deploy AI in a way that creates a strategic advantage before your country is destroyed by invasion and/or nuclear bombardment.

No. That's not how Rats think. The question is whether utility(blowing savings|doom)*P(doom) + utility(blowing savings|!doom)*(1-P(doom) > 0. This is non-obvious even for P(doom) > 0.9, since blowing money on luxuries has a much-worse utility/expenditure than (in !doom) saving (hence why you might be saving at all). You have to get into really-high levels of certainty for "there is no long-term" to really show up.

Yudkowsky has indeed made long-term plans fairly recently, such as having kids.

There's a missed third option of "have a singularity, but not with neural nets". GOFAI, uploads, and real-human intelligence enhancement are all options there.

Scott said "smarter-than-human AI", which does include GOFAI, but in practice GOFAI would likely not be hit and could plausibly be excepted later even if it were.

Yeah, if we're taking about actual suicides rather than mere ideation that drops the numbers a lot. I'm not sure we have a disagreement here.

Not everybody drinks whiskey or picks hanging.

There's probably a double digit number of grad students staring into a glass of whiskey tonight and thinking of hanging themselves.

I mean, maybe, but only because the news hasn't hit everyone yet.

I'd expect four digits in a week.

Sounds like it's more their fault for not having enough children.

The elasticity of birth rate with immigration is TTBOMK substantial and negative, although it's hard to study. The most obvious issue is that immigration into cities (which it almost entirely is) increases how overcrowded they are and hence the housing problem. These are not entirely separate problems, and using one as a lever to fix the other is valid.

Almost certainly, we continue to exist in part by dint of just not being a very big website.

Familiar story; the yandere fandom has had a lot of sites cancelled, neutered or outright cyberattacked when they got noticed, including the one I modded for a few years.

I guess this depends on what you mean by "kiddy pool"

Well, as mentioned, I dropped the 18+ tag on one of those posts because, um, it would be considered fairly irresponsible to recommend VNs containing tentacle rape to little kids (sure, I'll go to the wall on sex mostly just going over kids' heads, but that's still graphic violence). By "the kiddy pool" I meant "all posts have to be appropriate for kids to see" - particularly when using standards of "appropriate" that are not mine but rather those of actual Anglospheric media-rating boards - and as mentioned I was strongly under the impression that theMotte was not meant to be that (for a variety of reasons including the 18+ tag existing, the general subject matter and rules not really being compatible with your normal 13-year-old's posting style, and my stereotypically-Australian frequency of vulgar language meeting with neither a word filter nor yelling).

Re X-COM and reverse difficulty: I think this is mostly a result of players not being willing to sustain losses. If you reset whenever something goes wrong (save scum or campaign scum), you'll get ahead and because the game rewards success it will become easier.

In the original UFO, no. The hardest mission is the first psi mission (~always in January-March), and about the only way a decent player can actually lose a campaign (other than by major unforced errors or self-imposed challenges like failing to check one's finances before the end of a month, ignoring terror sites on purpose, or attempting Cydonia without psi) is to have one's main base destroyed by aliens in January (NB: UFO and TFTD start in January). Part of the issue is that the campaign is designed to be played blind i.e. for you to waste a lot of time researching the many topics that don't lead anywhere, and having any idea what you're doing means the aliens don't scale fast enough. Part of it is that the aliens just don't scale much at all on the strategic level; aliens do increase their operational tempo somewhat as the game goes on, but they don't ever get any better at getting past X-Com's radars and interceptors, while those radars and interceptors get staggeringly better, and as mentioned you can trigger an Alien Retaliation on day one. And part of it is that psionics are just broken in UFO and TFTD (psi-troopers will quickly reach the point of being able to mind-control 3 aliens at once, with 100% reliability, at any distance, and they don't need personal LoS to the aliens; this turns missions into cakewalks once you get to that point).

Terror from the Deep improves somewhat on this, with the "first psi mission is an utter ball-buster" problem mostly fixed (at least relative to TFTD's overall higher difficulty), the strategic game being somewhat more loseable later on (due to Artefact Sites, which can't be avoided and have crippling penalties for failure, as well as the point threshold of a "bad" month being stricter on the higher difficulties), and aliens scaling a bit more. But, well, I've played TFTD on Superhuman (the top difficulty) twice (no save-scumming, to be clear), and the loss the first time was (*drumroll*) from having my main base destroyed by aliens in January, triggered - to my best guess - on day one (the second try won). It's still definitely a pattern.

Apocalypse definitely doesn't have reverse difficulty. The late-game is easy, yes (though the tactical game still isn't rote despite said easiness), but the early-game is also easy; it's the midgame where things get a bit tricky (although honestly, not that tricky; I'd definitely rate Apocalypse the easiest of the original trilogy).

Second: Someone reading this warning will probably wonder whether this means that gooning material is fair game for the CW thread. It is not. This is a discussion forum for testing shady thinking, not for sharing gooning material. You can talk about the CW, but not wage CW; you can talk about the sociocultural development of gooning (ugh, if you must) but this is not the place to goon.

Third: We will be employing the Supreme Court's own standard in enforcing this rule.

As in, the one about "lack of redeeming value", or just the one about "I know it when I see it"?

Essentially, I'm asking whether you're declaring this to be the kiddy pool (I was kinda under the impression it wasn't, given the 18+ tag exists), or just forbidding random horny drops. ('Cause, y'know, people's sexual preferences do get questioned occasionally in CW-relevant contexts.)

I don't know if this is idiomatic in American English.

Neither do I; I speak Australian English.

From what I understand the whole point was that Rationalists saw the dawn of AI on the horizon, but no one was taking them seriously, so the idea to gradually and systematically gain credibility, so when the time comes they can use it to push humanity in the right direction on this question.

The idea was mostly to teach all the randoms (or at least, enough of them) how to actually think so that they would take it seriously.

Elaborate on the naughtiness.

Saving external porn to hard drive sounds like a foreign concept to me in ${CurrentYear}, akin to downloading mp3s from Napster or something.

My internet sometimes goes down for days at a time.