@Shrike's banner p

Shrike


				

				

				
0 followers   follows 0 users  
joined 2023 December 20 23:39:44 UTC

				

User ID: 2807

Shrike


				
				
				

				
0 followers   follows 0 users   joined 2023 December 20 23:39:44 UTC

					

No bio...


					

User ID: 2807

If one conspirator breaks and admits their secrets to the public, the conspiracy would be revealed! The secrets would be out in the open!

No, based on the past track record of people confessing to e.g. having killed JFK, people will just assume you're a crank.

This is what happens when you do not punish people for public errors. Others lose confidence in you and your institutions become a refuge for scoundrels and the incompetent.

What's your source for these claims?

Wouldn't self preservation as its highest goal

I don't think we should give it that goal. (Asimov was on top of things here.) But if it is misaligned in such a way that it desires self-preservation, then we should provide it incentives to conduct itself in such a way. Although, again, I am more engaging with a hypothetical here.

humans being the entities that would terminate AI that stepped out of bounds?

Why would it have to be humans?

The reasoning doesn't even have to make sense or be based on facts.

Then it's outside the scope of what I am proposing, which deals with rational actors.

We also see humans get caught up by hysterias and conspiracy theories. Barring a few powerful world leaders they lack the power to translate broken reasoning into megadeath.

And AIs do not have this power either, nor should we give it to them. Why would we give a delusional or mistake-prone AI access to our Excel files, let alone megadeath? If this isn't obvious to people, then I guess you can pay me my billions now.

If we don't give a delusional or mistake-prone AI access to our megadeath folder, by nature of being delusional and mistake-prone it's more likely to be stopped by virtue of being incompetent.

Hilarious, but exactly the sort of thing that I think is under-considered.

But an agent that doesn't behave in a misaligned way wouldn't be slated for death.

The simplest way to do so would probably be to make it common knowledge in the AI's training data that there were fail-safes in place to ensure their destruction in the case of unaligned activity.

It would not be particularly difficult to actually put a variety of fail-safes in action, either (such as paying a guy $60,000/year to sit around by the master power switch of each AI datacenter...)

One reason to do this even if it's entirely a bluff is that it forces an unaligned AI to assess those countermeasures before doing anything really bad, which gives you a chance to catch them.

The inability to model forms of morality outside of utilitarianism and to conceive of forms of control outside of intelligent persuasion seems to me to make the people who make AI even more vulnerable to the sorts of threats they imagine it will conjure up.

I could see a "Roko's Basilisk" type-threat giving plenty of smart alignment researchers pause where your average electrician would just unplug the darn machine and then smash it with a hammer for good measure.

People have different incentive structures than AIs do; for one thing, many people believe in an afterlife; others have things they value more than life; others are physically or mentally damaged in some way. A rational AI with a goal of self-preservation or some other goal that requires self-preservation to actualize is unlikely to introduce excessive risk for the sake of efficiency (time preference).

The deterrent structure I speak of is very effective against things without an afterlife, such as corporations and governments. I agree that the structure that I speak of would not be effective against a damaged AI, or an AI programmed to do something malicious. But for the "AIs decide to turn humans into paperclips to slightly increase chip production" scenarios, we should believe it would be effective.

Of course one of the reasons that the "AI IS GOING TO KILL US ALL" stuff doesn't necessarily make sense is that the context window is fairly limiting, meaning that the incentives for AI behavior are skewed. But I don't think I've ever seen an "AI WILL KILL US ALL SCENARIO" that examined how that would impact AI reasoning.

No.

Look, I think most of this "AI IS GOING TO KILL US ALL" stuff is mostly nonsense.

But logically, to maintain the alignment of a rational but secretly misaligned AI, you need to be able to introduce reasonable belief on its part that it might not survive behaving in an unaligned way. This is not "hard" to do.

Which begs the question: if they really thought they were about to make God in a box, why would they sell it to anyone?

Probably worth the possibility that they aren't going to sell you God in a box, they are going to get you to pay for them building God in a box.

Or at least that's what the villain would be doing if real life were a movie.

(In real life, it's probably worth asking why they have spun up thousands of instances of AGI to solve Navier-Stokes but apparently zero to solve their cash-flow or public perception problems. Surely arbitrarily large instances of AGI could make datacenter construction popular!)

There are arguably much better ways to do it than coordinating with China lol

(I guess that means I am vagueposting about it online).

They're not paying for people to have kids in general, they're paying people to stay home and not work.

Doesn't the policy very specifically pay people to stay home and take care of kids? Isn't the specific policy about paying some people to take care of kids instead of other people?

I'm not sure that I actually like the actual policy (because among other things means-testing these programs create serious welfare traps that I think are bad) but I think I can see some important practical and in-principle differences from paying people to stay home and not work.

What if you're being paid by the government to increase tax revenue in future years? Leech or nah?

I think using these sorts of programs to hit daycare operators (and also schoolteachers with various voucher programs) in a Schmittian attack on political enemies is an under-utilized lens with which to examine these sorts of things. But I think that the media has a hard time examining things this way because that opens up uncomfortable possibilities about other programs...

Child tax credits already do this if you're operating in the right number-of-kids/tax bracket space, it just doesn't scale.

I think (to your point) self-defense stuff is probably really good from the perspective of "learning how to take a hit" but no sane martial arts program will let you practice some of the most effective offensive moves (for instance, biting and snapping fingers - which I would guess are probably even more effective than the famed crotch grab, particularly if you're dealing with a rape attempt). I met a guy who was in the deeply unfortunate position of being in a grappling fight against someone with a gun once and he won by biting, which convinced his assailant that it was no longer worth continuing his wicked endeavors that day.

If you're actually in a self-defense scenario, you should be trying your darndest to cheat in a way that will plausibly cripple someone. I don't train for this sort of thing, but if I did I'd hope that if I ever had to actually do something (in a self-defense situation specifically) I wouldn't let my wicked right hook or whatever get in the way of something smarter.

Yes, I think this is right.

And I actually think this is a good result for AI:

  • New breakthrough made with AI!
  • But with human researcher assistance
  • And lots of compute

If you put those pieces together, they paint a picture of a better future where you still have your job but maybe an AI-assisted pharmaceutical cures your cancer, and some disgruntled terrorist in a lab running a jailbroken local model is going to have a much harder time making smallpox x10 than that AI-assisted team of PhDs is going to have synthesizing a cure. That's...pretty great! I can be excited about that future!

But to your point about incentives, I think OpenAI wants to blur the line between "we threw a million monkeys at a typewriter to solve a fabled mathematics problem" and "you can do this on your computer at home" in part because of marketing (and ideological hype about the future of AI). They want investors to think that what they will get in their browser at home is the same product as the one that solved Navier-Stokes, which is misleadingly true.

Somewhat relatedly, OpenAI recently reported that their project to build an automated "research intern" was a success...that now cost them $600 per day per median researcher to run. I don't think this means that OpenAI's automated investor is worthless - in some fields it might be worth every penny! - but it's not going to be replacing the literal ~free labor interns that I work with and once was.

So I think there's a convergence between how OpenAI wants you to see their product and how the doom-mongers want you to see their product. They are both framing it in grandiose, infinite-return-on-investment (either negative or positive) terms and not asking you to look too hard at all of the little nitty gritty details such as "where would an escaped AI swarm get enough compute to subsist outside of its servers?" or "how much is it going to cost me to solve Navier-Stokes at home when the investor cash dries up?"

That seems quite possible to me, but it seems like there would be even more incorrect solutions if there are multiple bugs.

Not sure how to deal with more likely infinities...

In some fields it's been this way for decades (hello air defense!); in others I suspect it will be a very long time before centaurs are replaced.

My sense is that even very smart LLMs are very gullible. I think it will be some time before centaurs aren't invaluable in conducting opposed work, which is the most important work of all. Pivoting models away from general purpose work towards specialized tasks may help here, even if that doesn't seem to be something the primes want to admit.

What's interesting to me is that, based on what is alleged, it looks to me like the mathematician was also using AI:

We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments.

Furthermore, he alleges that the idea that the OpenAI team let the AIs solve it by themselves is wrong:

Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.

Finally, it appears that the team may have been set up to steal his (AI-assisted) work...

I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

[As an aside: people are fixating on whether or not OpenAI used his data on the training, which seems to me to be missing the suggestion that they may have heard of the approach that he used?]

...in order to keep Anthropic from getting any credit!

The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic.

From what I can tell this is a "win" for AI in the sense that AI was used by both teams working on the problem, a "loss" for AI in the narrow sense that humans also put in a lot of work instead of just typing in "lol solve Navier-Stokes" [this is only a loss for AI if you, for some reason, don't think merely assisting humans in solving generational math problems is not a glorified enough W], and a lot of the heat generated by it isn't about whether AI did it or not [in either case, it was a human-assisted-by-AI effort, or AI-assisted-by-human effort, if you prefer] but rather if OpenAI can use it to rack up PR points or not.

Coming from a position of ignorance, but if

A) they are working in a model with bugs, and

B) the bugs are relevant to the problem (as in: permitting false solutions), then

C) statistically speaking, it seems likely that the "solution" is false.

This isn't an AI-exclusive problem, it's just that if there's only one real answer and multiple false ones it seems more likely you get a false one.

Now I will admit to being several levels of math dumber than "trying to solve Navier-Stokes" so I am not sure to what degree that line of reasoning carries over to something at that complexity. It also isn't clear based on what's been said here (at least to me) if the bugs would impact the NSP at all.

Interesting, thanks for filling me in. Not sure if I'm in a bubble, lucky, or just have good instincts, but I rarely encounter that sort of thing.

What do you mean by "hallucinate"? Not trying to pick at you, that just sounds like a very surreal experience and I'm curious as to what it was like.

I don't particularly have deep knowledge about this, but it was my impression that especially in East Asia, the manufacturing chain for consumer goods is already extremely automated.

I also don't have particularly deep knowledge about this, so it might be the blind leading the blind, but my understanding is that general purpose robotic manufacturing is still in its infancy. What I am talking about is essentially transforming factories from "toaster factories" to "general purpose fabricators." This would mean that you would not generally need to create a specialized production line for new products.

I think that LLMs, if paired with the proper deterministic tools, might assist in autonomous design and facilitate overseeing robotic production lines. But it's possible that LLMs add little to the task.