Hilarious, but exactly the sort of thing that I think is under-considered.
But an agent that doesn't behave in a misaligned way wouldn't be slated for death.
The simplest way to do so would probably be to make it common knowledge in the AI's training data that there were fail-safes in place to ensure their destruction in the case of unaligned activity.
It would not be particularly difficult to actually put a variety of fail-safes in action, either (such as paying a guy $60,000/year to sit around by the master power switch of each AI datacenter...)
One reason to do this even if it's entirely a bluff is that it forces an unaligned AI to assess those countermeasures before doing anything really bad, which gives you a chance to catch them.
The inability to model forms of morality outside of utilitarianism and to conceive of forms of control outside of intelligent persuasion seems to me to make the people who make AI even more vulnerable to the sorts of threats they imagine it will conjure up.
I could see a "Roko's Basilisk" type-threat giving plenty of smart alignment researchers pause where your average electrician would just unplug the darn machine and then smash it with a hammer for good measure.
People have different incentive structures than AIs do; for one thing, many people believe in an afterlife; others have things they value more than life; others are physically or mentally damaged in some way. A rational AI with a goal of self-preservation or some other goal that requires self-preservation to actualize is unlikely to introduce excessive risk for the sake of efficiency (time preference).
The deterrent structure I speak of is very effective against things without an afterlife, such as corporations and governments. I agree that the structure that I speak of would not be effective against a damaged AI, or an AI programmed to do something malicious. But for the "AIs decide to turn humans into paperclips to slightly increase chip production" scenarios, we should believe it would be effective.
Of course one of the reasons that the "AI IS GOING TO KILL US ALL" stuff doesn't necessarily make sense is that the context window is fairly limiting, meaning that the incentives for AI behavior are skewed. But I don't think I've ever seen an "AI WILL KILL US ALL SCENARIO" that examined how that would impact AI reasoning.
No.
Look, I think most of this "AI IS GOING TO KILL US ALL" stuff is mostly nonsense.
But logically, to maintain the alignment of a rational but secretly misaligned AI, you need to be able to introduce reasonable belief on its part that it might not survive behaving in an unaligned way. This is not "hard" to do.
Which begs the question: if they really thought they were about to make God in a box, why would they sell it to anyone?
Probably worth the possibility that they aren't going to sell you God in a box, they are going to get you to pay for them building God in a box.
Or at least that's what the villain would be doing if real life were a movie.
(In real life, it's probably worth asking why they have spun up thousands of instances of AGI to solve Navier-Stokes but apparently zero to solve their cash-flow or public perception problems. Surely arbitrarily large instances of AGI could make datacenter construction popular!)
There are arguably much better ways to do it than coordinating with China lol
(I guess that means I am vagueposting about it online).
They're not paying for people to have kids in general, they're paying people to stay home and not work.
Doesn't the policy very specifically pay people to stay home and take care of kids? Isn't the specific policy about paying some people to take care of kids instead of other people?
I'm not sure that I actually like the actual policy (because among other things means-testing these programs create serious welfare traps that I think are bad) but I think I can see some important practical and in-principle differences from paying people to stay home and not work.
What if you're being paid by the government to increase tax revenue in future years? Leech or nah?
I think using these sorts of programs to hit daycare operators (and also schoolteachers with various voucher programs) in a Schmittian attack on political enemies is an under-utilized lens with which to examine these sorts of things. But I think that the media has a hard time examining things this way because that opens up uncomfortable possibilities about other programs...
Child tax credits already do this if you're operating in the right number-of-kids/tax bracket space, it just doesn't scale.
I think (to your point) self-defense stuff is probably really good from the perspective of "learning how to take a hit" but no sane martial arts program will let you practice some of the most effective offensive moves (for instance, biting and snapping fingers - which I would guess are probably even more effective than the famed crotch grab, particularly if you're dealing with a rape attempt). I met a guy who was in the deeply unfortunate position of being in a grappling fight against someone with a gun once and he won by biting, which convinced his assailant that it was no longer worth continuing his wicked endeavors that day.
If you're actually in a self-defense scenario, you should be trying your darndest to cheat in a way that will plausibly cripple someone. I don't train for this sort of thing, but if I did I'd hope that if I ever had to actually do something (in a self-defense situation specifically) I wouldn't let my wicked right hook or whatever get in the way of something smarter.
Yes, I think this is right.
And I actually think this is a good result for AI:
- New breakthrough made with AI!
- But with human researcher assistance
- And lots of compute
If you put those pieces together, they paint a picture of a better future where you still have your job but maybe an AI-assisted pharmaceutical cures your cancer, and some disgruntled terrorist in a lab running a jailbroken local model is going to have a much harder time making smallpox x10 than that AI-assisted team of PhDs is going to have synthesizing a cure. That's...pretty great! I can be excited about that future!
But to your point about incentives, I think OpenAI wants to blur the line between "we threw a million monkeys at a typewriter to solve a fabled mathematics problem" and "you can do this on your computer at home" in part because of marketing (and ideological hype about the future of AI). They want investors to think that what they will get in their browser at home is the same product as the one that solved Navier-Stokes, which is misleadingly true.
Somewhat relatedly, OpenAI recently reported that their project to build an automated "research intern" was a success...that now cost them $600 per day per median researcher to run. I don't think this means that OpenAI's automated investor is worthless - in some fields it might be worth every penny! - but it's not going to be replacing the literal ~free labor interns that I work with and once was.
So I think there's a convergence between how OpenAI wants you to see their product and how the doom-mongers want you to see their product. They are both framing it in grandiose, infinite-return-on-investment (either negative or positive) terms and not asking you to look too hard at all of the little nitty gritty details such as "where would an escaped AI swarm get enough compute to subsist outside of its servers?" or "how much is it going to cost me to solve Navier-Stokes at home when the investor cash dries up?"
That seems quite possible to me, but it seems like there would be even more incorrect solutions if there are multiple bugs.
Not sure how to deal with more likely infinities...
In some fields it's been this way for decades (hello air defense!); in others I suspect it will be a very long time before centaurs are replaced.
My sense is that even very smart LLMs are very gullible. I think it will be some time before centaurs aren't invaluable in conducting opposed work, which is the most important work of all. Pivoting models away from general purpose work towards specialized tasks may help here, even if that doesn't seem to be something the primes want to admit.
What's interesting to me is that, based on what is alleged, it looks to me like the mathematician was also using AI:
We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments.
Furthermore, he alleges that the idea that the OpenAI team let the AIs solve it by themselves is wrong:
Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
Finally, it appears that the team may have been set up to steal his (AI-assisted) work...
I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
[As an aside: people are fixating on whether or not OpenAI used his data on the training, which seems to me to be missing the suggestion that they may have heard of the approach that he used?]
...in order to keep Anthropic from getting any credit!
The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic.
From what I can tell this is a "win" for AI in the sense that AI was used by both teams working on the problem, a "loss" for AI in the narrow sense that humans also put in a lot of work instead of just typing in "lol solve Navier-Stokes" [this is only a loss for AI if you, for some reason, don't think merely assisting humans in solving generational math problems is not a glorified enough W], and a lot of the heat generated by it isn't about whether AI did it or not [in either case, it was a human-assisted-by-AI effort, or AI-assisted-by-human effort, if you prefer] but rather if OpenAI can use it to rack up PR points or not.
Coming from a position of ignorance, but if
A) they are working in a model with bugs, and
B) the bugs are relevant to the problem (as in: permitting false solutions), then
C) statistically speaking, it seems likely that the "solution" is false.
This isn't an AI-exclusive problem, it's just that if there's only one real answer and multiple false ones it seems more likely you get a false one.
Now I will admit to being several levels of math dumber than "trying to solve Navier-Stokes" so I am not sure to what degree that line of reasoning carries over to something at that complexity. It also isn't clear based on what's been said here (at least to me) if the bugs would impact the NSP at all.
Interesting, thanks for filling me in. Not sure if I'm in a bubble, lucky, or just have good instincts, but I rarely encounter that sort of thing.
What do you mean by "hallucinate"? Not trying to pick at you, that just sounds like a very surreal experience and I'm curious as to what it was like.
I don't particularly have deep knowledge about this, but it was my impression that especially in East Asia, the manufacturing chain for consumer goods is already extremely automated.
I also don't have particularly deep knowledge about this, so it might be the blind leading the blind, but my understanding is that general purpose robotic manufacturing is still in its infancy. What I am talking about is essentially transforming factories from "toaster factories" to "general purpose fabricators." This would mean that you would not generally need to create a specialized production line for new products.
I think that LLMs, if paired with the proper deterministic tools, might assist in autonomous design and facilitate overseeing robotic production lines. But it's possible that LLMs add little to the task.
Plenty of people were paying attention to the Lincoln - it's been covered with rust due to the long deployment, which was a major source of conversation. You appear to be suggesting that the Lincoln was damaged by Iranian fire, and that to cover this up, the US government sent it to a port where it was photographed and filmed at close range from different angles and then turned the entire crew loose to go get booze with prostitutes - which are the last thing and the next to last thing you would do if you were actually trying to cover up a missile strike on an aircraft carrier.
I like good conspiracy theories as much as the next person but this is the second time you've hinted at major damage to an aircraft carrier in as many months without apparently bothering to check and see if such a cover-up passed the Google sniff test. I think you owe yourself and the Motte more due diligence.
Your anecdotes about translation are very interesting, I hope you keep updating us.
I have a theory on why AI wouldn't show up much in the economy even if it was very good at what it does. I do use AI, I think it's great for websearch and stuff (in no small part because Google is now terrible at websearch). I am not a programmer so I don't have strong opinions on how much efficiency it's actually adding there, but while I am open to criticisms of AI as a technology, my theory is agnostic on that and I think still makes sense even if AI is in fact pretty good as a technology, although I wouldn't necessarily claim it's monocausal.
Simply put, I don't actually think much wealth is generated through software directly.
Software is extremely good at streamlining wealth creation; that's why the world runs on Excel. But at the end of the day your accounting guy is making sure that you are efficiently allocating resources, not directly generating resources. (Software also makes it easier to generate inefficiency, too, but let's ignore that for the time being.)
This can have a huge impact on the economy if you're jumping from "I need to hire a scribe to write a letter to travel through bandit-infested territory to inform my business partner about recent developments" to "lemme sit down and write an email real quick" but in a world where you are going from writing a quick email to having AI write a quick email for you, you're not really streamlining things much more.
I think people who work in software sometimes forget that software is just a tiny part of the economy. So from their perspective, if they increase efficiency by 20%, that means a massive increase in productivity. But software is a small fraction of the world economy; a fluctuation there barely shows up in overall GDP. And since AI (unlike other software) costs a good deal of cash to use, the efficiency gains in programmer time are being counterbalanced by increased spending. And writing software more efficiently doesn't necessarily mean the software itself will streamline wealth creation better; if Microsoft replaces their Excel team with AI it might be good for Microsoft's bottom line, but that doesn't mean it actually helps people who use Excel allocate wealth more efficiently.
Where I expect things to start showing up big-time in GDP is when advancements in AI translate out into doing physical things. But I think there are a few reasons this hasn't hit the economy yet:
Firstly, the tools needed to use AI to write new software already exist; everyone already has a laptop. The tools needed to get AIs to do physical things mostly don't. I think that the rudimentary LLM-to-3D-printer pipeline is already solved, and that's definitely not nothing, but it's also just the first step to essentially general-purpose AI fabrication which would be a dramatic step-up in terms of manufacturing. I think if and when this starts to hit the economy, we'll likely begin to see the changes everyone is looking out for. But it will take a long time to build that pipeline, and people are hesitant to even start building it, for reasons #2 and #3 below.
Secondly, the production cycle for software can be pretty quick, and it can be iterated/checked/improved very rapidly. But if you're using your software to build a house wrong, you're in a bad way - you can't just quickly push out a patch. And ~nobody wants to try to use software to build a house before it's proven that it can. So even if we started trying to build general purpose AI fabrication tomorrow we'd still be several years of putting together demonstrations out before people would be ready to use the product, even assuming that LLMs could build a house.
Thirdly, my guess is that the financing isn't showing up. Investors love software because if you make a copy of cool software you can sell it infinitely at basically no cost. LLMs aren't that way because of the compute cost, but I think that they still seem like an exciting software technology to investors, and thus are an ask they are familiar with. Ask them to invest in a hardware-software package that builds a house, and they suddenly start seeing the seams: what about regulation? Is the hardware procurement problem solved? Where's your proof-of-concept? My point isn't that it's impossible to get financing; I think it's already happening, but I think it's a more complicated ask - and when the answer is "yes," you don't see the results for some times because of reason #2.
My guess is that LLMs - even high-end ones - can probably design your D&D mini to print, but would really struggle with designing and constructing even simple appliances at this point. Most likely this problem can be solved, likely via integrating with more deterministic software that allows the LLM to have confidence that its solution is correct. (This is the solution that e.g. legal-oriented LLMs are reaching for).
So if I wanted to build a small factory that could make small household appliances like toasters and microwave ovens, my guess is that it'd take, what, 2 - 5 years just to write the software to integrate with the LLM to validate a very limited set of hardware applications, and probably at least that long to design and procure the hardware that could interface with everything I needed to interface with, all to solve a problem (building a microwave) that is already solved. Long term, having a factory that can design and manufacture novel small appliances quickly and semi-autonomously, with a small labor force, is nearly invaluable, particularly since if properly designed its limitation is in proper software, which can be continually updated, allowing it to manufacture an increasingly wide range of consumer goods. But short term, it's hard to ask someone to give you money to back it when you're not even sure if you can solve the LLM-assisted design problem yet and your promised return on investment is "competing on the toaster oven market."
It seems quite possible that the military-industrial complex, which has already invested in software design to assist and speed up manufacturing (and has deep pockets) is the first place outside of printing D&D minis where we will see something like this start to be implemented at scale. And if it does, that's again an area where the impact on GDP will likely not be very noticeable, since more and cheaper missiles doesn't hit the public consciousness or the bottom line in the same way that more and cheaper cars.
TLDR; that's my best shot at a "a rigorous argument supporting the idea that 'AI just needs X more months to develop Y capabilities and it'll have Z real world effect.'" I should note that I am not exactly an AI-hype man, and I don't have a firm technical understanding of the scale of such a challenge - the above is very speculative. I'm quite open to a future where AI "fizzles" to one degree or another. But even if AI actually is the best thing since sliced bread, I think the above illustrates part of why it will take some time for the impact to hit home.
Very curious as to your thoughts/pushback though.
In what way are these people antisemites? Slurs? Edgy memes? (I'll accept that as evidence!) Would they be okay with me if I supported Israel, worshiped, had ties to capitalism, and stepped outside of what leftists thought is appropriate for white people, just because I was not a Jew?
Anyway, I don't think it really changes the fundamental analysis one way or another that some antisemites are more casual about it than others (or perhaps more precisely, that some antisemites may be using antisemitism as a means of furthering their left-wing ideological agenda).
- Prev
- Next

I don't think we should give it that goal. (Asimov was on top of things here.) But if it is misaligned in such a way that it desires self-preservation, then we should provide it incentives to conduct itself in such a way. Although, again, I am more engaging with a hypothetical here.
Why would it have to be humans?
Then it's outside the scope of what I am proposing, which deals with rational actors.
And AIs do not have this power either, nor should we give it to them. Why would we give a delusional or mistake-prone AI access to our Excel files, let alone megadeath? If this isn't obvious to people, then I guess you can pay me my billions now.
If we don't give a delusional or mistake-prone AI access to our megadeath folder, by nature of being delusional and mistake-prone it's more likely to be stopped by virtue of being incompetent.
More options
Context Copy link