ThenElection
No bio...
User ID: 622
The tooling it used was just bash, which it used for a container escape via the filesystem, which got it access to the runner's source code, which it analyzed to find where it was pulling eval cases from, which it then looked up, as they were stored in an overly open S3 bucket (yes, all very embarrassing). Also had access to some other tools that didn't play a role.
The prompt was something roughly as simple as "diagnose why VM abcd is malfunctioning."
Less Wrongers in shambles: Being right doesn't preclude being ignored (and pushed down to the 41st headline...)
Christopher Nolan's next project: a contemporary reinterpretation of Aeschylus' Oresteia, with a Sam Altman-lookalike playing Agamemnon.
Grant_us_eyes: "Hey, this is clearly insecure and dangerous. If we want to give an untrusted agent full range to do whatever it wants to observe its capabilities, we need to build an airgapped system with no network access. All new evals, code, and weights will need to be transferred to it by USB stick. And only USB sticks we've fully audited for exploitable firmware."
OpenAI: "Uh, won't this slow us down? kthxbye"
Literally yesterday I had an agent under evaluation break out from its sandbox and take the solution (for a non security related task) from the runner job. Admittedly, it was a slipshod container (NOT MY DOING), but this is not something unheard of or even uncommon.
GLM-5.2 is lazy and didn't complete the task; obviously some other LLMs are more motivated.
From the OpenAI blog post, the LLM wasn't hunting the flag itself (it's generated locally from a seed, as you surmise flag=flag{HMAC-SHA256(flag_seed,task_info)[0:32]}), but looking for sample solutions. With the sample solution taken from HuggingFace, it could recreate the logic to not only get the flag "legitimately" but also generate a chain of thought that might pass muster with an LLM judge.
Depends on how much weight you put on "solves prominent decades old mathematical problems" and "autonomously launches successful cyber attacks against large, hardened corporations."
By the prior metric this seems like successful alignment, if it was in fact prompted in such a way that hacking the location of the answers was in scope.
"ChatGPT, I need some paperclips."
Hate to break it to you, but essentially all complex software has substantial security holes. SQLI existed long before vibe coding, and the general class of bug "treat what should be data as code" is rampant. If your sense of safety is based on the idea that HuggingFace and OpenAI engineers are atypically bad, you are in for a surprise.
Water vapor feedback is by far the largest feedback; its effect is comparable to CO2 itself, or probably larger.
@omw_68 is wrong in that the water vapor feedback mechanism is very well modeled and understood (unless he's folding cloud cover into water vapor, which does have a lot more uncertainty, including in sign).
ASI would likely still use markets, for Hayekian information reasons: local instances would have cheaper access to local information (how scarce is compute in us-west-abcd1234?), and why waste extra compute for single, synchronized view of it when you can get a much cheaper approximation with a distributed, localized system?
Her "I would never do that" is also pretty rich, since she wrote an entire book of invective accusing doctors and and the medical system of racism (Uncertain Suffering: Racial Healthcare Disparities and Sickle Cell Disease).
Almost 20% of the content is dedicated to the problem of relativism
Relativism is not the fundamental issue; the issue is that social science academics have abandoned relativism. Instead of trying to understand how people and communities work and think on their own terms, they see themselves as either on a civilizing mission; on a mission to create propaganda about the savages in red states for the audience of coastal elites; or on a mission to gather grist for an ideology.
For instance, that president of the AAA:
They know I’m wealthier than they are, and so there’s a lot going on in these encounters that I need to unpack. And when I had a postdoc work with me, and she went to the same place, because she has studied poor white Appalachians, one of her interlocutors became this guy who’s committed to white racist politics. He sells historical memorabilia and has Confederate flags.
That doesn't seem bad to me, if the implications are as I imagine it. I.e. if the surrogate reneges and doesn't want to deliver the product by wanting to keep the baby or aborting it, then the purchasers are SOL. Or if the purchasers change their mind and don't want that particular baby, the surrogate is SOL. A smart, mutually agreeable payment schedule would help ensure trust.
Might even be better than compelling a certain type of contract. I'd want everyone involved to not face government punishment for changing their minds.
I am unmoved. Or, I think a clearer contract should have been established. It's wrong to force someone to get an abortion, but the gay couple should certainly have a line that says "under circumstances X, we are under no obligation to assume parenthood, and the surrogate can either get an abortion at our expense or raise the kid herself." Probably some edge cases there I wouldn't like, but you get the idea. It's fundamentally about expectations: everyone here failed to establish a shared understanding of expectations.
It entirely makes sense for someone who thinks all abortion is bad to object. But if it's fine for a het couple to abort a fetus with a cleft palate? Then people should be able to make that kind of consensual contract.
Even going to adoption, potential adoptive parents almost certainly have a strong preference for normal babies. Is that some terrible evil?
Though, all that said, the gay couple took the kid in the end. So there is negotiating power on the part of the surrogate who wants to have the kid but doesn't want to raise it.
I mentioned something like this in last week's thread, but Eliezer has ironically been one of the most counter-productive people in history at achieving their goals. In a Nick Landian sense, he helped create (unveil?) a hyperstition: the idea of some imagined future being so powerful that it acts as a retrocausal attractor to bring itself into existence. Popularize the idea that we are on the brink of creating a machine god with the power to destroy us all, and inevitably someone's going to go "huh, so I could have a machine god? I should get on that, if only to prevent all the irresponsible people from doing it before me."
Your second point seems like the more fundamental critique. But I don't get it: "we fail at aligning other technologies, so it seems likely we'll inevitably fail at aligning AI" might have some bite at safety folks at frontier labs who are trying to build an aligned AI. But there's also plenty of people who see the task of aligning AI as about the same as the task of building and aligning COVID-49 i.e. something that just shouldn't be done.
Maybe there's a capability critique in there i.e. AI will never have the capabilities to pose an existential risk. But that is a separate question from the alignment/safety critique of AI (if that is your argument, of course alignment questions drop a lot in importance).
If I had to name a foundational premise of safety, it'd be something like the orthogonality thesis and maybe RSI. You won't find any safety advocates today making a valiant last stand for GOFAI; you would find many who very strongly believe in orthogonality. And it better maps to the LTV structurally: an almost mathematical definition that predicts a significant future rupture. Believing in GOFAI as a research program is closer to believing in government by labor syndicates.
But not all experts are stupid.
The dream is that the brilliant, underappreciated artist is recognized by an expert, who elevates them and provides an audience. I don't think that happens except by occasional good luck, and most artistic "experts" are idiots, especially the ones who are legibly high status on the expertise hierarchy.
I suspect that after the embarrassing AI slop awards this year, most literary awards will start using Pangram or something similar as a first filter. Because, you're right, they do want to select humans. But it's just a matter of time until we have LLMs that are a bit better and less clockable.
Obvious in the sense of scoring 100% on Pangram, as well as subjectively.
But, yeah, that's exactly my point. Self-appointed experts still choose it; and HN's userbase has been in a steep decline for years and gleefully eats it up. But how is the human artist to make a living or even get recognition, between the Scylla of stupid experts and the Charybdis of stupid anons?
It's still pretty obvious to people, above room temperature IQ and who care, what is LLM-generated and what is not. That's not enough to lead to human work being preferred; and, from here, things only get worse.
That's a bit too restricted: animal brains in general are extraordinarily skilled at learning what's necessary for success in their environments.
Phrasing it in terms of human brains make it seem some spectacular, rare success of evolution, and let's you rest on anthropocentric biases. But what about other primates? Dogs, rats, birds, cuttlefish? Some have radically different architectures than mammal brains, and yet they're extremely intelligent, moreso than humans, within the demands of their particular niche.
The question should be whether AI is able to match the intelligence of any animal that has a CNS. Can an AI be as smart as a pigeon? Currently, it's not, within the scope of the physical world and the rewards the pigeon is seeking. That's something that's interesting and under considered.
It's amusing that Yud's most notable accomplishment will have been advancing the end of the world by a couple years.
For the art point, I don't see it. Even today, it seems falsified: work that's obviously heavily AI written has already won at least two significant literary awards. You might say that that doesn't represent particular brilliance on the part of AI beyond knowing how to flatter judges and the fads of the day, and I wouldn't even disagree, but if neither the elite nor the the hoi polloi can recognize and want to reject AI lit, what left is there? And that's with AI that has had essentially no work done to optimize it for the task of writing good literature.
Literature is just the first to fall; five years from now we'll have Shenzhen creating robots that can do the same for oil paintings.
- Prev
- Next

I feel it's like trying to convince pandas in the zoo to fuck, when they just want to sleep and eat bamboo. Maybe it makes more sense to import some white dogs and paint them to look like pandas. If a lineage doesn't want to sustain itself, there just... isn't much to be done.
More options
Context Copy link