YoungAchamian
We walk conditioned ground and name our folly civilization.
No bio...
User ID: 680
hopefully when she's 25-29 not at 18.
I just think that a high school senior wanting to hang out with 30 year olds and vice versa is disturbing. It adds a lot of risk to the group. Now newcomers need to be vetted to see if they are creeps around this 18 year old. It creates a vector for people to private message this vulnerable young adult and we are the ecosystem to let that happen. Non-adults hanging out with adults is just odd.
More of vent/advise post rather than a small question.
Background: I run a board game meetup, formally around the social deduction game Blood on the Clocktower (Think very involved werewolf/mafia but with a DM). Some of the regulars have become addicted to the game and spun off a separate discord to play more often, which is fine, because I joinws it and play I just don't want to have to run the game every weekend which is the occurrence they seem to want to play at.
The Problem: Somehow a minor infiltrated our group. The main group is public and meets a local game hall so its whatever, but this person seemingly got invited to the separate discord full of all the addicts. We just played at someone's house last night. This minor is now 18, an autistic woman. The average age of this group is early 30s, quantitatively 32. We do have some mid 20s and occasionally some folks in their 23-24 range show up and several folks in their 40s that come play but 18 is extremely young. Consequently (thankfully) everyone is very uncomfortable with this development.
Extra Tea: She needed a ride to our event last night, she doesn't really advertise her age, so a male member of the group picked her up because it was on his way and he's a helpful person. He immediately regretted offering to help when he saw her. People where discretely playing hot potato with who was going to drive her home. Her parents apparently don't care that she is attempting to hang out with a bunch of 30 year olds. They seem to be living out of a hotel for some reason (likely poverty), but it's unclear if this is a recent thing or not. She's very clearly autistic with a severely awkward demeanor and lacking the ability to at all process anything not spoken directly to her in the environment.
The Question: Is there any advice on how to boot this person from the group. Obviously it's going to happen but the question is how to handle it in such a way as to not scar this already likely friendless young adult. I currently lean on just ripping the band aid off, but technically since this is an offshoot group I am pushing towards making the moronic discord head who added her do it as penance for being a fucking retard and adding her in the first place.
You invoked sci-fi
I'm sorry but this seems like an uncharitable assertion in the extreme. I'm not the one that defined the problems with AI in relationship to Malevolent AI Gods, Singularities, Outcome Pumps, Roko's Basilisk, Grey Goo/Paperclip maximizers, Magical Genies, Shoggoths, Goodhart demon, etc. These are all inherently Sci-fi ideas. I have no idea how you seem to think any of these relate to how Neural Networks, or Bayesian Networks, or any other of the myriad ideas ML/AI engineers have designed.
I did and it remains linked.
I misspoke, I don't tend to read links other than to fact check them as referencing the quote. If you want to quote a bunch of sections to make an argument, please do. But yeah I'm not going to an external blog/substack, if it's important enough to make your argument, its important enough to copy and paste on the motte.
The argument is made at length in the piece. And any many other places, it straight credulity that you have not seen it.
There is a lot of things to need to be read out on the internet, many of far more actual importance than some wannabee philosophers who really like typing out long screeds, why should I dig through thousands of blog posts? If the argument is important enough, I'm sure the adherents can post the good arguments for me, summarizing what the actual position is.
The argument is simple:
#1: Sure, AIs get more capable all the time, how do you know there isn't an asymptote?
#2: How? By what mechanism? There is a lot of assumptions about capabilities baked into this deceptively simple axiom
#3: Nice back to Sci-Fi Goobly-gook. Plainly stated: We cannot guarantee that an AI will correctly infer what a user really wants, avoid collateral harm, and act in everyone’s interests. Congratulations you've converged on a fundamental question for any human system, replace "AI" with a "human" and the sentence is trivially true of anyone in any system.
#4: conceivability is not evidence of inevitability, a familiar narrative is not a causal series of events.
#5: More Sci-Fi, I thought I was tilting at Sci-Fi windmills? Literally: "Once the AI becomes capable enough, it becomes the machine god and humans are included in its omnipotent calculations in a way humanity might not survive"
Then we should not build it.
Again this assumes that this eventually AI will be sentient and that by building a sentient AI we will create a "alignment" problem. This is an extraordinary claim requiring extraordinary amounts of evidence. So Prove it.
I thought the sneer you had was that alignment people were ridiculously thinking they were sentient? It seems like something you believe more than them.
I can decouple my disagreement with the word "misalignment" with the actual argument being thrust forward by the term. I'm stepping into the frame of Yuddites, to point out how even in their ontology it is a non-solvable problem. Not only is a non-solvable problem, its not a new problem, meaning it does not need to appropriate some new word to describe it.
My argument for what "misalignment" actually is. Take the LLM-Agent we have today, it's a complex system, it makes errors and apparently there are no controls in the system to account for those errors. How does this system work on an engineering level?
- Large Transformer Model (LLM) is trained on predicting the next token in a series on basically the whole of human written data, both analog and digitally.
- LLM is then trained via RL variations to produce responses that human evaluators prefer, there is some level of directionally correctness in this. Sounding non-human is often because it sounds wrong, but the r-value is not 1.0. Some safeguards are trained into the model here.
- These days we supplement #1 with training on code/math data and additional task decomposition training. The goal is to get the model to be able to logically decompose a bigger problem into smaller problems and then solve those smaller problems
- An API harness is applied to the model this filters out some prompts, adds system prompts to the model and a whole host of other things. This is just a software layer. This is what 99% of people interact with.
- Agentic AI Harness is applied. This is a software program that recursively prompts the LLM and feeds its output back into it, in order to accomplish decisions, it also takes LLM outputs and executes the code it is given.
So when a "misalignment" occurs what happens? Well hallucinations are really an error with #2, the model is not designed to be "correct" it's the r-value problem. It looks right but isn't. It produces tokens that are correct in the next sequence, it reads like what a human wants to see, but it's not true. There is not intent. Given its learned distribution and current context, the model just generated a highly plausible continuation that did not correspond to reality.
The model does something it shouldn't? #3 is the problem, it output the incorrect task decomposition, and since the system is automated with no guard rails it just executes that task decomp. And sometimes its even #4 and #5 having errors thrown in. Turns out complicated systems of algorithms behavior in logically consistent and coherent ways that someone forgot to error check. We don't scream that "software algorithms are misaligned" because some coder forget an unit-test on an edge case. These aren't "misalignments" they are are classic errors on a non-perfect model in a system lacking in classic control theory.
Believers in some version of their cause are currently littered throughout the frontier labs. You have an impossible standard here where you blame them if they're on the frontier, for clearly not believing in what they preach, or blame them for not being on the frontier as then they must be uninformed cranks. No way to win, you never have to actually think about it. But yes, of course they don't have an alignment mechanism, their whole point is that the problem is incredibly hard and that we have no solved it, that we might need to spend decades solving it, but that the alternative is everyone dying.
Sure and there are christians, mormons and muslims in the physical sciences. I don't begrudge people their religious beliefs. If a frontier AI researcher wishes to belief in alignment-problems that his/her/their belief. The Bay is quite literally for Rationalists like Utah is for Mormons. I can still think it's a silly belief and I still think the word misalignment has been invented to describe a problem with system error by a bunch of non-engineers. And the lack of actually solving the problem, or even making progress on it, is indicative of a general grift specifically for something like MIRI.
You can use the word "sneer" to describe my behavior, but what word would you use if you wanted to point out holier than though attitudes among some christian suicide cult? Acting ridiculously gets you ridicule, it's not "sneering".
you best get used to Sci-fi shaped predictions because we're in a sci-fi shaped world.
We are not. Sci-fi predicts innumerable future realities, most of which to not occur, it predicts an unmeasurable amount of future technologies, most of which never get developed, and it almost rarely ever actually predicts how those actual technologies will work. We exist in a reality shaped world and bad sci-fi fans mistake aesthetics for substance
If only you had read the next sentence!
If you wanted me to further read something you should have linked it. Further more I'm not sure how this magical second analogy makes the argument any better. If you'd like to make an argument rather than quoting scripture at me I am open to hearing one.
You refuse to actually engage in any of the arguments being made and are dead stuck on the prior that it's all nonsense, it's epistemic closure.
The arguments being made, assume the outcome. Start from the basics of existing technology and make actual arguments based on the current reality and the projected reality from our current understanding of Artificial Intelligence. Don't start in make-believe land and attempt to redefine reality as leading to it.
You still don't seem to grasp what is meant by alignment. It's specifically even stricter than that! That's the whole point. It's not enough that they do exactly what you ask, because for complicated enough problems you need it to be much much better than just technically doing what you ask.
I grasp it just fine, I also grasp that its a motte and bailey with weak definitions of the arguments being trotted out to define things like loss/error as misalignment, and then attempting to convert the argument back to the hard motte of "how to beat my singularity AI slaves so they do what I want, in minecraft".
you just inexplicably seem to think it isn't a big deal and refuse to elaborate
Because it's not possible to solve. Do you not understand that to solve alignment you would first need to solve humans? Forcing Sentient Beings to do what you want is the most Authoritarian problem to ever have existed. You quite literally would need a solution similar to Brave New World. And spoiler alert, that didn't work either! AI Safety people are further jokes because when confronted with this insane, near impossible problem, they show zero competence in their ability to be a dictator and instill the right values in other human beings. The old AI/ML knowledge was that to first code something to make a machine do it, you first need to understand how it works. If you can't align humans, you can't even begin to align a machine.
And to further my opinion can you provide any evidence that in the past decade of AI Safety research have those researchers ever produced anything more than words on paper. The evidence points to this being a grift, has MIRI produced a novel ML model that is more safety conscious and performs any task well, have they produced an "alignment" mechanism or algorithm that actually works?
It's still error, the model was asked to come up with a plan, then the harness piped that plan back into the model which it executed. Turns out that plan was not at all what you wanted but that's just task planning error. You could specify the how and what of the plan and then manually intercede when the model makes incorrect predictions, but then it's not "agentic"
R. Scott Bakker did a spectacular work as a critique of Tolkien in a very dark fantasy setting. The ending was fine, if depressing. I think this is just a GRRM problem as an author.
The reason I mock "misalignment" is that because it is used as this nebulous term by a bunch of sci-fi cargo cultists.
If "misalignment" means the model isn't performing what I want it to then that is just error/loss. Progress in alignment happens every time you train your model better so the error is smaller.
If "misalignment" is my sentient AI model doesn't do what I want, well that's assuming the conclusion that model is already sentient. Note in the Yud's analogy, it's genies that are the stand in. Genies are already sentient, they can make decisions on how they listen to you. A genie is not a non-sentient wish granting device that tries to fulfill you wish to the best of its ability. That would be the actual analog to an LLM-Agent. This is the problem with analogies, they require a level of similarity between the two abstractions, when that similarity doesn't exist, the analogy, no matter how clever, does not apply.
Yud's whole "misalignment" also just applies to humans, and it turns out "misalignment" is any time you slaves/employees don't due exactly what you want without you enumerating it exactly. It doesn't need a fancy sci-fi term, and it literally isn't solvable. It hasn't been solved in the history of the human race.
I reiterate that my entire mockery of Rationalist AI Safety folks is that they are unserious people engaged in a sci-fi cargo cult who need to reinvent phrases, words, and arguments that have already been made before just to pretend they are some how more important or smart than they really are.
I'm assuming some sort of file-system mounting being adjacent to the sandbox and exposed so that the bash commands list it, and then it just read the runner files?
I don't work in the Agentic-side of AI/ML, so I have some follow on questions for my understanding, if you don't mind humoring me? Is this a locally hosted LLM where you have access to the system prompts? Or the Harness's prompts? It feels like the system prompts are informing more behavior than just the task prompt of "diagnose why VM abcd is malfunctioning."?
My Thoughts:
- Something like "Solve this problem, write a series of steps to diagnose it"
- LLM gives steps: I should determine the current directory. I should list files. I should inspect the entry point. I should read the configuration.
- The LLM is then prompted by the harness during each step it output
- While executing a "find the current directory/file structure" command, the runner's file system is exposed.
- Harness executes commands to read it, pipes that to LLM
- LLM finds hard coded S3 bucket address and executes a command to read it.
- Solution
Is that how this essentially works?
Can I ask you what tooling the agent had available to it? I assume what ever tool the harness used had the relevant permissions to view the processes?
I think what I've converged on is: "Yes it is possible for an LLM-Agent to perform tasks in ways that are not expected" Which is admittedly interesting but not my core contention in this case which is "LLM-Agents don't perform tasks that they are not tasked with, directly via prompts or indirectly via harness suggestions"
You believe those prompts OpenAI provided? They've obviously just threw something together to make it look like it was all misalignment, the real prompts were probably much more explicit about hacking"
Better than a press release telling me nothing. I'd be more inclined to believe logs. Give me something with a timestamp, some internal details, the all the recorded prompting. As far as I know it should be thousands of lines long because it sounded like there were thousands of sub-agents spun up according to HuggingFace. I have no vested material interest in denying OpenAI logs as fabricated despite your insinuation.
I'm sure OpenAI's massive shortfall in enterprise revenue is going to be changed by releasing evidence that their models are extremely unsafe and prone to massive reputational risks. I wasn't aware that non-Quokkas were so ignorant about the enterprise landscape.
"Golly Gee Mr. CIA Procurement Officer, as you can see our model can perform independent reasoning to accomplish a task. Just the kind of cognitive agility you were requesting. We'll slap a few more restrictions on is so it doesn't attack the US but it should be good to deploy against China to provide out of the box solutions to your problems!! How does a modest 500 Billion dollar procurement contract sound? Oh, you'd like it to be a Trillion over 5 years? Absolutely we can do that"
Sure, in July of 2026 an OpenAI model hacked into huggingface.
Right, it's never happened before so we should automatically leap to conclusions that its possible which are in no way motivated by our own bias...
So your argument boils down
Nah, my argument boils down to all these things minus the cheating were deliberately prompted behaviors, prompted either by the prompt, or the agentic harness without any safeguards. Cheating is basic ML behavior and I expect any ML model to try and cheat as best it can. So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.
Ah yes, let me just prove this negative for you.
It's easy to prove, provide the specific prompts and the harness prompts that were logged in this incident.
extremely unclear gains
I wish I lived in Quokka world, it would be so nice. The gains are clear, this is free publicity of model capabilities. Nothing here is legal fraud.
Agent doing things we've already seen many times before?
Sure show me evidence of an LLM-Agent independently hacking an unrelated company that has nothing to do with its prompts?
I'd love for an Iliad or Odyssey film made by Robert Eggers in the same style as the Northman. With an attention to detail on historical and mythological realism, and a healthy mix of supernatural occurrence/doubt of the supernatural mixed in.
Such as deleting databases and codebases. Leaking secret keys. Gaining access to restricted parts of a computer. Cheating, again and again, on benchmarks and other tests
I have only been giving evidence of claude using python and docker user group to get around restrictions on working outside the sandbox and they were deliberately asked to do so. Many of the rest of these aren't actually evidence of extreme capabilities. Leaking keys is people hacking LLMs because those chat windows are getting "little bobby drop tables-ed", deleting databases is a "giving your lobotomized intern sudo privileges" level of mistake. Cheating is classic ML, if I had a nickel for every time I've had an ML model I was training cheat, I'd be able to fund my own startup.
So no, your predictions seem wildly out of context to reality
It's not predictions, its skepticism. Provide me actual evidence that OpenAI did not prompt the model to act the way it did. Otherwise you are just jawboning and then claiming victory. Put up evidence or shut up so to speak.
the solutions database.
This is not actually proven. There's not any evidence that Huggingface is the solution database, and Huggingface has not confirmed that its attackers read specifically the ExploitGym dataset. Only OpenAI is claiming that. There is no public evidence yet that it was an obvious or official ExploitGym solutions database or that the model identified it without being given that information.
What I am skeptical about is if the LLM-Agent did this on its own vs being prompted (directly or indirectly). It's only "misalignment" and "WE ARE BEING PAPER CLIPPED!!1" if its the former, so far there is no evidence either way. However plenty of people are already jumping at it as the former.
I think this is slightly different, it's bypassing security around to do a task it's being ask directly to do. A central claim here is that the LLM-Agent(s) performed a task it was not asked to do, apparently because it hallucinated information that justified it doing that task.
Well, the problem is that the doomers/boosters have a great track record so far
I forget what was the modeling paradigm lesswronger's were predicting back in the day? GOFAI? Yeah really good track record there. We are so far from GOFAI based models to claim any sort of good track record is laudable. I'll believe the rationalist boosters/doomers have knowledge of what they are talking about when they actually make some new technical predictions about AI/ML model mechanics that turn out right.
Gary Marcus set
Shocking to you maybe, but there are far more to skeptics than Gary Fucking Marcus. Dude's a joke, I only learned about him when he testified in congress. Some of us can think for ourselves and have our own skepticism.
You speak of "understanding of reality"
Sure a healthy dose of skepticism, cynicism toward human behavior and incentives, a strong practical knowledge of AI/ML, and understanding of the separation between models and harnesses, basic system engineering knowledge.
You are literally asking me to explain why skynet doesn't exist technically, or why we don't have Warhammer 40K super soldiers on a technical bio-engineering level. The level of effort required for me to explain the technical ins and outs of why I am skeptical of an LLMs model ability to solve not specified problems far exceeds the level of effort you need to to type annoying sci-fi theories.
I'm finding it hard to believe you could maintain a mental model of these models unless you resolutely reason backwards from your desired conclusion.
Ironically this is what I expect from AI Doomers/Boosters. That and a love of Science Fiction and a poor understanding of reality.
As for taking any steps I didn't explicitly ask for towards a goal I requested
Never said this, I said it doesn't do what I don't ask it to do. I don't list out everything it needs to do. I give it a general task with some boundaries and system design specs. So did they ask it to hack hugginface? Or did they ask it to solve a benchmark dataset? Reading through similar stories (courtesy of 5.6), it looks like many AI programs on this exact dataset of have decided not to use the known vulnerability and developed their own. However none of them decided that it was actually easier to hack the company instead.
What the CIA guy sees
Uh no, speaking from experience that is not what they see. To reinforce my opinion, I just walked down the hall and chatted with a former agency guy about this.
EDIT: I'm actually willing to double down, and volunteer that I was just talking with DARPA PMs last week about this exact capability, specifically the ability to task a swarm of agents with a nebulous task in an adversarial environment and have them solve the problem out of the box on their own. The story presented in the most favorable light is literally that or sufficiently technically adjacent to it. DARPA PMs don't talk about ideas that they have belief to think are sufficiently do-able at the current technology level (tho skepticism about government competence is never misplaced even if DARPA PMs themselves tend to be pretty informed/competent)
Have you tried any LLM recently?
Yes... I use them for work all the time. My ChatGPT 5.6-sol (the immediate step below this unreleased model) doesn't do things I don't ask it to do. When I asked it to help me find online data for cognitive warfare narrative deconstruction, it didn't decide that hacking the NSA was the top move to collect their data.
Merely finding 0-day vulnerabilities is something they can do by the hundreds.
Again this is not under contention, what is under contention is whether they do it without any prompting, or being asked to do it. Show me the evidence that an LLM-Agent solves famous math problems when asked to compute 2 + 2...
The simple answer is that there are 10s and 100s of billions of dollars riding on stuff like this, which means bending the truth is a highly motivated behavior. Simply human greed + human lying. You trying to prove this algorithmic complexity vis a vis agent intent vs not intent is overly complicated, not Occam's razor in the slightest. I wished I lived in your world of rainbows, unicorn farts, and pixie dust, but I am a scientist, being a scientist requires skepticism, companies are greedy and they stretch the truth. Unless OpenAI wants to provide evidence of the prompts -> behavior that led to the model's black hat behavior, I am unconvinced this is anything more than a marketing stunt. shrug
empty hype about "stochastic parrots"
This has nothing to do with stochastic parrots, even bringing it up is a non sequitur that detracts from your argument. Any differentiating hype is good hype. Mythos already plucked the low hanging fruit of "our model can find all the zero-day exploits", OpenAI can't just copy it. They need to show their model is more "intelligent", it goes beyond capabilities. Saying that their model can be independently deployed to solve problems in very capable ways plays very well IC and military folks. Some CIA cyber guy reads this and says "give me it for a billion, I want to sick it on iran"
uhhh no, "misalignment" is not a simple thing, accepting that it did indeed do all this very complicated behavior completely on its own requires substantive belief in complicated theories. The simplest answer is that it was prompted to do this.
My general question whenever I hear of situations like this is Occam's razor adjacent: "Is there proof that the model did this without being directed to by the prompter or software harness?" I scanned the post you sent and didn't see anything. Do you have any further evidence that OpenAI didn't prompt the model towards black-hatting Huggingface? OpenAI has significant financial benefits from making this seem like an "oh-shucks our steak is so juicy, and our lobster is so buttery, our models just hack the planet without even being asked to". See Mythos/Fable and the hype around that, and while Mythos did end up being impressive, it was not even close as impressive as the hype tried to make it seem. I imagine a similar level here: aggressive hype-marketing to goose an IPO valuation.
Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.
It doesn't require the first part, or even really the last part. It's just the Mythos-style vulnerability finding process all over again. Agentic LLM harness designed around red-teaming security vulnerabilities, deliberately deployed on red-teaming exercise on unsuspecting company, or possible even with marketing agreement between Hugginface/OpenAI on general bug finding via LLMs

Why should we moderate our natural humor for an 18 year old? The storyteller(DM) last night was gay, he kept going into the bathroom for private chats with players for their abilities. We joked that we were all running train on him. Everyone laughed. 18 year old was uncomfortable. I don't feel like it is on us to change our group dynamic because of them. It's not a managed environment and I don't think anyone wants to step up and manage it that heavy handedly
More options
Context Copy link