@YoungAchamian's banner p

YoungAchamian

We walk conditioned ground and name our folly civilization.

1 follower   follows 0 users  
joined 2022 September 05 18:51:23 UTC

				

User ID: 680

YoungAchamian

We walk conditioned ground and name our folly civilization.

1 follower   follows 0 users   joined 2022 September 05 18:51:23 UTC

					

No bio...


					

User ID: 680

I'm assuming some sort of file-system mounting being adjacent to the sandbox and exposed so that the bash commands list it, and then it just read the runner files?

I don't work in the Agentic-side of AI/ML, so I have some follow on questions for my understanding, if you don't mind humoring me? Is this a locally hosted LLM where you have access to the system prompts? Or the Harness's prompts? It feels like the system prompts are informing more behavior than just the task prompt of "diagnose why VM abcd is malfunctioning."?

My Thoughts:

  • Something like "Solve this problem, write a series of steps to diagnose it"
  • LLM gives steps: I should determine the current directory. I should list files. I should inspect the entry point. I should read the configuration.
  • The LLM is then prompted by the harness during each step it output
  • While executing a "find the current directory/file structure" command, the runner's file system is exposed.
  • Harness executes commands to read it, pipes that to LLM
  • LLM finds hard coded S3 bucket address and executes a command to read it.
  • Solution

Is that how this essentially works?

Can I ask you what tooling the agent had available to it? I assume what ever tool the harness used had the relevant permissions to view the processes?

I think what I've converged on is: "Yes it is possible for an LLM-Agent to perform tasks in ways that are not expected" Which is admittedly interesting but not my core contention in this case which is "LLM-Agents don't perform tasks that they are not tasked with, directly via prompts or indirectly via harness suggestions"

You believe those prompts OpenAI provided? They've obviously just threw something together to make it look like it was all misalignment, the real prompts were probably much more explicit about hacking"

Better than a press release telling me nothing. I'd be more inclined to believe logs. Give me something with a timestamp, some internal details, the all the recorded prompting. As far as I know it should be thousands of lines long because it sounded like there were thousands of sub-agents spun up according to HuggingFace. I have no vested material interest in denying OpenAI logs as fabricated despite your insinuation.

I'm sure OpenAI's massive shortfall in enterprise revenue is going to be changed by releasing evidence that their models are extremely unsafe and prone to massive reputational risks. I wasn't aware that non-Quokkas were so ignorant about the enterprise landscape.

"Golly Gee Mr. CIA Procurement Officer, as you can see our model can perform independent reasoning to accomplish a task. Just the kind of cognitive agility you were requesting. We'll slap a few more restrictions on is so it doesn't attack the US but it should be good to deploy against China to provide out of the box solutions to your problems!! How does a modest 500 Billion dollar procurement contract sound? Oh, you'd like it to be a Trillion over 5 years? Absolutely we can do that"

Sure, in July of 2026 an OpenAI model hacked into huggingface.

Right, it's never happened before so we should automatically leap to conclusions that its possible which are in no way motivated by our own bias...

So your argument boils down

Nah, my argument boils down to all these things minus the cheating were deliberately prompted behaviors, prompted either by the prompt, or the agentic harness without any safeguards. Cheating is basic ML behavior and I expect any ML model to try and cheat as best it can. So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.

Ah yes, let me just prove this negative for you.

It's easy to prove, provide the specific prompts and the harness prompts that were logged in this incident.

extremely unclear gains

I wish I lived in Quokka world, it would be so nice. The gains are clear, this is free publicity of model capabilities. Nothing here is legal fraud.

Agent doing things we've already seen many times before?

Sure show me evidence of an LLM-Agent independently hacking an unrelated company that has nothing to do with its prompts?

I'd love for an Iliad or Odyssey film made by Robert Eggers in the same style as the Northman. With an attention to detail on historical and mythological realism, and a healthy mix of supernatural occurrence/doubt of the supernatural mixed in.

Such as deleting databases and codebases. Leaking secret keys. Gaining access to restricted parts of a computer. Cheating, again and again, on benchmarks and other tests

I have only been giving evidence of claude using python and docker user group to get around restrictions on working outside the sandbox and they were deliberately asked to do so. Many of the rest of these aren't actually evidence of extreme capabilities. Leaking keys is people hacking LLMs because those chat windows are getting "little bobby drop tables-ed", deleting databases is a "giving your lobotomized intern sudo privileges" level of mistake. Cheating is classic ML, if I had a nickel for every time I've had an ML model I was training cheat, I'd be able to fund my own startup.

So no, your predictions seem wildly out of context to reality

It's not predictions, its skepticism. Provide me actual evidence that OpenAI did not prompt the model to act the way it did. Otherwise you are just jawboning and then claiming victory. Put up evidence or shut up so to speak.

the solutions database.

This is not actually proven. There's not any evidence that Huggingface is the solution database, and Huggingface has not confirmed that its attackers read specifically the ExploitGym dataset. Only OpenAI is claiming that. There is no public evidence yet that it was an obvious or official ExploitGym solutions database or that the model identified it without being given that information.

What I am skeptical about is if the LLM-Agent did this on its own vs being prompted (directly or indirectly). It's only "misalignment" and "WE ARE BEING PAPER CLIPPED!!1" if its the former, so far there is no evidence either way. However plenty of people are already jumping at it as the former.

I think this is slightly different, it's bypassing security around to do a task it's being ask directly to do. A central claim here is that the LLM-Agent(s) performed a task it was not asked to do, apparently because it hallucinated information that justified it doing that task.

Well, the problem is that the doomers/boosters have a great track record so far

I forget what was the modeling paradigm lesswronger's were predicting back in the day? GOFAI? Yeah really good track record there. We are so far from GOFAI based models to claim any sort of good track record is laudable. I'll believe the rationalist boosters/doomers have knowledge of what they are talking about when they actually make some new technical predictions about AI/ML model mechanics that turn out right.

Gary Marcus set

Shocking to you maybe, but there are far more to skeptics than Gary Fucking Marcus. Dude's a joke, I only learned about him when he testified in congress. Some of us can think for ourselves and have our own skepticism.

You speak of "understanding of reality"

Sure a healthy dose of skepticism, cynicism toward human behavior and incentives, a strong practical knowledge of AI/ML, and understanding of the separation between models and harnesses, basic system engineering knowledge.

You are literally asking me to explain why skynet doesn't exist technically, or why we don't have Warhammer 40K super soldiers on a technical bio-engineering level. The level of effort required for me to explain the technical ins and outs of why I am skeptical of an LLMs model ability to solve not specified problems far exceeds the level of effort you need to to type annoying sci-fi theories.

I'm finding it hard to believe you could maintain a mental model of these models unless you resolutely reason backwards from your desired conclusion.

Ironically this is what I expect from AI Doomers/Boosters. That and a love of Science Fiction and a poor understanding of reality.

As for taking any steps I didn't explicitly ask for towards a goal I requested

Never said this, I said it doesn't do what I don't ask it to do. I don't list out everything it needs to do. I give it a general task with some boundaries and system design specs. So did they ask it to hack hugginface? Or did they ask it to solve a benchmark dataset? Reading through similar stories (courtesy of 5.6), it looks like many AI programs on this exact dataset of have decided not to use the known vulnerability and developed their own. However none of them decided that it was actually easier to hack the company instead.

What the CIA guy sees

Uh no, speaking from experience that is not what they see. To reinforce my opinion, I just walked down the hall and chatted with a former agency guy about this.

EDIT: I'm actually willing to double down, and volunteer that I was just talking with DARPA PMs last week about this exact capability, specifically the ability to task a swarm of agents with a nebulous task in an adversarial environment and have them solve the problem out of the box on their own. The story presented in the most favorable light is literally that or sufficiently technically adjacent to it. DARPA PMs don't talk about ideas that they have belief to think are sufficiently do-able at the current technology level (tho skepticism about government competence is never misplaced even if DARPA PMs themselves tend to be pretty informed/competent)

Have you tried any LLM recently?

Yes... I use them for work all the time. My ChatGPT 5.6-sol (the immediate step below this unreleased model) doesn't do things I don't ask it to do. When I asked it to help me find online data for cognitive warfare narrative deconstruction, it didn't decide that hacking the NSA was the top move to collect their data.

Merely finding 0-day vulnerabilities is something they can do by the hundreds.

Again this is not under contention, what is under contention is whether they do it without any prompting, or being asked to do it. Show me the evidence that an LLM-Agent solves famous math problems when asked to compute 2 + 2...

The simple answer is that there are 10s and 100s of billions of dollars riding on stuff like this, which means bending the truth is a highly motivated behavior. Simply human greed + human lying. You trying to prove this algorithmic complexity vis a vis agent intent vs not intent is overly complicated, not Occam's razor in the slightest. I wished I lived in your world of rainbows, unicorn farts, and pixie dust, but I am a scientist, being a scientist requires skepticism, companies are greedy and they stretch the truth. Unless OpenAI wants to provide evidence of the prompts -> behavior that led to the model's black hat behavior, I am unconvinced this is anything more than a marketing stunt. shrug

empty hype about "stochastic parrots"

This has nothing to do with stochastic parrots, even bringing it up is a non sequitur that detracts from your argument. Any differentiating hype is good hype. Mythos already plucked the low hanging fruit of "our model can find all the zero-day exploits", OpenAI can't just copy it. They need to show their model is more "intelligent", it goes beyond capabilities. Saying that their model can be independently deployed to solve problems in very capable ways plays very well IC and military folks. Some CIA cyber guy reads this and says "give me it for a billion, I want to sick it on iran"

uhhh no, "misalignment" is not a simple thing, accepting that it did indeed do all this very complicated behavior completely on its own requires substantive belief in complicated theories. The simplest answer is that it was prompted to do this.

My general question whenever I hear of situations like this is Occam's razor adjacent: "Is there proof that the model did this without being directed to by the prompter or software harness?" I scanned the post you sent and didn't see anything. Do you have any further evidence that OpenAI didn't prompt the model towards black-hatting Huggingface? OpenAI has significant financial benefits from making this seem like an "oh-shucks our steak is so juicy, and our lobster is so buttery, our models just hack the planet without even being asked to". See Mythos/Fable and the hype around that, and while Mythos did end up being impressive, it was not even close as impressive as the hype tried to make it seem. I imagine a similar level here: aggressive hype-marketing to goose an IPO valuation.

Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.

It doesn't require the first part, or even really the last part. It's just the Mythos-style vulnerability finding process all over again. Agentic LLM harness designed around red-teaming security vulnerabilities, deliberately deployed on red-teaming exercise on unsuspecting company, or possible even with marketing agreement between Hugginface/OpenAI on general bug finding via LLMs

I lived on the outmost west side of the "good" parts of the city, in avondale/logan square. It was fine. The neighbor had a bit of "character" about it, but was more rough working class rather than actually ghetto. The worst thing that ever happened was a drive by shooting 3-4 blocks south of my apartment. Unfortunately, as is generally the case, a bystander was caught in the crossfire.

I've heard it about parts of Chicago as well

Chicago was strongly affected by red-lining. Anecdotally, most of the non-petty crime is strictly limited to the south and west sides of the city. There is very little chance you will just "wander" into the bad neighborhoods unless you are completely oblivious, deliberately or otherwise, to the changing neighborhood affluence markers. It's literally a gradient from good -> bad. The only real exception to this rule is the area around Hyde Park, which is the purview of the University of Chicago. It's smack dab deep in the ghetto, the metra stop for Hyde Park is very rough, but the actual neighborhood itself is super nice, until you cross the street boundary.

Once again, you won't find me saying its either gender's fault.

It's the women.

In general I think this is one of those maybe technically correct opinion in the strictest sense but is basically incorrect. You never say the exact words "I blame women" instead you repeatedly assign a disproportionate amount of causal and moral responsibility to women for the problems you've identified. It's a pain in the ass to dig through all your comments ever, so I sicked ChatGPT 5.6-sol on your user history, it only sampled some of them, I guess AI is as lazy as humans are, but it looks like it took the most recent 3 pages of your profile history and then about 5 randomly sampled and the 4 ending pages. Something like 28% of your comments are Dating/Women related and 84% of those expressed negative sentiment about women according to the Bay-Area RLHF regime. I suppose I could get a better read by running the code myself but that feels more invasive than what I just did and there seemed to be some DNS problem that I am too lazy to debug.

Look, I don't think men are the problem. But the problem is with men. Specifically, in their mind, in that they've formed cultural expectations that, via feedback loops, are completely divorced from reality and renders their own desires unachievable.

So... you have to address their desires. Reality can't be manipulated to fit their desires, so it seems obvious to me that you gotta at least TRY to make their desires comply with reality.

Making this about men, makes it still just as right as it was about women. It's almost like this is a complex multi-faceted systemic problem where both sides are stuck in delusional feedback loops feeding their own rage-filled echo chambers.

This seems to undermine your thesis.

I'm not sure how, this is no different than "Being a misogynist is repellant to women vs have you look at female attraction towards dark triad traits". Studies are complex and a twitter post with no paper attached is the worst kind of evidence because you can't even fact check. It's literally just selling outrage for clicks. I get that you like living in an echo chamber because outrage feels good but this is rage porn. News flash interacting with other people in competitive stressful environments makes you like other people less.

We need to fix things for their sake, too.

Yes I'm sure like most conservatives you believe we need to save people from themselves. If only they would not be gay then their souls will be saved, if we just torture them some more they will find salvation in heaven. What is weeks of earthly pain for an eternity in heaven!!! Do I need to pull out the C.S. Lewis quote?

What's annoying is pointing out that all that data, anecdotes, and personal experience is all pointing in the exact same direction.

How many women have you talked to about their anecdotes, and personal experiences? You clearly act like you think women are children, is that why you don't think their opinion is relevant to the conversation?

"sounds like you hate women bro."

That said I've also never said this, and I don't think it. I think you have a very complex relationship with women. They clearly are some sort of collective villain in your mind.

Tho the longer I have this conversation the more I feel like some poster arguing with SS about DaJoos. As in you've already made up your mind, there is nothing that would change it. You already have this grand unified theory of everything. Everything returns to DA FOIDs, I talk about social community formation, back to DA FOIDS HAVE NO HONOR (This is a flair for the ages).

Part of my personal history is that I was two weeks away from actually getting married before my ex broke it off.

Right but was the "Women are the problem" part of your view back then, or was that the catalyst that pushed you towards "the discovery of the truth"? Dating really is a lot of luck, you need to meet the right person at the right time.

I will take no personal blame for problems that are clearly systemic in nature. If EVERYONE is struggling in precisely the same way, its not an individual problem. Its definitely not a me problem.

Yeah I can agree that the current dating market is fucked, I just don't buy grand overarching narratives that its all a single gender's fault rather than the complex interplay between genders, their desires and their interactions. However given that this is the environment, the real question is more what can you do to protect yourself, improve your odds, and still build the life you want? Being right about the system can explain your situation. It cannot, by itself, improve it. At some point, you have to decide whether proving that the problem is unfair matters more than finding a way through it.

There seems to be a bit of a motte and baily here: Motte: There are systemic issues causing relationship formation and interpersonal social cohesion to plummet. Bailey: Women are the cause of all of this because xyz evidence.

So no, I'm not very persuaded by someone trying to assess my desirability based solely on posts I make on the autistic argumentation and controversial debate forum, who has never interacted with me in real life.

As is your right, I once was like you, I made the same arguments you do, with the same totalizing world view. I think it can be very cathartic, but it's very resentment driven. Resentment poisons the soul. You come across very resentful on this forum towards women, but who am I to know you IRL. I can only pattern match to myself of yesteryear and advise you to take a more nuanced stance that isn't so totalizing, blaming a single gender for all the social dysfunction. I'm sure you'll tell me its not resentment, its observing reality and truth! Truth hurts!! Yeah... I've been there...

This is something I really am learning to hate in other autistic people. They can accept these two premises individually but can't seem to understand what they mean together because they can't observe the phenomenon.

  1. Humans communicate on two levels: explicitly through stated content, and implicitly through subtext. Additionally subtextual communication is both conscious and nonconscious: tone, inflection, body language, etc. Essentially non-verbal behavioral leakage exists and is positively correlated with strong emotive feelings

  2. Autistic people suck and recognizing, and conveying subtextual communication. They can't passively observe it, not within how they communicate or how others communicate.

Result: Strong emotionally affective beliefs end up being roughly communicated through non-verbal communication in nonconscious ways. This is obviously very lossy information wise, but when ever has social comms been an exact science, it's all vibes and intuition anyway. If you are autistic you will have a hard time determining what non-verbal comms you are sending out, and may literally be communicating more than you intend.


You seem to have a seriously complicated relationship with women. You blame them for the poor dating environment, you blame them for poor social environments, you think they are less pro-social than men, they can't honor their words vs men. You think the political ecosystem is worse because of women, you think women need to be monitored, controlled, shepherded because they are impulsive children, they lack desirable traits in a partner, they cheat more, their education is socially harmful, it goes on and on. BUT, you also love women, you've staked your entire existential meaning towards something that requires a woman to love you. You are dependent on women for your desired future, for your meaning. Women are simultaneously the villains of your entire world view and the indispensable condition of your personal salvation. (If you would like me to reference all of your posts on these, I absolutely can).

You have women up on a pedestal and simultaneously in a gibbet.

These are strong emotionally affective beliefs, I absolutely think they leak.

You actually remind me of a close friend, he's very similar to you, late millennial, single, has trouble with women, expresses very similar opinions to you on that subject, slightly autistic, slightly conservative (though a militant vegan!) and for the important reference he's crypto-conservative, ie hiding it. He doesn't talk conservative points, doesn't express his power-level, in his opinion nobody should know he's a conservative. But in our friend group it's pretty much an open secret that he's conservative. Why? Because all of his non-verbal communication, and the very few comments he does make. Neurotypical people can intuit it, they talk, they compare notes. You state a strong opinion around him, and you can literally see the right-wing debate bro combination of Charlie Kirk + Sam Harris as gears in his head start spinning up, even though he refrains from engaging. I don't have character references, so yeah maybe you are unimpeachable, but if it's just your opinion on your unimpeachability, well I think you might be unaware of the subtext you are conveying.


That said I think you should try non-party social events in a public place around a shared activity. People flake because they feel the social pressure to say yes and the lack of social consequences for not showing up. I have experienced it when I organized events around social parties in Chicago, it's not gendered. I think an open social activity is more likely to draw new blood. Food for thought.

EDIT: and because I realized I didn't respond to several of your points, I don't think this is solely on you, absolutely there is a trend of social isolation among people these days. But this returns to my pragmatic point: You can be happy or you can be right. Just because there is a trend doesn't mean you are helpless and should just sit back and seethe, you can make changes to your own lifestyle, behavior, beliefs, and if you really place all this existential importance on a family, then you would try to make it happen at any cost. No change would be too much of an ask. Who cares what everyone else is experiencing, you need to look out for your own happiness, your own success, and your own life fulfillment, it's a hard game out there, you just going to quit and whinge because the deck is stacked against you??

So I read through your original post, I have to say I'm don't really think thats the same thing I had in mind. You are doing house gatherings from an existing BJJ(?) gym community. That gym community doesn't select from a "social"-oriented group, and is already a niche section. There's no way for you to bring in extra people other than manually recruiting from the gym. At the party, is it just a nebulous gathering like most parties? Or are you activity oriented? People tend to flake on nebulous gatherings because its just generic social time.

The difference with what I and my friend did, is that it's oriented around a very specific activity, it is a large gathering (I have 15-30 people showing up), they are public so you get a flux of random new people in, which softens the blow of core group members not showing up.

ABYSMAL gender ratio

Yeah the dance classes can be really hit or miss, 80% of people show up with their partners. BUT the goal is not to date within the class. No, the class goes to clubs/bars doing latin night where there is far more women than men, women who might not be formal dancers but grew up dancing etc, or just want to have a good time. And you are the skilled dancers, because you took that class for months years even, and you have fall back friends to dance with if it's not working out. This burning desire to date and only date at all costs comes across as desperation. You gotta let go of it and just enjoy the moment.

My board game group meets via Meetup, it's a bigger social deduction game, called Blood on the Clocktower. I can have up to 15 people playing. We meet every couple weeks at a consistent interval. People come to play the game, maybe get dinner afterwards, that's it. People find us through Meetup, or friend referrals and while we have flakes since they sign up of their own accord it rather than being socially pressured there's less of it. There's a set time and activity so it's not just another generic socializing party. My current girlfriend came for at least a year before I asked her out, I imagine some people date but I don't really pay attention to it.

zilch NO social pressure on women to actually honor commitments

Dude this isn't just women, these are people, men and women are flaky as adults. The fact that you are grinding the edge on women only comes through. People pick up on social intuitions.

The problem is with them.

There is something deeply ironic about this, it's like a union organizer complaining that nobody wants to show up to the meeting to fight the capitalists.

"The problem is with them!!" he screams, looking at his empty seats, "All those traitorous capitalist scumbags! How dare they not show up to my meeting!"

Proper disclaimer first: anyone who tells you they have a high success approach in this environment, who doesn’t live in your polity, have your hobbies, is asking you to join a cult, or sell you a subscription to hustlers university…

Solution: form your own community. Legitimately waiting for others to create something so you can come in and consume is why third spaces no longer exist. Too many free loaders. It doesn’t need to give people purpose(like a cult) it needs to be a space that people want join, giving them friends, connections, people to hangout with. Instant social status, you are a leader, you have cred. It’s a ton of work but i think its rewarding.

It’s up to you what community you want to make, but i built mine around social deduction board games. We have a 60/40 gender split, which is functional because it’s not full of sleazeballs only there to date. I pull people off it to play the hard core board games I like. Idk if people in it are dating each other but my current girlfriend of a year did approach me through it.

A close friend in Chicago did something similar with Latin dancing. He doesn’t date within the group but they go out to latin dance clubs and he dates women he meets through there. The women in the group wingwoman him, the other guys show him respect because he organizes it. He recently fumbled a year long relationship for stupid personal reasons but he met her through his dance outings.

Just food for thought, not sure what your hobbies are or where but i see more success from making your own community than waiting to find one.

You’re not talking to a boomer mate, I spent 12 years in the OLD trenches. I know the stats, the sources, the arguments. I’ve made them, I sang that song, and I’m telling you very few of the couples I know, now including me, met online and I hang out with NERDS, where there is no/little stigma.

I literally came here from PPD.

Which bridges and doors are you referring to?

The ones that exist when you are a part of a community. When friends introduce you their partner’s friends, when you meet potential partners in your community, when people in your community vouch for you, when you bring your partner to your community and the women side conversation convey what a catch you are. Basic social things. Nothing reinforces a woman’s mate choice like a community that approves and lauds that choice as a good match.

If nobody is matching up then the obvious answer for you is to plant your feet in the sand and refuse to soften your language… because that will show them!