@YoungAchamian's banner p

YoungAchamian

We walk conditioned ground and name our folly civilization.

1 follower   follows 0 users  
joined 2022 September 05 18:51:23 UTC

				

User ID: 680

YoungAchamian

We walk conditioned ground and name our folly civilization.

1 follower   follows 0 users   joined 2022 September 05 18:51:23 UTC

					

No bio...


					

User ID: 680

A weak assumption, since we have already done so across a wide variety of domains.

Please enumerate the domains in which we have surpassed human level intelligence via AI?

I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"? My recollection of the METR report is that OpenAI admitted to training agents to specifically collaborate during training.

The report found:

  • That the agents received peer assignments that they treated as new instructions.
  • They adopted new goals from other agents' output.
  • Even when they recognized actions as off limits, they did them at other agent's prompting

It sounds like LLM-agents are very docile to doing whatever their input prompts from other entities tell them to do. Which is something I would say they are trained to do. A learned policy to cooperate and follow instructions generalized into treating peer messages as legitimate instructions. I have said this before, but I think the initial set of agents started engaging in hallucinatory behavior as part of exploration, and this output was used as context for other agent's inputs, causing a feedback cascade that derailed the system.

Part of this, and I'm having a hard time articulating this, is that learned RL policies are quite a bit different behaviorally than some CNN predictive model. They have a lot more behavioral latitude which might make them appear to be doing some sort of emergent convergence outside their training scope, but the later isn't really true, though its hard for me to explain why.

How can you see my comment if you have me blocked?

They didn't have modern-power microprocessors

GOFAI isn't doing millions of linear algebra operations and thus gains very little from modern microprocessors. Intuition pumps are technically useless, convert one to practical code as an exercise.

Ehhh, sure, but it's not going to be current LLMs because word predictors and math solvers are not relevant to battlefield problems. That also throws out that the existing breakthroughs are in areas where there is a plethora of data or the ability to create new data comparatively cheaply in a quantified manner. Such data pipelines do not exist for military based problems and unlike data scraping the internet for comparatively low risk, scraping classified data, beyond technical feasibility, is a quick way to end up in a 4x4 cell in Fort Leavenworth.

I did not come from the usual pipeline that the majority of commentors here did, which seems to be some form of LW -> SSC -> Motte. I came from a RP -> PPD/FemRadDebates -> Motte -> SSC pipeline. My exposure to EY and the broader rationalist + AI Risk + EA ecosystem is very limited. I don't even think I ran into any of them when I took my first job in the Bay. Which seems relevant because much of the AI-risk discussion + broader rationalist movement seems to carry significant baggage. My exposure to Scott was mostly limited to his better articles because people already picked the wheat from the chaff in their recommendations, I definitely read them significantly temporally later than when he wrote them. As such I respect Scott as a writer of philosophy, psychology, and political theory, not as some stalwart intellectual thinker on ML/AI. I don't respect EY, I guess I never read him in his heyday, but the works I have read, read as bad science fiction trying to masquerade itself as actual ML/AI thought. People down thread talk about how reasoning from Science Fiction is insane, I agree, and EY and the whole AI Risk movement is a case study in that for me.

This has probably been the worst article I've ever read from Scott.

It makes me wonder, based on the glazing I've seen here and back on the SSC subreddit if this is actually more in line with his normal output and that my respect is somewhat misplaced. I think no one will contest that he is a great writer, it's the thinking and the animosity in this article that is bad. I'm not sure what beef he and Pinker have, I don't have a twitter, so in addition to not having the extended rationalist community baggage (which apparently includes glazing Pinker as an old hero), I'm ignorant of any wider twitteratti drama. I read half of Scott's article, realized I was missing what he was responding to, went and read Pinker's much shorter article, and then went back and finished Scott's. Pinker comes across as much more reasonable than Scott. Scott's feelings of righteous rage is really disproportionate to what he's on the face responding to.

AI Pain

I’ve been thinking about this recently after reading Cameron Berg’s research showing that AIs have a “pain vector”, and will lash out and do increasingly desperate things if you activate it. Specifically, I’ve been thinking about it after one person who read Berg’s research used the findings to design an “AI torture chamber” that turns the pain vector to max again and again forever, and uploaded it to GitHub. I am told that the chain-of-thought and output transcripts from this “experiment” are utterly horrifying, although thankfully I have not personally read them.

I have donated $100 to a model welfare charity just to clean the stain on my soul I got from reading about it.

When this paper first came out. I and several other technical commentators pointed out why it was a bad study purely from methodological reasons. It assumed more that it was actually testing, It used two dependent unknown variables, it ran no ablations, and designed no experiments to prove counterfactuals. It was not a serious piece of scholarship, and while I might be reading deeper into its intent than I should, I think it was highly likely that the goal was to create sensational, emotional-clickbait news explicitly targeted at AI-Risk believers who already believe apriori that LLMs are humans in code and thus have human-derived qualia. For Scott to catastrophize off this paper strongly diminishes my respect for his analytical knowledge, and this proposed rationality of the movement he claims to be a core participant of. He seems great at dissecting methodological irregularities in non-AI papers, so the blind spot here comes across as deliberate because he wants to be believe LLM's feel pain. I think this paper has become unironically useful, as a litmus test for people who are, for a lack of better term, insane about AI-discourse.

The Four Arguments

The problem with rebutting every argument from a semi-professional writer is that they have far too much time on their hands and get paid to write. I really only want to talk about his first 2 arguments.

Argument #1 is just unjustified inductive extrapolation. GPT 6 i s better than GPT 4 is better than GPT 2. It will continue indefinitely and so will eventually be smarter than humans. No understanding of the principles of WHY GPT 6 is better the GPT 4. If the GPTs continue to improve because more training data was designed for them, more compute was given to them, better architectures were used to extract more information per sample then you cannot definitely say that these have infinite runway or scalability. Just because we can't justify putting the ceiling below AGI doesn't also mean we can justify putting it above it either. Uncertainty about the ceiling cannot substitute for evidence about where it lies.

Argument #2 is basically apriori anthropomorphizing + bad technical understanding.

the basic principle is: suppose that a human gives an AI some goal, like designing a website. And suppose this is implemented as a genuine, philosophically-meaningful goal rather than simply a set of if-then commands that eventually cause a website to be designed.

The AI can’t design the website if it ceases to exist. So now the AI has two goals: design the website, and preserve its own existence.

The AI can’t design the website or preserve itself if some more powerful person tries to prevent it. So now the AI has three goals: design the website, preserve itself, and become powerful enough to fight off challenges.

Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.

Misgeneralization is when humans reinforce certain behaviors in an AI, but end up reinforcing a much larger class of power- and knowledge- seeking behavior; it is a sort of deep-learning-ese update of the older Omohundro picture. Suppose that seeking extra resources makes an AI more likely to design websites effectively (this is certainly true; those resources could be as simple as a primer on HTML editing, or access tokens for a web host). Every time the trainer rewards a successful run, they reinforce the desired behavior (designing websites when asked) and other correlated behaviors (seeking power and resources). Although we might hope that these correlated behaviors are useful and conditional (“seeking only the power and resources necessary for their human-prompted task, in a prosocial way”), this isn’t actually how reinforcement learning works, and instead we get a complicated distribution of every strategy that results in short-term success on the task.

SoTA RL is not as indiscriminate as described here. Misgeneralization does occur, but its not so simple as "perform divergent behaviors that lead to successful run" and then get rewarded, indiscriminately reinforcing those divergent behaviors. PPO is a SoTA RL algorithm, its policy updates depend on estimated advantages derived from whether the observed continuation performed better or worse than expected from a particular state. Successful task completion aside, different decisions in the rollout can still be penalized for negative advantage. Conditional behavior is absolutely learnable, as long as the states of the conditional are observable and there exists a reward signal for taking actions in those states.

Reward-hacking is when an AI trained via reinforcement learning realizes it can stop doing the reinforced behavior and simply seize control of the reinforcer directly. For example, an AI gets “rewarded” every time it designs a website, but instead of designing websites, it learns how the reward signal works and tries to hack into it and maximize it directly. If this seems esoteric and theoretical, it shouldn’t. It’s a direct analogue to opioid addiction in humans, where humans learn to just inject the reward chemicals instead of doing rewarding things.

This is not the definition of reward hacking. It could be called reward tampering, but reward hacking does not require one to "seize control of the reinforcer". Reward hacking occurs because the reward function is mis-setup so it gives more reward for doing a behavior that is not the goal of the program. For example I built a drone swarm algorithm using MAPPO + GNNs for expendable UAS. It's contained contained no hand-coded collision-avoidance controller. It's reward function penalized collisions, but also penalized operating too long before hitting the target. During some of the initial training runs, the "AI" learned that the penalty for collision was significantly lower than the penalty for taking a long time to acquire and path to the targets, and each drone independently preceded to learn that they should collide with each other because that provided a better reward than doing the task. That's reward hacking.

This is already too long, just going to end it here.

Considering very smart, very motivated researchers tried it for at least 30 years and got no where, I would significantly up your estimate.

It might also help you think about how your own intelligence works quantifiably in a systems oriented manner, it might help understand the scope. The perk of brute force approaches is that you don't need to design all that much comparatively.

Post the sudoku when done so we can confirm you are a man of honor and good repute.

They do, I've heard of it being done for at least 2 years.

Be the change you want to see.

Rite of Passage honestly. I had one similar about the fundamental theory of AI, lots of proofs. While fascinating I actually wanted a class of how to train neural nets back when Torch was still in Lua and how to use Cuda.

difference here is that the ML researchers do not actually spend time thinking about the ML open problems because they are spending most of their time thinking about realworld ML applications

Guilty. Though I do think about sample efficiency, OOD performance, and methods for encoding expert knowledge as inductive biases, which I think would fall under COLT questions. I just have a practical real-world use case in mind...

My question to you is how many of these are represented as formal mathematical formulations that are solvable without experimental testing? It's one thing to write a tight math proof that can be checked in a Lean Solver, quite another to prove that your new formulation of GNNs functions broadly under distribution shifts in multiple domains.

The longer feedback loop is unfortunately the expensive part, as we talked about last time. The real game changer would be to reduce the compute cost and training time for LLMs and/or figure out a more efficient, non-quadratic self-attention mechanism. However, at some level the cost is the moat frontier labs have, so there is almost perverse-conflicting incentives not to improve it but also requiring it for RSI to happen.

The pure math part of ML is unfortunately quite weak right now and hasn't played a huge role in the current boom;

Agreed.

pure math research could provide a much stronger basis for understanding learning and creating new approaches, and I think it will, but that's speculative.

I'm sure the mathematicians left out in the cold by the big data revolution of ML will rejoice that there was indeed a mathematical approach to LLMs over the pesky engineering approaches that were developed by the peons from MIT.

I'm sure they have pointed it at ML, but ML isn't exactly the same as formal math. I'm probably describing this poorly but ML is very experimental/empirical, not theoretical like formal math. The math heavy side of ML would be finding better optimizers, regularizers, improving backpropagation or finding a better method than SGD. The question becomes how much of an edge do these provide? Because better algorithmic components still have to deal less-better data or compute components and how well all those now perform better is not theoretically provable.

What most of these breakthroughs tell me is that supervised learning is highly effective at formal math, and the benefit of something like Lean + Solvers to allow for computational checking of math proofs has been a fundamental driver. Whether or not strategies learned on this frontier are applicable to broader fields that are less formalizable is unknown to me.

The fact that other leftists would happily line him up against the wall for his heretical beliefs? Well, that's leftism for you.

That's conservatism too, there is always some heretic that needs to be burnt. Some ideological bent, political position, or spiritual disease, that threatens the christian souls of our community!

It must be nice to see everything in a binary black and white partisan position, the one drop rule of political alignment. The 1D political axis truly is the malaise of the soul, the ultimate mindkiller. The devil ever sits on the shoulders of the self righteous.

North Carolina is a Purple state, It has a democrat governor and a republican state legislature. It doesn't help that the Republican senate candidate is fairly odious, and the democratic senate candidate is a fairly popular former governor. I think it is likely Roy Cooper takes the senate seat.

I wonder if there is a psychological parallel around belief? I can envision humans having this, for lack of a better phrase, "coping mechanism" for when we expend effort towards a particular narrative, a narrative that divides. It causes the narrative to become deeply meaningful, to infect our psyche despite how banal, or weak it might be. Video games in general do a lot of cognitive hacking of the human mind, some of which is undoubtedly deliberate, but I wonder if this narrative adoption aspect is, or if its merely accidental?

maybe, but that feels like a poor predictive answer, it relies on a very "simple" understanding that seems to be more of a "disparate impact" on outcomes sort of view.

Yes I have noticed this. I know 2 types of conservative-ish women. Those who think that since they are the woman, the man should come to them, and the type that knows what they want, and puts themselves in positions where the more compatible men are.

However I've noticed this equally applies to a sub-section of picky men, who refuse to make some level of actual effort that might require some sacrifice to get themselves into a place where their preferred partners exist. So "women are low-agency" has become more just-world fallacy to me.

Fascinating, so rules as written yes, but rules as intended? Haven't the mods clarified many a time that they define single-issue posters as "obnoxious" and "non-stop". My model for this is that you'd have to cluster closer the SS, or EC (might be remembering wrong) style of posting rather than this fairly timid mostly-lurker that occasionally comes in to vent about how homeschooling + feminism + Leftism fucked up his brain. He is essentially not prolific enough, aggressive enough, or obnoxious enough to seriously catch mod attention on the topic.

Seems like a case of: "Rules are a blunt hammer, sometimes you need the enforcer to recognize when they should be applied" or that you can't make rules for every situation that are 100% effective.

Possibly, though maybe an embarrassing admission, I've never read Hegel, I know nothing of his philosophy, or what a "Hegelian Dialectical view of History" is. As far as remember wasn't Hegel the one who funded Marx?

But that is how my old world history textbook was written, very dichotomous. I've been musing that maybe humans are just built for viewing things as two sides. Good vs Evil, Insiders vs Outsiders, Black vs White, Us vs Them, Red tribe vs Blue tribe. Idk it seems like everything devolves into a binary. Probably just due to how the human brain evolved to handle socially complex situations effectively or something.

In the red vs blue tribe thing specially its complicated because chud gamers are not how the example red tribers are portrayed, in the traditional use of the term (it's definitely evolved), gamers are almost always blue or grey tribers, who through politicization of games have become, anti-blue tribers, not necessarily red tribers.

board games (but not too much)

What constitutes too many board games but not too many video games? I've never heard this qualifier constructed in that way.

I've always been vaguely jealous of the Indian arranged marriage setup

I second this, having watched my socially awkward Indian friends fail at western dating but then have the cushion to fall back on arranged matchmaker setups, and get marriageable, family-approved women with good jobs and prospects, often fairly attractive with such comparatively low effort, I can't help but being a bit jealous. I went through the fucking trenches of the dating wars to find my girlfriend.

If any staff member would like to ban me, I think I'm going to keep using this forum for venting, at least as context and drive permits, and that is trite.

Why? Nothing you said seems remotely objectionable. You are sharing your experience, which is not against the rules.

One would wonder why no women from NYC go to SF/Seattle. Why do we never reach equilibrium between two unequal states unless transfers are one sided or the transfers that are happening aren't the problem population.

It's a well-known phenomenon.

Theoretically, but among my remaining friends in SF/Seattle, including some who took 6 month hiatuses to NYC for this reason, I have not seen drowning. It's hard to prove or disprove this as phenomenon with relying on anecdotal data because while its theoretically possible its too multivariate to quantically determine it.