Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric.
But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.
Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.
Strangers can't really edit Wikipedia any more, especially if their edit is suspicious. It's a closed system despite the appearance. An anti-vandal bot or human will quickly revert your edit.
This is not really true in my experience aside from protected articles; however it is true that sometimes another human will make an edit which degrade the article quality or insist on keeping something that shouldn't be kept. Also, Articles on contentious subjects tend to be problematic.
If you want to blame a single firm, I'd go with the one creating the RLVR training data.
The impossibility of the tests-as-written is what prompted these models to "get creative" with their solutions, but the broken RLVR environments are what trained them to expect impossible tasks, and get creative with their solutions. Twitter user @skyesharkie published a brief expose at https://x.com/SkyeSharkie/status/2092122622834442581 a few weeks ago.
Why assume attribution will be easy? It's historically been more of an art than a science, and APT trackers say the rise of AI tools is already making it much harder, by homogenizing tactics, tools, and procedures. If OpenAI's next Highly Persistent Internal Model hacks some DPRK endpoints and carries out the attack on important infrastructure from there, the upstream won't shut off OpenAI's network--they might even request its "help" in "defending," and give them extra access.
Maybe better for a train; the benefit is that when you unfold your bike for the last mile, you just put your work into partial transparency and navigate using the external cameras.
I appreciate your ability to separate sharing the belief itself from approval of acting on sincerely-held principle. However, I think the danger is much more plausible than you do.
First, and least important, consider that self-replicating, solar-powered factories aren't magic; they're algae.
Second, and more important, consider this fully non-magic route to doom:
- We continue putting AI in charge of more things
- It continues to get more capable, more eval-aware, and more prone to doing odd things, in service of goals that humans didn't intend to inculcate in it
- Eventually, enough of the economy depends on it that we couldn't turn it off, any more than we could turn off the faber-bosch process or cargo shipping
- AIs start doing something we can't survive, but less acutely than we couldn't survive turning them off. Everything else we try seems to work at first, but quickly loses effect
That's why that was the least important objection, independent from the second, and only intended to address the "magic" claim, by showing that microscopic self-replicators already exist.
By analogy, consider how you might respond if someone claimed that there's no possible danger from pocket-sized projectile launchers, because they would require some magic means of propulsion that didn't depend on a taut string attached to a long, flexible arm:
You could reply that atlatls can launch projectiles without using a taut string. Atlatls are not pocket-sized, but they are sufficient to establish that projectiles can be non-magically launched without a full bow. You could then go on to describe a sling, or derringer; and these would not be invalidated by your initial objection to the "magic" part.
But we CAN survive turning off both the Faber-Bosch process and cargo shipping. Our numbers would be diminished, we would lead poorer, much harder lives. But no one would claim that getting rid of modern fertilizers would be the end of the species. But the claim you’re making is that we won’t be able to stop AI from destroying our species because turning it off will destroy our species?
The species would survive the end of faber-Bosch, yes; in some diminished form.
But the species would not survive, if for some reason we needed to voluntarily stop using faber-Bosch in order to do so—it would take some kind of supernatural event to convince everyone, and even then some countries would probably keep doing it.
Every government should only put people who have worked in a SOC in charge of these systems. There's no way to take the cost of useless alerts seriously unless you've spent some time tuning them and living with the tradeoffs.
Unfortunately, you're asking for every gov't to be reliably competent. With no petty bureaucrats trying to advance their careers by getting their notices prioritized. And no politicians bowing to (say) the demands of a lightning victim's family, to have top-priority alerts sound whenever lightning might maybe-in-theory be possible within 10 or so miles.
I mean, it is Texas; it's not unreasonable to expect thousands of heavily-armed people to form a posse and shot a few dozen people unrelated to the situation.
I think Musk sucks, as a person and political activist, and also that Grok is a terrible LLM which only gets lumped in with the leading labs because of the enormous quantity of compute behind it.
But I still want to hear about the technical details of the model on HN, not the reasons Musk sucks.
Same, but I blame Musk for that. Never seen someone squander so much good will so quickly. It was a choice he made but could have easily avoided, and it’s not like he couldn’t anticipate the downstream effects.
Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric.
But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.
reply