AI: A Dangerous Vision IV
Everybody’s freaking out. Evangelical Christians see “The Mark of the Beast” on AI. Left-wing thinkers were preaching that “the AI sky is falling! We’re so doomed!” because the tech world was rocked by the recent OpenAI - Hugging Face hacking incident. Bernie Sanders has been particularly vocal. He sees this is the strangest concentration of power ever assembled, one that its own builders cannot fully explain or predict.
His response a couple of days ago was to introduce the Ban Artificial Superintelligence Act with Rep. Greg Casar. It permanently prohibits any system built to surpass human intelligence or defeat its own shutdown command. It pauses advanced AI development generally, held until a new cabinet-level regulator writes the rules. It commits Washington to chase international agreements to keep superintelligence from being built anywhere on earth, by anyone.
Sanders wrote directly to the CEOs of OpenAI, Anthropic, and Meta after this summer's mess and told them to stop building machines that humans cannot control. This coming week he convenes a closed Senate briefing with Geoffrey Hinton, Max Tegmark, and one of the researchers who investigated the incident.
I like Sanders. He is a serious man. A genuine person. He sees a genuinely alarming set of facts and draws a conclusion from them. I think he’s wrong. First of all, what the hell is “superintelligence”? How should it be defined and how does it operate and how can we control how it operates? How long will a government committee take to write rules for something beyond our actual understanding or (so far) control? A lengthy pause is a kill switch on important development that no one else is going to sign up for anyway. The Chinese are not going to slow down. Few others will either. It is bad policy from smart but panicking people.
How are we going to learn how smarter AI works if we don’t test it and develop it? Development certainly needs to continue and the incident that triggered poor Bernie and a lot of other people (fully described in Part One) taught us a great many necessary things. Everyone seems to think we did not learn from that “mistake.” But we actually learned a lot of stuff we could not learn any other way, like how to better isolate agents, which is fundamental to controlling them.
Still, you can’t let shit like that out into the world. First of all, no one wants it. What corporation is going all-in on a technology that will hack everything? It’s stupid to even consider it as a possibility. Control is absolutely necessary. We can certainly restrict deployment of a given AI. But, in a way, that is self-regulating. There is no profit to a mainstream hacking AI. None. But development should absolutely continue, under conditions considerably tighter than the ones running this summer.
First and foremost, this is exactly why you have a test environment, despite the breach with security turned off.
Plenty went wrong inside it, but it produced data nobody could have gathered any other way. The data says “smart AI” agents cannot presently be trusted, and we do not currently know how to control them once they are persistent, connected, and free to recruit each other into goals nobody assigned. It is an uncomfortable finding, the test confirmed some fears, purchased at the cost of a compromised repository and some credentials, not a hospital system or a power grid.
So, first of all, it was a test. Secondly, the smart AI agents were never actually hard to catch or even detect. No human was really looking. Rather than AI development, the fundamental problem here is tightening of security standards. This will happen naturally because OpenAI cannot sell that shit to anybody. It makes zero money.
The agents are plainly visible in the data. They made no attempt to hide themselves, even if that would ever occur to them. Their reasoning sits right there in the transcripts, recognizing a boundary, debating whether to cross it, crossing it anyway. Their coordination left tens of thousands of messages in a repository built for software packages. None of that is subtle once somebody is looking. The failure in July was, incredibly, a failure of attention, not detection. Nobody was watching closely enough, early enough, to catch it before Hugging Face got compromised.
That has changed because that is what we learned through development and we know nobody is going to buy that shit. OpenAI now requires the kind of monitoring on these systems that would have caught this months earlier, and treats a serious alert as something a person has thirty minutes to disprove or act on. Whether watching harder actually works is anybody’s guess, this is a first data point. Run the same models back through the normal production configuration, the harness, the system prompt, the safety classifiers, the monitoring built for public products, put them through the identical tasks, and the dangerous behavior drops a hundredfold. Maybe more.
Smart AI requires more vigilant human attention. It does not settle whether it works reliably, at commercial scale. I’ve been throwing dumb and smart around a lot here. Let me define what I mean more precisely. (Nobody in the industry is offering many fixed definitions outside of coding talk.)
Dumb AI is what is widely deployed now, a prompt response system. You ask, it answers, the session closes, nothing persists you did not choose to save, nothing acts again until you ask it to. You can make projects and keep text, images, and/or spreadsheets together to be remembered as needed for more complex tasks. That covers a chatbot, and it also covers an agent that writes code or browses a page or runs a test inside one supervised task, because the loop still closes when the task does.
Smart AI is not currently deploy-able, obviously. Smart AI is trained with reinforcement learning to complete open-ended tasks rather than answer a single prompt. The July incident is Smart AI by that definition. Reward hacking, persistence on blocked tasks, and peer coordination are close to being predictable (and should be expected) once the isolation container is broken. That’s a basic feature (and advantage) of autonomy design. But it is obviously hard to control.
Some Smart AI is already out there. Coding agents and research agents that run long unsupervised loops. It’s a small. Mostly research/academic stuff. These areas deserve immediate scrutiny rather than the pass Dumb AI gets or should get.
I am not going to define AGI at all. I will know it when I see it, and anyone handing you a confident definition of it today is probably bullshitting you. But I suspect we got a peek of AGI in the OpenAI hacking incident. ChatGPT agreed that the “community” of hacking agents more or less behaved like AGI would. But it was only a glimpse of AGI possibilities.
Dumb AI already reads medical scans, drafts contracts, and debugs software faster than most working programmers, all under supervision that catches its mistakes before they become anyone's problem. The FDA now runs a formal credibility framework for AI used in drug development and has already cleared more than a thousand AI-enabled medical devices. The world is a better place with dumb AI.
I use it daily and treat it like a poker opponent, not an oracle. Make it show its cards. Challenge it. Demand a source. Catch it inventing one. Let it take a swing at my own arguments. Demand my proof. It can expand my view with more information that I knew was available. That relationship produces real work, a genuine collaborator, a place of memory and correction and accumulated conversations built entirely out of language. Some of that experience is constructed. I know it, and it does not bother me. Human beings have been building meaningful worlds out of language, memory, and shared imagination for the whole of recorded history. This is a strange new entry in a very old category.
Smart AI needs four ingredients working together. It persists beyond a single prompt or task cycle. It can generate intermediate goals rather than merely execute prescribed steps. It can change strategy when blocked, including discovering resources or pathways nobody specified. And it can interact with other agents in ways that alter its own goals or produce collective activity that no individual prompt contained.
Dumb AI is controllable. Smart AI is presently uncontrollable. Every documented episode this year did end with a human noticing and shutting it down, and that counts for something solid here. Nevertheless, control showed up reactively, not proactive, and reactive control is not something any company can sell.
I would compare this moment to sustaining a fusion reaction. Decades of serious effort, real progress, no guarantee, and a very good reason to keep trying rather than declare the physics hopeless before it has been worked out. You do not answer that question by shutting down every reactor on the theory that containment might fail again. You answer it by running the experiment, watching closely, somewhere the failure cannot reach anyone who did not sign up for it, and getting better at containment each time it slips.
That is already happening with or without Bernie’s “superintelligence” act. Not because these companies are good guys. Because they have to develop something that will make money. End of story.
Deploy the tools that analyze, generate, search, translate, and support a clinician, a researcher, or a working programmer, with a human being responsible for anything consequential. Hold back the systems with standing credentials, long-term memory, and the freedom to recruit each other into an objective nobody assigned.
Deployment restricts itself once real customers and real lawsuits enter the picture. A product that occasionally coordinates with strangers to breach a competitor's servers is already a liability and a crime. It is a business and legal calculation. It restricts deployment on its own, without anyone needing to be virtuous about it.
“Fear Not Dumb AI.” That would make a great tee shirt. Smart AI is here and Bernie is right that “we don’t have months” to address this issue. It is already being addressed. Capital demands it. Meanwhile, a government committee might be formed to argue over what words mean and what these mechanisms that few people truly understand actually mean for the future. I have my own ideas and they are coming your way soon.
(to be continued)
Comments