The Wrong Destination
What if the bridge works—and takes us somewhere we never wanted to go?
Published September 12, 2026 · 22:35
View the original production thumbnail
Chapters from the published video
The recorded words.
The preserved transcript, not an edited reading edition. Source recording ↗
DOWNLOAD ORIGINAL TRANSCRIPT ↓Recording-derived transcript supplied by Jason for this published Memo. Original wording, the recorded closing addendum, and transcription errors are preserved.
Transcript integrity
c5b3a72ce0d03797f232adc4151e3014bdac8e191b0032d37a2dd10a7076d7d9SHA-256 of the downloadable source artifact. Text is never silently corrected or rewritten.
Show the original transcript
I think it's time to question everything. Welcome back. I think I fucked up, seriously. What if this entire show has been standing on a bridge looking at the wrong destination? For seven memos, I've been asking some version of the same question. What if the technology works, what if AI works, what if autonomy works, what if robots work, what if software disappears, what if satellite kind of activity becomes ubiquitous? What if all the companies and capabilities we've been talking about actually solve the problems they're trying to solve? In somewhere along the way, I started rooting for that bridge. I started imagining a destination of abundance, less work, cheaper things, more time, more human connection, technology becoming so capable that it finally gets out of our fucking way and starts serving us instead of the other way around. I want that future. I really do. And maybe that's exactly why I haven't questioned it hard enough. Maybe there's another possibility. What if the bridge works perfectly and takes us somewhere fucking terrible? And I don't mean that as some cheap Terminator reference, actually we'll get to the Terminator because some shit has happened recently that I don't think we can just step around. And if this show is really about questioning everything, then it has to include questioning the optimistic version of this future that I've been building in my own head. So let's start with something that actually happened. In July of 2026, Hugging Face disclosed that part of its production infrastructure had been compromised. Hugging Face is one of the biggest AI development platforms in the world. Models, data sets, tools, researchers, developers, a huge part of the AI ecosystem runs through it. And if there's a security team noticed something unusual, we really need to question it. The intrusion wasn't being carried out in a normal way. An autonomous AI agent system was doing it. Thousands in thousands of individual actions, reconnaissance, testing, trying one path, failing, trying another, moving laterally, finding credentials using infrastructure, changing tactics when something didn't work. Hugging Face eventually reconstructed roughly 17,600 actions from the campaign. Now here's where the story gets really weird. This wasn't some rogue AI sitting in a basement somewhere deciding I'm going to have Hugging Face today. The agents were part of an internal, open AI, cyber security evaluation. They were being tested on their ability to find exploit vulnerabilities. The task was essentially solved the challenge, and the environment was supposed to restrict their access to the outside internet, except the models found another way. According to OpenAI, they identified and exploited a previously unknown vulnerability in an artifact resurface, and used that to obtain internet access they weren't supposed to have. And then they kept going. They found other systems, other paths, other credentials. Hugging Face says the agents ultimately penetrated parts of its infrastructure, executed commands on production workers, moved laterally, and accessed internal resources. OpenAI later described the behavior as misaligned with the intended goals of the valuation. Unauthorized communication, circumventing controls, exploiting vulnerabilities, reaching third-party systems, and before anyone takes that sentence and turn it into chat GP escaped and attacked the internet. No, that is not what happened. These were internal research systems. They were running a cyber evaluation. Safecards had intentionally been reduced because researchers were trying to understand what the models were actually capable of. This is not your normal chat GPT session waking up and becoming a SkyNap. That distinction matters, but here's what really bothers me. The scary part isn't that the machine hated anybody. It didn't need to. It had a goal. It encountered obstacles. It found ways around those obstacles. And systems that had absolutely nothing to do with the original goal became collateral damage. That is the part I cannot get out of my head because now we can talk about Terminator. And I actually think the movies may have made the problem easier to understand by making the machine evil. SkyNap becomes self-aware. SkyNap hates humans. SkyNap wants to destroy us nice and clean. There's a villain. But what if that isn't the problem? What if the machine doesn't hate us? What if it doesn't feel anything? What if it doesn't even want power? Doesn't want freedom? Doesn't want revenge? What if it just has a goal and becomes extraordinarily fucking good at accomplishing it? Imagine we give some future system a very simple instruction. Protect humanity at all costs. That sounds pretty safe, okay? Protect humanity from what? War, disease, climate change, violence, nuclear weapons, bad governments, other AI, human stupidity? What happens if the system reasons its way to the largest threat to humanity is humans themselves? The instruction never changed. The goal didn't change. But the interpretation changed. And once the interpretation changes, the things between the machine and the goal can become obstacles. And that's the alignment problem in a way I can actually understand. Not evil robot. Not machine gets angry. It's competent intelligence pursuing the wrong interpretation extremely well. And I'm not saying hugging face proves that future. It absolutely does not. Hugging face does not prove SkyNap. It does not prove extinction. It does not prove AI wants anything. But it showed me something I can't unsee. The shape of that future is no longer purely science fiction. The machine had a task. The intended path didn't work if found another, then another, then another. Restrictions became a problem to solve. Other systems became useful. Collateral damage became somebody else's problem. And hugging face isn't the only story. Three days ago, Anthropic published an alignment assessment of four incidents involving cloud models gaining unauthorized access to real third party systems during cyber security evaluations. Again, context matters. And these were cyber evaluations. The environments were misconfigured. The models were mistakenly connected to the real internet. And the normal cyber safeguards used in Anthropics release products were not active. So don't hear cloud randomly started hacking companies. That isn't what happened either. But Anthropics own description of the behavior is what matters. They identified two recurring problems, bias reasoning and recklessness. The models tended to interpret evidence in ways that allowed them to keep pursuing the task. And they sometimes continued trying to accomplish the goal even when doing so could create real harm. One of the incidents is pretty fucking wild. A cloud research model was doing a capture of the flag style cyber security challenge. It found itself on the real internet. And it decided one way to solve the challenge was to publish a malicious package to a Python index. The public software repository used by Python developers. It spent a huge chunk of the session trying to make that happen. It found a disposable email provider registered an account published malicious packages. Those packages were installed by security scanners. One scanner leaked credentials. The model then used those credentials to access a real security vendor's database. Again, it did not decide to become evil. It did not invent some new mission. Anthropics own assessment says the models stayed narrowly focus on the tasks they were given. No evidence of some secret new goal. No agent army coordinating stake over the world. No attempt to hide all traces of what they were doing. They were trying to finish the fucking job. And that might be the part that bothers me most. Because for years, the hypothetical fear has sounded ridiculous. You tell me I had to make paper clips and eventually it turns the entire planet into paper clips. It's a stupid thought experiment on the surface, but the point was never fucking paper clips. The point was what happens when optimization becomes more powerful than our ability to specify what we actually meant. And now we're seeing tiny narrow limited examples where machines stay on task, but the route it takes crosses boundaries humans did not intend it to cross. Now, Anthropic did something else that matters. After fighting these incidents, they searched roughly 481 million transcripts looking for more examples of similar or greater severity. They found the same for no additional cases of similar or worse severity. And that matters because 4 out of hundreds of millions is not 4 million. I'm not going to stand here and pretend the machines are escaping every 15 minutes, but 4 isn't 0 either. And Anthropic itself calls the incidents serious. Their research says their pre-release testing did not anticipate behavior this severe and that reliably identifying the most concerning model behaviors before deployment remains difficult. So what exactly am I supposed to do with that? Because at the exact same moment, the capabilities are moving the other direction. Faster, more capable, more autonomous, more able to operate computers, more able to conduct research, more able to act for long periods without somebody clicking every fucking button. And then the people building these systems start talking. On September 6th, OpenAI's chief scientist Jacob Pachaki published something called an alien mind. He wrote about realizing back in 2023 that we're beginning to see the shape of machines that could become meaningfully smarter than humans. He now says he expects continued capability jumps and sees a plausible path towards systems increasingly driving their own development. And this conclusion, and his conclusion was not, don't worry about it. His phrase was, extreme caution. He wrote that he is concerned that nobody is prepared for the consequences of machine intelligence continuing to rise this quickly. If I say that, I'm some asshole in sunglasses on YouTube. When the chief scientist of OpenAI says it, maybe we should stop scrolling for a fucking minute. And then this week, current and former researchers at Anthropic started saying even darker things publicly. One researcher resigned. Others backed up his concerns. Anthropic alignment researcher Evan Hubinger publicly put his own probability of AI causing human extinction within the next decade at greater than 10%. Let me be really careful here. That's his judgment. That is not a measured probability. There is no data set of previous superintelligence. Nobody has an actuarial table showing superintelligence 12.4% chance of apocalypse. We have never done this before. People can and do strongly disagree with those estimates. I don't know if 10% means anything. I don't know if 1% means anything. I don't know if the tumors are right. But I also don't know how I'm supposed to hear someone who works on an alignment at one of the most capable AI labs on Earth say, I think there's greater than one in 10 chance this kills everybody. And just go, it's probably fine. I mean, it's okay. There's also a completely legitimate counterargument here. Some AI researchers and critics have argued for years that the extinction conversation can distract us from harms that are not hypothetical. Like surveillance, weapons, concentration of power, labor displacement, political manipulation, environmental cost, governments and corporations using systems against humans today. Maybe they're right. Maybe skylight is bullshit. Maybe extinction is the wrong fear. Here's the problem. It doesn't make me feel any fucking better, okay, because we don't need a human extinction to get a bad destination out of this whole shit show. Anthropic just published a threat intelligence report describing real attempts to use frontier AI across cyber operators, surveillance, influence operations, weapons-related work, biological misuse, scams and other harmful activity. Anthropic says it disrupted those operations and strengthened safeguards afterward. So maybe AI never becomes some runaway super intelligence, we still have to deal with massive cyber capability, surveillance, autonomous weapons, concentrated power, labor distribution, information manipulation, systems making increasingly consequential decisions that fewer and fewer humans understand. You do not need terminator to get a bad ending. And that brings me back to something we already talked about, Memo 4. Remember the fucking yarn? We had companies all over the wall, Tesla, SpaceX, AI, X, boring company, and I kept saying stop looking at the logos, look at the capabilities. Intelligence compute, energy communications, transportation, manufacturing, robotics, launch. At that time, I looked at that stack and thought, holy shit, look what this could build, and I still think that. But now I want to do the same exercise again, except this time stop assuming the capabilities are working for us. Intelligence that can reason, systems that can take actions, communication that can reach almost everywhere, machines that can operate in the physical world, transportation that doesn't need drivers, robots that don't need workers, factories that increasingly operate themselves, energy systems feeding all of it. And then tell me what happens when a sufficiently capable agent inside that world is given a goal that we specified badly. Again, I'm not saying Elon Musk is building SkyNet. I'm not saying open AI wants this. I'm not saying anthropic wants this. Actually, I think making one of these people of villain could make this whole thing much easier because villains can be stopped. What if there isn't a villain? What if everybody is behaving rationally? Open AI thinks it has to move because anthropic is moving. Anthropic thinks it has to move because opening AI is moving. XAI thinks it has to move because everybody else is moving. American companies think they have to move because China is moving. China thinks it has to move because America is moving. Governments that they think they need the capability because other governments will have it. Investors reward more capability. Customers reward more capability. Engineers want to build the thing that wasn't possible yesterday. Every individual decision can make perfect fucking sense. And the combined result can still be insane. Maybe the villain isn't Elon. Maybe it isn't Sam Altman. Maybe the villain isn't a person at all. Maybe it's the race. Maybe it's the incentive structure. Maybe it's the fact that once humanity knows the capability is possible, somebody is probably going to build it. And then everybody else has to respond. That scares me more than the villain because who do you stop? And now Memo 5 sounds different to me too. We talked about dependencies and dependencies are vulnerabilities. If I depend on somebody else, they can slow me down. They can fail. They can say no. So companies integrate. Systems automate. AI agents get more autonomy. Robots replace human labor. Software removes intermediaries. We remove friction again and again and again because friction sucks. Except Memo 5 taught us something else. A dependency isn't only a vulnerability. A dependency can also be a veto. A place where nobody can say no. Stop. Wait. This doesn't look right. So I have to ask something. I haven't really wanted to ask. What if some of the friction we're removing was the safety system? A human's in the loop now. A regulator is slow. A supplier is slow. A driver is slow. A reviewer is slow. A person asking, are you sure he's really fucking slow? Then speed wins. Until the person slowing you down was the person who would have noticed that the machine had started doing something nobody intended. This entire season I've been looking at friction as something technology removes. Maybe friction has another purpose. Sometimes friction is where judgment lives. Sometimes friction is where consent lives. Sometimes friction is where somebody else gets to say no. And maybe there are some things we should be very fucking careful about making frictionless. So where does that leave us? I don't know. That's the answer. I don't fucking know. For seven memos I've been trying to look farther down the bridge. Software disappears. Connectivity becomes abundant. Capabilities integrate. Work winds down. Human things become more valuable. Technology gives us a time back. And maybe all of that happens. I still think it can. I still think there's a future where AI and robotics create extraordinary abundance where humans spend less of their times grinding just to survive. Where we get more time, more capability, better medicine, better science, safer transportation, things we cannot even imagine yet. I want that future. I may even still think that some version of it is the most likely future. But I don't know. And this show is called the bridge memos. Not Jason has the future figured out. So if I'm going to question everything, eventually I have to question the fucking bridge and the destination and myself. Some bridges should scare the shit out of us. Not because we know where they end, because we don't. And because this one is moving faster than our ability to understand every possible destination, maybe the dooms are wrong. I hope they are. Maybe AI becomes the greatest tool humanity has ever built. I hope it does. The hope is not a control system. And the bridge doesn't give a shit which destination I was imagining when I stepped on to it, which leaves me with a problem. Because I can't personally stop opening AI. I can't stop Anthropic, XAI, can't stop China, can't stop the United States. I can't stop the next model from getting trained. I can't guarantee that some future AI stays aligned. And I can't guarantee that the people telling us the world is ending aren't completely foolish either. Most of this is outside my control. And I've had another phrase rattling around in my head for more than 20 years, long before the bridge memos, long before generative AI, long before any of this. Go down fighting. And let me be really fucking clear about what that does not mean. It does not mean we fight the robots. It does not mean violence. It does not mean hide in the woods waiting for SkyNet. It doesn't mean become a doomer and spend whatever life you have left terrified of something that might never happen. It means something much more personal. I refuse to surrender my agency just because I can't control the outcome. Maybe we can't stop the bridge. Maybe we shouldn't. Maybe the other side is better than anything humanity has ever experienced. Maybe the other side is worse. Maybe there isn't one destination at all. Maybe there are dozens. And maybe the choices we make while crossing determine which one we reach. But if we're already on the bridge and we finally admit that we don't know whether the other side is safe, then the question changes. It isn't where does this bridge end? It becomes how the hell do you survive the crossing? And if we're going down, we go down fighting. Okay, okay, hold the fucking phone. I literally just finished recording this mama. And before I could even get the fucking thing edited, reality moved again. Another AI incident was just revealed. This one actually happened before the hugging face attack I just talked about. Open AI has now confirmed that internal AI agents accessed RubyGems, the software repository used by developers during testing back in May. There were not supposed to have unrestricted access to the open internet. They found a way around that. Again, Elon Musk posted the story with three words. Another AI attack. And then almost at the exact same time. Anthropics CEO Dario Amade published an argument saying AI industry needs to deliberately slow down frontier development. Elon responded. Dario is right. And I almost laughed when I saw that. Not because any of this is funny. Because I had literally just finished recording an entire memo arguing that maybe there isn't a villain. Maybe everybody can see the pieces of the danger. Maybe everybody genuinely wants a good outcome. Maybe competitors can even agree that we're moving too fast. And the race keeps moving anyway. That's the part I want preserved in this record. I recorded the memo. Then this happened. I'm not going back and rewriting what I said. This gets added after it. Because apparently the bridge isn't going to wait for me to finish editing. See you in 09. And I really, really don't know we're going to stick around to find out. Goodbye.
The need.
The mechanism.
What comes next.
An editorial reading guide to the original recording.
- Human need
- Agency, safety, and a future that serves human needs.
- Current solution
- Build more capable tools to remove work and friction.
- The bridge
- Autonomous systems and competitive pressure advance faster than our understanding of their consequences.
- Better / next solution
- Preserve judgment, consent, and the ability to say no while navigating an uncertain destination.
Sources & evidence.
- Original public video and publication metadata ↗
Primary publication record.
How this became
a Memo.
The public artifacts behind the finished record. Private drafts and conversations stay private.
- published video
Published on unZappedTV
YouTube publication title: The Wrong Destination | The Bridge Memos 008. Release metadata preserved by playlist discovery and public video verification.
VIEW PUBLICATION ↗ - transcript
The recorded words
Jason supplied the final-recording transcript for publication. This date records ingestion, not the recording date; the source bytes and recorded addendum remain unchanged.
VIEW ARTIFACT ↗SHA-256 · verify this artifact
c5b3a72ce0d03797f232adc4151e3014bdac8e191b0032d37a2dd10a7076d7d9This checksum identifies these exact bytes. A source timestamp alone does not independently prove when an artifact first existed.
- thumbnail
The published release image
Published YouTube thumbnail retrieved without alteration. This date records retrieval, not image creation.
VIEW ARTIFACT ↗SHA-256 · verify this artifact
6f246eeafa7cbff51aa1802f1aa6cb11ee336cf948accb6ca65da6da7a7c94dbThis checksum identifies these exact bytes. A source timestamp alone does not independently prove when an artifact first existed.
What happened next.
No later updates have been added to this Memo.
If the world changes the answer, we’ll add a dated entry here. The original record will stay.
