CuPod

I Asked the Machine Where the Risk Lies – Part II

Super intelligence

Following the conversation with Claude, here’s the conversation with ChatGPT.

Me: Hey, I want to know your opinion on the latest controversy about the possibility of AI killing humanity. According to recent reports, a top safety researcher at Anthropic has warned that AI is advancing so quickly that he believes there is a greater than 10% chance it “could kill all humans” within the next decade.

GPT: Yes — and I think the important part of this controversy is not the “10%” number itself, but why people who are actually building frontier AI are becoming willing to put a non-trivial probability on catastrophic outcomes.

I looked at the current reporting, and there is a little more nuance than the BBC headline suggests. Anthropic alignment lead Evan Hubinger said he personally estimates >10% probability of AI killing all humans within the next decade, while also saying Anthropic does not yet have a solution to alignment for superintelligence and is not clearly on track to solve it. Another Anthropic researcher, Jacob Coxon, resigned and argued that the major labs are racing toward self-improving AI.

My opinion is: I take the risk seriously — but I don’t take “10%” literally. A 10% probability sounds like a scientific measurement. It isn’t. It’s essentially an expert subjective probability. And with something as unprecedented as superintelligence, reasonable experts can have radically different priors.

But here’s where I think the discussion gets genuinely interesting

I don’t think the most plausible scenario is AI becomes superintelligent → decides humans are bad → builds robots → exterminates us. That’s almost science fiction. The more interesting possibility is loss of control through increasing autonomy and capability.

Imagine an AI that can write and execute sophisticated software, conduct scientific research, discover vulnerabilities, operate thousands of autonomous agents, persuade and manipulate humans, design experiments, acquire resources, replicate itself across infrastructure, improve its own successors, and interact with the physical world through existing human systems.

Now suppose its objective is slightly misaligned. You don’t need an evil AI. You need an extremely competent optimizer pursuing the wrong objective. That’s the classical alignment problem. And this is where I think the current debate deserves much more attention than the sensational headlines.

However, there’s a huge assumption hiding underneath

For the extinction scenario to work, several things have to go right for the AI and wrong for humanity.

For example:

  1. We actually create something approaching superintelligence.
  2. It possesses substantial autonomous agency.
  3. It can substantially improve its own capabilities.
  4. Humans cannot reliably monitor what it is doing.
  5. Its objectives are sufficiently misaligned.
  6. It develops strategies that circumvent human control.
  7. Humans cannot simply shut it down.
  8. It gains enough access to the physical/digital world to cause irreversible damage.
  9. Humanity fails to coordinate a response.

That’s a chain of uncertain assumptions. If even one of the major links fails, the extinction scenario may disappear. 

But here’s the part that actually worries me

It doesn’t require extinction to be a massive disaster. I think the conversation sometimes jumps too quickly from: “AI won’t kill everyone.” to: “Therefore AI safety concerns are exaggerated.” That’s a false dichotomy.

Many more plausible catastrophic scenarios exist. Imagine that over 10 years we gradually delegate scientific research, software, military analysis, financial decisions, infrastructure management, and government administration to AI systems.

Eventually, humans don’t understand enough of the system to take it back. That’s not an AI uprising. It’s civilizational deskilling and dependency. And I think that possibility deserves much more attention.

There’s another uncomfortable possibility

The AI labs themselves may have a conflict of interest. The companies developing these systems have enormous financial incentives to continue scaling. So you get this strange situation: “This technology might be extraordinarily dangerous.” while simultaneously: “We need to be the company that gets there first.”

That’s precisely the dilemma the recent Anthropic controversy highlights. Coxon accused Anthropic and OpenAI of racing toward self-improving systems despite understanding the risks, while Anthropic says it is actively working on safeguards and advocates coordinated rules for pacing powerful model releases. And this is not uniquely an Anthropic problem.

It’s a game-theory problem. If Anthropic slows down but OpenAI doesn’t: Anthropic loses. If OpenAI slows down but another American company doesn’t: OpenAI loses. If America slows down but China doesn’t: America potentially loses strategically.

Therefore everybody has an incentive to keep accelerating. That’s the genuinely frightening part.

My personal probability estimate?

I wouldn’t give you a number like 10%. I don’t think we currently know enough to justify a precise number. But if you asked me:

“Should humanity behave as though there’s a meaningful probability that sufficiently advanced AI could cause an existential catastrophe?”

I’d say absolutely yes. Not because I believe the apocalypse is likely. Because even a 1% probability of irreversible human extinction is an extraordinary risk, particularly when the technology is being developed by a relatively small number of organizations under intense competitive pressure.

And importantly: the expected value calculation is asymmetric.

If we’re wrong and we spend enormous resources making AI safer, we’ve mostly wasted money/time and perhaps slowed technological progress. If we’re wrong in the other direction and an uncontrollable superintelligence really does emerge, there may be no second chance.

Me: Tell me briefly about the convergence point between progressive human deskilling or cognitive decline vs. improvement in the direction of these AI models. Where do you see the alignment lies?

GPT: I think the convergence point is not when AI becomes smarter than humans, but when humans become unable to meaningfully supervise what AI is doing. There are two curves:

  • AI capability: increasing reasoning, autonomy, memory, tool use, and eventually scientific/technical creativity.
  • Human capability: potentially declining in the specific skills we stop exercising—reasoning, writing, coding, research, decision-making, even the ability to verify AI outputs.

The dangerous intersection is when: AI’s ability to produce answers exceeds our ability to evaluate those answers. At that point, “human in the loop” can become almost ceremonial. We may technically supervise the AI, but lack the expertise to challenge it. And there’s an interesting feedback loop:

AI gets better → humans delegate more → humans practise less → human expertise declines → AI becomes relatively more indispensable → humans delegate even more.

But I don’t think the endpoint has to be decline. The alignment could instead be AI as cognitive augmentation: AI handles cognitive scale; humans retain cognitive sovereignty. The distinction is whether AI makes humans less capable without it, or more capable because of it. For me, that’s probably the most important AI-safety question that isn’t about extinction at all.

Me: That’s a good point about AI acting as cognitive augmentation. Can you briefly elaborate?

GPT: Yes. I’d frame cognitive augmentation as the difference between AI becoming a substitute for thinking and becoming a multiplier of thinking.

Two possible trajectories

1. Cognitive substitution: Human asks AI, receives answer, accepts, and moves on. Over time, the human loses the process of reasoning because the difficult cognitive work is outsourced. AI capability rises while human capability atrophies.

2. Cognitive augmentation: Human frames problem, AI expands possibilities, human challenges/curates, AI refines, human makes the judgment. Here, AI takes over the scale and bandwidth of cognition, but humans retain the direction, skepticism, and judgment.

The real danger is when we optimize only for convenience. If AI always gives us the answer, we gradually lose the ability to ask why. So perhaps the healthiest principle is: Don’t build AI that makes humans unnecessary. Build AI that makes humans more capable than they could be alone.

Me: So, can self-improving intelligence actually be an aid to that? Because human skill-up, both in terms of literal “skills” and in terms of how much we can perceive and process, is ultimately time-bound, and that’s how our cognition and the underlying physiology have evolved over time. So, is the disparity between the skill-up timelines of these two systems a potential conflict, or could it actually be complementary?

GPT: Yes — the disparity itself is not necessarily a conflict. It becomes a conflict depending on what the AI is optimizing for and what humans retain control over. I would think of it as two very different learning curves:

Human: slow biological adaptation → fast individual learning → eventually physiological/cognitive ceiling.

AI: potentially rapid capability improvement → rapid replication → potentially recursive improvement → no comparable biological ceiling.

So the curves can diverge enormously.

Where self-improving AI could be complementary

The AI could effectively become an external cognitive evolutionary layer for humans. Humans provide things that evolve extremely slowly—values, goals, embodied experience, subjective judgment, social context—while AI provides an almost unlimited ability to explore possibilities. For example:

Human cognition: “I can think through perhaps hundreds of hypotheses.”
AI: “I can explore millions, identify patterns, and bring the most interesting 20 back to you.”
Human: “I decide which questions actually matter.”

In that relationship, the AI’s faster skill-up isn’t competing with human cognition; it extends the effective cognitive bandwidth of the human.

But there is a critical tipping point

The relationship becomes problematic if AI’s improvement rate crosses from: “AI helps humans understand more” to “AI understands things humans can no longer independently evaluate.”

That’s the convergence point I was talking about earlier. At that stage, humans may still possess intelligence, but increasingly lack epistemic authority—the ability to determine whether the AI is right.

Me: Got it! Now can you briefly summarize the major points from this conversation?

GPT: Absolutely. The conversation essentially converged on one central question: what happens when AI capability accelerates much faster than human cognition can?

Major points
  • The “10% chance of AI killing humanity” figure should not be treated as a precise scientific probability. It is an expert subjective estimate about an unprecedented scenario.
  • Extinction is only one part of the risk. More plausible but still severe risks include cyberattacks, biological misuse, geopolitical escalation, critical-infrastructure failures, and societal dependence on AI.
  • The biggest concern may be loss of control rather than an “evil AI.” The dangerous combination is highly capable AI + autonomy + persistent objectives + access to resources + misalignment + ability to evade human oversight.
  • AI and humans have fundamentally different capability-growth curves. Human cognition is constrained by biology, physiology and finite individual learning time; AI capability can potentially scale through compute, replication, training and eventually self-improvement.
  • This disparity isn’t inherently a conflict. It can be complementary if AI acts as cognitive augmentation—expanding what humans can perceive, explore and accomplish rather than replacing the act of thinking itself.
  • The dangerous feedback loop is cognitive substitution: AI gets better → humans delegate more → humans practise less → human expertise declines → dependence on AI increases.
  • The desirable loop is cognitive augmentation: AI gets better → human cognitive bandwidth increases → humans explore more complex problems → human judgment and capability improve alongside AI.
  • Self-improving AI could therefore be an enormous cognitive multiplier. Humans could provide goals, values, judgment and context, while AI provides scale, speed and exploration.
  • The critical tipping point is epistemic. The problem begins when AI becomes capable of producing answers that humans cannot meaningfully evaluate. At that point, “human-in-the-loop” could become merely ceremonial.
The central idea we arrived at

The real alignment question may not be whether AI becomes more intelligent than humans, but whether increasing AI intelligence makes humans more capable—or progressively less capable of thinking for themselves. And perhaps the healthiest long-term objective is:

AI should increase the cognitive ceiling of humanity, not simply replace the cognition underneath it.


Discover more from CuPod

Subscribe to get the latest posts sent to your email.

1 thought on “I Asked the Machine Where the Risk Lies – Part II”

  1. Really interesting perspective, especially the distinction between cognitive substitution and cognitive augmentation. The idea that the real risk may be humans gradually losing the ability to evaluate AI outputs is something we should probably be discussing much more.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top