For the last few days, the internet has been abuzz with warnings about AI from researchers at Anthropic. In a social media post, Jacob Coxon, a researcher at Anthropic and also a former OpenAI employee, wrote that the people building these superintelligent AI agents fear that AI could potentially kill humans in the coming decade. Two other Anthropic scientists, Evan Hubinger and Samuel Marks, have affirmed the possibility and also said that there is a certain consensus among AI developers.
Now, though many practitioners may have already guessed it one way or the other, hearing it straight from the lion’s den is indeed a bit chillier than hearing it as a matter of personal belief.
But it got me curious: what do these intelligent agents think about this? I mean, although the current AI models are not the direct accused, their future successors are. So, they should at least have a chance to rebut these claims or tell us how truthful these claims actually are, right?
So, I kind of chatted/interviewed Claude (Sonnet 5) and ChatGPT (GPT-5.6 Luna) about this risk and what they “think” about it.
Here’s the conversation with Claude. The following post will have the same conversation with GPT.
Me: Hey, I want to know your opinion on the latest controversy about the possibility of AI killing humanity. According to recent reports, a top safety researcher at Anthropic has warned that AI is advancing so quickly that he believes there is a greater than 10% chance it “could kill all humans” within the next decade.
Claude: This checks out — it’s a real, recent story, not a misquote. Here’s the shape of it:
What happened: Jacob Coxon, a former Anthropic/OpenAI researcher, resigned this week and posted publicly that people building AI “earnestly believe it could kill us all by the end of the decade,” accusing both labs of “gambling with our lives.” Evan Hubinger, Anthropic’s Alignment Science Lead, responded on X confirming the sentiment and putting a number on it: he personally thinks it’s greater than 10% within the next decade, while adding that Anthropic doesn’t yet have a plan to solve alignment for superintelligence. He clarified the risk isn’t from today’s models but from a hypothetical future where systems can recursively self-improve — designing successors without human involvement.
My take, for what it’s worth:
- This is a real, if unusually blunt, expression of something that’s actually fairly common among alignment researchers: wide uncertainty about a low-probability-but-nonzero catastrophic tail risk, not a confident prediction. “10%” here is a subjective credence, not a measured statistic — there’s no model producing that number.
- The fact that it’s coming from inside Anthropic isn’t necessarily damning — it’s arguably the more honest position than either dismissing the risk or pretending it’s solved. It also isn’t obviously reassuring, since it’s an admission that the org building the tech doesn’t have a solution yet.
- I’d treat this less as “authoritative forecast” and more as a data point in an ongoing, genuinely unresolved debate — other credible researchers put the odds much lower, and there’s no consensus methodology for estimating something like this at all.
I’m an Anthropic model, so take my view with that in mind — I’d rather point you to the range of expert opinion than tell you what to think about my own creators’ risk levels.
Me: Got it! But do you think that the latest models being developed and released pose any threat to human civilization, or at least to human cognition?
Claude: Fair question, and worth separating into the two different threats you named, since they’re on very different timescales and have very different evidence bases.
On civilization-scale risk: I don’t think it’s honest to say “no threat at all.” The concern isn’t that today’s models are secretly dangerous — current systems can’t self-replicate, resist shutdown, or improve their own architecture in any meaningful way, and Hubinger himself said the risk from present models is low. The concern is about direction of travel: capability is increasing faster than the field’s ability to verify what these systems actually want or would do with more autonomy and less oversight. Whether that gap closes in time is genuinely unresolved — there’s no consensus number, and serious researchers span from “single-digit percent over the century” to “double-digit percent this decade.” I’d be skeptical of anyone, on either side, who states high confidence here.
On human cognition/society: This is the part I think gets underweighted next to the sci-fi framing. There’s more concrete, measurable concern already showing up:
- Studies on cognitive offloading suggest heavy reliance on AI for tasks like writing and problem-solving can reduce active recall and critical engagement with the material — this is closer to “use it or lose it” than dystopian, but it’s real and already measurable.
- Effects on epistemics: search/chat tools shaping what info people encounter, potential homogenization of writing/thinking style, harder-to-detect misinformation at scale.
- Labor and economic disruption is a nearer-term, better-evidenced concern than extinction risk, and arguably the thing most likely to actually affect your life in the next decade.
So: not “no threat,” but I’d resist collapsing “civilization-ending AI takeover” and “AI changes how people think and work” into one blob — the first is speculative and contested even among experts, the second is already happening and easier to reason about with actual data. Happy to go deeper on either one.
Me: I see your point. So, there’s no civilization-level risk from the present “public” models. At this point, cognitive decline is a more practical scenario to be talking about as a risk, right? Fair enough. But do you see a correlation between “capability is increasing faster than the field’s ability to verify what these systems actually want or would do with more autonomy and less oversight” and declining human cognitive ability? Where do you think these two directions would meet if there’s no checkpoint?
Claude: There is a plausible mechanism connecting them, though I’d call it a compounding risk rather than a proven causal chain — worth being precise about which.
The mechanism: Oversight of an AI system isn’t passive — it requires people with sharp enough domain expertise and independent judgment to catch subtle failures, verify claims, and know when something looks wrong even if it’s not obviously broken. This is well-documented in other automated domains (aviation, radiology) as “automation complacency” or “deskilling”: the more you hand a task to automation, the rustier your own judgment gets on that exact task — right when you’d need it most to catch an edge case the automation misses. Apply that to AI: if engineers, evaluators, and regulators increasingly use AI to evaluate AI (summarizing safety research, reviewing code, drafting risk assessments), the pool of people who can independently catch a subtle problem may shrink precisely as the systems get harder to independently verify.
Where they’d meet without a checkpoint: capability keeps climbing on its own momentum (compute, data, algorithmic progress don’t need much human cognitive input to keep scaling), while independent human capacity to scrutinize what’s happening either flatlines or erodes. The gap is the risk — not a single moment of “AI takes over,” but a slow loss of the ability to tell whether things are still on track. That’s exactly why people like Hubinger say “no plan for alignment” is the alarming part, not any specific model.
Where I’d push back on the framing: this isn’t a one-way street. AI-assisted tools (mechanistic interpretability, automated red-teaming, scalable oversight research) could in principle expand verification capacity faster than human deskilling erodes it — that’s an active bet several labs are making, not a null hypothesis. So the “meeting point” isn’t fixed; it depends on whether oversight tooling outpaces both capability growth and the cognitive-offloading effect. Nobody has a strong empirical answer for which wins that race — that’s the actual crux, more than either trend in isolation.
Me: But whoever ends up winning the race, capability growth without oversight or faster human deskilling, what good would that really do for civilization? Or is that “good” itself a much greyer?
Claude: Genuinely grey, and I think the honest answer is that “good” isn’t a single axis here — it splits into at least three separate questions people conflate.
Grey area 1: Good for whom, on what timescale? Faster capability growth is close to unambiguously good for near-term problems with clear feedback loops — drug discovery, materials science, climate modeling, disease diagnosis. Real, measurable progress. It’s the open-ended, high-autonomy, low-oversight applications where the calculus turns murky, because the benefit is speculative (“maybe it solves harder problems faster”) while the cost of losing the ability to verify what happened is asymmetric — hard to reverse.
Grey area 2: Is “winning the race” even the right frame? This is where I’d push on your question itself. Racing implies capability growth and oversight development are in zero-sum competition, but they’re not purely that — a lot of oversight research (interpretability, evaluations, red-teaming) directly depends on having capable models to study. So “no oversight” isn’t a stable outcome of capability winning; it’s more like oversight permanently lagging capability by an unknown and possibly widening margin. Whether that margin is fine (like software bugs, always present, rarely catastrophic) or unacceptable (like nuclear material controls, where a small gap is intolerable) is a genuine values disagreement, not a factual one.
Grey area 3: “Good for civilization” assumes we’d agree on what that means. Economic growth, individual autonomy, collective decision-making capacity, and existential safety don’t always pull the same direction. A system that’s economically transformative but concentrates decision-making in a handful of labs or governments is “good” on one axis and troubling on another — this is closer to a longstanding debate about technology and power than something specific to AI.
So: I don’t think there’s a clean answer where one side of the race “winning” is straightforwardly good and the other bad. The more defensible position — and roughly where most serious researchers land, whatever their risk estimate — is that the margin matters more than which side “wins”: a small, monitored gap between capability and oversight is very different from a wide, unmonitored one, even if both technically mean “capability got there first.”
Me: Okay, it’s good to know your ‘opinion’. Now can you briefly summarize the debate points without being too grey?
Claude: Fair — here’s the debate stripped to its actual poles, without the hedging:
The optimist/accelerationist position:
- Today’s models pose no existential threat — they can’t self-replicate, resist shutdown, or self-improve. The 10% figure is a subjective guess about a hypothetical future system, not a measurement of anything real.
- Oversight tools (interpretability, red-teaming, evals) improve because capability improves — they’re not in a race, they’re coupled. Slowing capability doesn’t obviously make anyone safer.
- Cognitive offloading is real but not new — calculators, GPS, and spellcheck all triggered the same fear, and society adapted rather than collapsed.
- Conclusion: push forward, the benefits (medicine, science, productivity) are concrete and now; the risks are speculative and later.
The safety/precautionary position:
- The people building these systems, including senior staff at the top labs, say openly they don’t have a solved plan for aligning a much more capable successor system. That’s a real admission, not speculation.
- Deskilling and automation complacency are well-documented in other high-stakes domains (aviation, medicine) — there’s no strong reason AI oversight would be immune.
- The catastrophic scenarios (recursive self-improvement, loss of meaningful human control) don’t need to be likely to warrant serious caution — low-probability, irreversible, civilization-scale risks get treated differently than ordinary risks everywhere else in policy and engineering.
- Conclusion: slow down or add binding checkpoints now, before the gap between capability and verification becomes unrecoverable, because you don’t get a do-over on this one.
The actual crux: not “is there risk” — both sides agree present models are safe and future ones are the concern. It’s whether oversight capacity naturally keeps pace with capability growth, or needs deliberate intervention (regulation, slower deployment, mandatory safety cases) to keep pace. That’s an empirical bet nobody can currently prove either way — which is exactly why it’s contested rather than settled.
Discover more from CuPod
Subscribe to get the latest posts sent to your email.