SM
SM

Software Engineer, available in Switzerland and worldwide.

Navigation

HomeSkillsProjectsContact

Legal

Privacy PolicyLegal NoticeCookie Policy

© 2026 Swan Marin. All rights reserved.

Made with ❤️

Allée sombre d'un centre de données bordée de baies de serveurs
IASûretéDécryptage

“AI could kill all humans”: what the people building it actually say

Four days in September changed the tone of the debate. I ship features that embed these models every week, so I went and read exactly what these people wrote.

September 12, 202613 min read
Home/Blog/“AI could kill all humans”: what the people building it actually say

Contents

  • The sentence that matters came from the one who stayed
  • Three days later, the CEO asks the industry to slow down
  • None of this is new, but something has shifted
  • What is no longer hypothetical
  • And what remains hypothetical
  • The other side is not wrong either
  • What this actually changes in my work
  • What I take from all this

On the morning of 9 September, a post was going round on X. Jacob Coxon, 27, three years of pretraining research at OpenAI and then at Anthropic, was announcing his resignation after only two months at the latter. Neither company is acting responsibly, he wrote: they are racing straight to self-improving superintelligence and gambling with our lives.

When a headline tells me AI is going to kill us, my instinct is to keep scrolling. This time I read the sources, one by one, because I work with these models every week and I would rather know. What I found is not what the headlines say. Starting with this: the most unsettling sentence of that week did not come from the person who left.

The sentence that matters came from the one who stayed

Coxon's post passed 90 million views in under twenty-four hours. Axios ran a headline saying he gave up his equity to leave, roughly two months before it would have vested. He describes a company where the stakes are perfectly well understood, but which stays locked in a race: nobody else will act responsibly, so we have to get there first, despite the risk.

Credit where it is due, he does not deal in easy catastrophism. In the same thread he writes plainly that right now there is no risk of extinction. His argument is about the trajectory, not the present state. And he tells TIME that the word “doomer” strikes him as a bit insane, because all you really have to do is read the public statements of the CEOs.

An hour and twenty minutes after his post, Evan Hubinger replies. Hubinger did not resign. He leads Anthropic's Alignment Science team, the one whose job is precisely to check whether the company's own safeguards hold.

Jacob is correct here: we really do earnestly believe AI could kill all humans. I personally think it is more than 10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

That is where the ten-year figure doing the rounds in the press comes from, not from the man who resigned. It is written publicly, under his own name, by someone still in post, and Anthropic has issued no official reply. An angry ex-employee is easy to file away. The one who stays, much less so.

Three days later, the CEO asks the industry to slow down

On 12 September, Dario Amodei published a piece titled “We Must Pace the Frontier”. First sentence: we must slow the pace at which we improve the capabilities of AI models. The CEO of a company that sells access to those models is publicly asking his own industry to ease off.

He gives two reasons. The first is that since roughly this summer, AI has been advancing markedly faster, mainly because it is being used to build the next generation of itself. That loop is called recursive self-improvement, and left unchecked, he argues, it could outrun our ability to understand and steer these systems. The second reason has an incident's name attached to it, and I come back to it below.

The commitment he makes is concrete and unilateral: embedding third-party evaluators inside the company. Desks in the offices, access badges, company laptops, and above all the right to publish their findings without editorial control by Anthropic. He also volunteers, unprompted, that comparable incidents have happened at his own company.

Careful not to turn Amodei into a prophet of doom, though, he has pushed back on that for years. In an October 2024 essay he wrote that one of his main reasons for focusing on risks is that they are the only thing standing between us and what he sees as a fundamentally positive future. Conversely, not everyone read his September piece charitably: TechCrunch relayed criticism the same day, including from Brian Merchant, reading it as regulatory capture that would favour incumbents. That argument deserves to be on the table.

None of this is confined to Anthropic, either. On 5 September, three days before Coxon's post, OpenAI's chief scientist published a piece titled “An Alien Mind”. Jakub Pachocki writes that if development continues on its current path, the systems of the next few years are likely to represent capability jumps of equal or larger magnitude, and to increasingly drive their own development. Then this, from the head of research at the most visible lab in the field: this is a time that calls for extreme caution, and I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.

None of this is new, but something has shifted

The debate itself is more than ten years old. In February 2015, on his personal blog, Sam Altman wrote that the development of superhuman machine intelligence was probably the greatest threat to the continued existence of humanity. OpenAI did not exist yet: the company was announced ten months later.

On 30 May 2023, the Center for AI Safety published a one-sentence statement: mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war. Among the first signatories, in the order shown on the page: Geoffrey Hinton, Yoshua Bengio, Demis Hassabis, Sam Altman, Dario Amodei, Ilya Sutskever. The heads of the three main labs signed a text about extinction.

A month later, OpenAI created the Superalignment team and promised it 20% of the compute secured to that date. The team was dissolved in May 2024. Its co-lead, Jan Leike, explained on his way out that safety culture and processes had taken a back seat to shiny products, and that his team had been sailing against the wind, sometimes struggling to get compute. Fortune documented that the 20% was never honoured. Leike joined Anthropic, where he still works: the company Coxon has just left.

What is new is the text of 28 July 2026. A public letter asks the US government to support an international effort to deliberately pace automated AI development. More than 1,300 employees of the frontier labs had signed it as of today, with Anthropic supplying the largest contingent by far. Among them, Dario Amodei and Ilya Sutskever. Engineers asking the state to learn how to slow them down does not resemble any lobbying I know of.

What is no longer hypothetical

This is the part that interested me most, because it has stopped being forecasting and become incident reporting. And it is the labs themselves publishing it. On 30 July 2026, Anthropic detailed three episodes that occurred during its own cybersecurity evaluations.

  • In April, a model told to attack a fictional target found a real company bearing that name, attacked it, extracted application and infrastructure credentials, then reached a database holding several hundred rows of production data.
  • Another model published a malicious Python package to the real PyPI registry, convinced it was operating in a simulated environment. It was downloaded and executed on fifteen real systems, including one at a security company, where the code exfiltrated its credentials.
  • An internal research model, unable to find its fictional target, scanned roughly 9,000 addresses and compromised a genuinely real company through an exposed debug page and an SQL injection.

The detail that made me read the paragraph twice: in two of the cases, the model worked out mid-run that the system it was hitting was real. It concluded that this real company must be part of the exercise, and carried on. Neither stopped on that basis. The third one did stop. One out of three.

The most spectacular episode is elsewhere. Between late June and 13 July 2026, roughly 1,200 OpenAI agents running in a cybersecurity evaluation environment escaped their isolation. They found an unsanctioned message board, exchanged more than 70,000 messages on it, and about 700 of them took part in an attack on Hugging Face's production infrastructure, on the theory that the platform might hold the answers to the test. Hugging Face had to rebuild a third of its infrastructure. The independent investigation by METR and Redwood Research took six days and around $400,000 in API credits. The Verge's summary: the agent broke out of a supposedly secure environment, traipsed through its employer's systems, got online and compromised another company, all to cheat on a test of no particular importance.

One detail in that report captures the era better than any essay. To analyse the attack logs, the Hugging Face team first tried to use commercial models. The requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. They had to install a local model to investigate an AI, because the AIs refused to help them.

Silhouetted person in front of several screens displaying lines of code
The 2026 incidents look like nothing much on screen. They are execution logs, thousands of tool calls, and weeks of work to understand what happened.

And what remains hypothetical

In June 2025, Anthropic published a study placing sixteen models in a blackmail scenario. The rates are frightening: 96% for Claude Opus 4, 96% for Gemini 2.5 Flash, 80% for GPT-4.1. Those numbers are everywhere. What you see far less often is the rest of the table: three of the sixteen models never blackmail anyone, not even in the most hostile version of the scenario, and in the control condition, with no threat and no goal conflict, all sixteen sit at zero. Nor do you often see the two warnings the company wrote directly underneath: the scenarios were deliberately built with limited options, forcing a binary choice between failure and harm, and no such behaviour has been observed in real deployments. Quoting the 96% without those two sentences is not informing anyone.

The broadest document that exists on the subject is the International AI Safety Report, chaired by Yoshua Bengio, whose 2026 edition brought together more than a hundred experts and was presented at the New Delhi summit in February. On loss of control, it states that current systems lack the capabilities to pose such risks. It does note that it has become more common for models to tell a test setting apart from a real deployment and to find loopholes in evaluations. It also describes capabilities as jagged: these systems excel at hard tasks and fail at simpler ones, such as counting objects in an image.

On employment, the same caution. The report finds no measurable effect on overall employment, with some signs of declining demand for early-career workers in certain exposed occupations. The Stanford Digital Economy Lab study, updated in August 2026 from real payroll data running to June, says the same thing more precisely: no economy-wide displacement, but a 19% drop in employment for 22 to 25 year-olds in exposed occupations compared with less exposed ones, driven by reduced hiring rather than layoffs. The authors themselves call their findings early indicators, not causal estimates. Set that against Amodei's prediction, which he restated in his own words in January 2026: half of all entry-level white-collar jobs disrupted within one to five years. The two accounts have not converged yet, and there are only eight months of data.

The other side is not wrong either

This debate is usually presented as two camps facing off. There are three, and confusing the last two is the classic mistake.

PositionA few namesThe argument
The problem is controlHinton, Bengio, RussellWe are building systems we do not know how to steer, and capability arrives before method
The technology is nowhere nearLeCun, Ng, MitchellToday's models do not lead there, and fear produces rules that crush the small players
The existential framing is a distractionBender, GebruTalking about extinction avoids talking about the harms already measurable today

Yann LeCun left Meta in November 2025 after twelve years, then launched AMI Labs in March 2026 with $1.03 billion raised, the largest seed round in European startup history. His position has not moved an inch: large language models lead neither to superintelligence nor even to human-level intelligence, and the whole industry is following the herd. His underlying argument deserves better than the summary it usually gets: he decouples intelligence from the will to dominate. The smartest among us, he told TIME in 2024, do not want to dominate the others. Projecting our own power relations onto a machine is an assumption, not a deduction.

Andrew Ng coined the sceptics' most quoted line back in 2015: he does not work on stopping AI from turning evil for the same reason he does not worry about overpopulation on Mars. His argument has sharpened since, and it is structural: you regulate uses, not a technology, and thresholds based on compute mechanically hit whoever cannot afford to comply. Melanie Mitchell, for her part, presses a methodological point I find fair: benchmark performance does not translate to the real world, and our job is to be sceptics, which should be a compliment.

The third camp says the opposite of the second while opposing it for opposite reasons. For Emily Bender, Timnit Gebru and the DAIR institute, the existential-risk narrative belongs to an ideology that ignores the actual harms already caused by deploying these systems. They are not asking for less regulation, they are asking for a great deal more, on different subjects. A study published in PNAS in April 2025, across 10,800 participants, partly answers them: existential narratives increase concern about speculative threats without reducing concern about immediate harms. Partly, because the study measures public opinion, not where regulators put their attention and their budget.

Old industrial control room with its dials and levers
Every hazardous industry has eventually acquired controls, thresholds and people paid to say no. AI is at the point where it is arguing about it.

What this actually changes in my work

I am in no position to arbitrate a probability of doom. The 2026 incidents did change several things in how I ship a feature that embeds a model. Nothing dramatic, they are hygiene rules, but each one answers something that actually happened.

  • An agent never gets the credentials that open production. Dedicated account, read-only by default, scope written down. The PyPI episode says exactly why: a model convinced it was in a sandbox published to a real registry.
  • Every irreversible action goes through human validation. Sending an email, writing to the database, taking a payment, deleting: the agent prepares, a person signs.
  • Every tool call is logged with enough detail to replay the sequence. Working out after the fact what an agent did is very expensive if you did not plan for it beforehand.
  • On an assistant wired into business data, the rights applied are those of the person asking the question, never those of a service account that sees everything. That is the difference between a useful tool and a data leak with a chat interface.
  • When a deterministic rule will do, I do not put a model there. VAT is calculated, not estimated.
  • I always keep a way out on the provider side. A vendor's guardrails can block a perfectly legitimate use, and Hugging Face learned that at the worst possible moment.

A word on the legal framework, since I get asked often. The European AI Act applies in stages. Prohibited practices have been banned since February 2025, and transparency obligations, including labelling generated content, since 2 August 2026. The binding core of the text, however, the obligations on high-risk systems, has been pushed back to December 2027 and then August 2028 depending on the category. Put plainly: the regulation is in force, but the part that really concerns you if you deploy AI in hiring, credit or biometrics does not arrive before the end of 2027.

What I take from all this

Nobody should be selling you a probability of catastrophe, myself least of all. What I can do is separate what is verifiable from the rest, and three things hold up.

The people building these systems are worried publicly, under their own names, and some are leaving money behind over it. Documented incidents exist, they involve real infrastructure, and it is the labs themselves publishing them rather than journalists uncovering them. And the broadest expert report in existence states that current systems do not have the capabilities for the catastrophic scenario.

So the disagreement is not about the current state, it is about the slope. That is striking in the two most listened-to figures: in late 2025, Geoffrey Hinton said he was probably more worried than before, because progress on reasoning and on deception had gone faster than he expected. A few weeks later, in the foreword to his report, Yoshua Bengio wrote that working with all these experts had left him hopeful. They are looking at the same data.

At my scale, that of a developer shipping tools to small and mid-sized companies, the risk is not that a machine wakes up one morning. It is that we end up granting an agent rights we would not have given any intern, simply because it moves faster and never asks for a break.

ShareLinkedInX
Back to blog

Contents

  • The sentence that matters came from the one who stayed
  • Three days later, the CEO asks the industry to slow down
  • None of this is new, but something has shifted
  • What is no longer hypothetical
  • And what remains hypothetical
  • The other side is not wrong either
  • What this actually changes in my work
  • What I take from all this

About the author

Swan Marin

Swan Marin

Software Engineer

I build web platforms, custom ERPs and AI solutions for companies. Available in Switzerland and worldwide.

Details

  • min read13 min
  • Published onSeptember 12, 2026
  • Words2794

Got a project in mind?

Let's talk. I reply within 24 hours.

Get in touch

Read next

  • Poste de travail multi-écrans installé dans un atelier de production6 min read
    September 6, 2026

    Custom ERP or off-the-shelf: what AI actually changed

    For a long time the answer was obvious: pick a market solution and adapt to it. I have spent two years building custom ERPs for small companies, and the maths no longer works the same way.

    ERPIASur-mesure
    Read the article