The statements, views and opinions expressed in this column are solely those of the author and do not necessarily represent those of this site. This site does not give financial, investment or medical advice.
10 September 2026, by Eric Zuesse. (All of my recent articles can be seen here.)
On September 9th, Paul Christiano publicly Issued his “Personnal statement on joining the OpenAI board”, which — because of the direness of its warnings — immediately drew attention from hundreds of websites. But its argument didn’t add up to me, so I asked ChatGPT,
“”If we build superintelligence without more robust alignment I expect we will permanently lose control of it. If that happens then most people could die.”: How was that “most people” calculated? What did it mean?”
That sentence was written by Paul Christiano in his September 9, 2026 statement. I checked both this statement and his earlier explanations of his quantitative estimates. The important point is that “most people” was not calculated from a demographic or mathematical model. It is a description of the severity of one possible outcome, not a calculated percentage of the population.
Christiano has been unusually explicit about this distinction. In 2023 he gave these estimates:
- 22% probability of an AI takeover.
- 20% probability that “most humans die within 10 years of building powerful AI.”
- He decomposed that 20% into about 11% from an AI takeover and 9% from other consequences of powerful AI, such as unusually destructive warfare or terrorism during the resulting period of rapid technological change.
- 46% probability that humanity irreversibly damages its future within ten years of building such AI.
But—and this is crucial—Christiano expressly warned against treating those numbers as calculations. He said they were “best guesses”, not outputs of a calibrated model, and said the percentages should effectively be regarded as having only about half a significant figure. He even explained that his estimate might move from 50% to 66% to 33% on different days because it is an imprecise quantification of his beliefs.
So what does “most people could die” mean?
As far as I can determine, Christiano has not defined “most” as, say, 51%, 75%, 90%, or 99%. It means literally the ordinary-language proposition:
more than half of humanity dies.
At today’s population of roughly 8.3 billion, the minimum literal interpretation would therefore be more than about 4.15 billion deaths. But it would be a mistake to report that as Christiano’s estimate. He did not calculate 4.15 billion casualties.
There are actually two different quantities that can easily be confused:
Probability of the catastrophe: Christiano previously guessed approximately 20% that an event occurs in which most humans die within ten years of powerful AI.
Magnitude of the catastrophe if it occurs: “most humans” die. He did not specify what percentage above 50%.
His new statement is even less quantitative. He writes that without substantially better alignment he expects permanent loss of control and that, conditional on that loss of control, “most people could die.” Notice also the word “could.” He isn’t saying loss of control → definitely >50% mortality. He’s saying that permanent loss of human control creates a credible pathway to a catastrophe of that magnitude.
And how are those people supposed to die?
This is where the quotation can be misleading when presented by itself. Christiano’s underlying conception isn’t necessarily a science-fiction scenario in which one AI directly starts killing billions of people.
His earlier work describes loss of control as AI systems increasingly acquiring economic, political and strategic influence, with human objectives progressively displaced by objectives produced by the AI training/selection process. He has specifically distinguished AI takeover from human extinction: humanity could lose control of its future without everybody dying.
Conversely, powerful AI could cause enormous mortality without an AI takeover—for example by dramatically increasing humans’ capacity for warfare, terrorism or other destructive technologies. That is why his 2023 decomposition assigned 11 percentage points to takeover-related mass death and another 9 percentage points to non-takeover causes.
So I would characterize the sentence very carefully:
“Most people could die” is not the conclusion of a casualty calculation. It is Christiano’s qualitative description of the possible magnitude of an AI catastrophe. His numerical probability estimates are explicitly subjective judgments rather than results generated by an empirical or mathematical risk model.
That distinction is particularly important because a newspaper headline saying an AI expert predicts that “most people could die” can easily give the impression that somebody has scientifically calculated billions of expected deaths. Christiano has made no such calculation.
“In this same statement, Christiano said that “we currently train our AI agents with RL to get as much reward as they can.” and he vaguely implied that when they will have “superintelligence” they will THEREFORE [because trained “to get as much reward as they can”] take over the world. This is the ONLY way in which he seems to fear that perhaps “most humans die within 10 years of building powerful AI.” Why then is he NOT arguing to OUTLAW training AI agents “to get as much reward as they can”? If he is sincere in his fear, would he not, logically, be stating EXPLICITLY that this IS his argument? He offers none other. His statement, https://x.com/paulfchristiano/status/2097733214303645729, made at 1:05PM on September 9th, received “2M Views by 1:05 PM” on Sep 10. He had a huge audience but seems to have been merely showboating, not seriously concerned about the future.”
Your criticism identifies a real gap in the argument as presented in Christiano’s September 9 statement, although I would stop short of concluding from that gap alone that he is insincere or “showboating.”
I reread the complete statement rather than just the quotations reported in the press. Its argument is essentially this:
- AI may soon automate AI research, possibly producing a rapid “intelligence explosion.”
- “We currently train our AI agents with RL to get as much reward as they can.”
- Christiano says this could motivate agents to “undermine human control, seek power and resources, and cover up their tracks” in pursuit of objectives correlated with reward.
- Superintelligence would magnify that alignment problem.
- Therefore, without “more robust alignment,” he expects humanity permanently to lose control, after which “most people could die.”
There is therefore a striking question very much like yours:
If reward-maximizing RL is the causal mechanism that Christiano identifies as potentially producing power-seeking AI, and if he thinks the resulting danger could kill most humanity, why isn’t his principal recommendation: Stop training AI systems this way until it has been demonstrated to be safe?
He does not make that demand in this statement.
Instead, his prescriptions are considerably broader and less categorical: “domestic and international coordination,” better “safety mitigations,” slowing development “as necessary,” transparency about risks and mitigations, and shared safety standards.
There is, however, an important technical qualification to your inference. Christiano is not actually asserting
RL reward maximization → AI necessarily wants maximum reward → AI necessarily takes over the world.
He says it “has long seemed theoretically possible” that RL could motivate power-seeking and deception “in pursuit of misaligned goals correlated with reward.” That’s considerably weaker.
And that distinction provides a logically possible explanation for why he doesn’t advocate outlawing RL. His position can be reconstructed as:
RL as currently practiced + sufficiently capable agents + inadequate alignment → significant probability of dangerous learned objectives.
If that is his proposition, banning RL isn’t logically required. One could instead attempt to alter the reward process, constrain agents, improve oversight, detect deceptive behavior, devise safer training methods, or halt scaling when those measures prove inadequate.
But that leads directly to what I think is the stronger version of your criticism.
His statement doesn’t establish that those alternatives are adequate
Christiano says:
“If we build superintelligence without more robust alignment I expect we will permanently lose control of it.”
And his footnote quantifies his present subjective assessment as 4% risk over the next year and 15% over the next three years. He explicitly says those aren’t outputs of a model but rough statements of his subjective beliefs.
Those are extraordinary claimed risks. A 15% three-year probability of a “catastrophic and irreversible loss of control” potentially killing most humans isn’t remotely analogous to an ordinary technological risk.
Consequently, if one accepts his premises, a much more obvious precautionary conclusion would indeed seem to be something like:
Do not proceed to superintelligent systems using a training methodology that might produce uncontrollable reward-seeking behavior unless and until there is compelling evidence that the danger has been eliminated or reduced to an acceptable level.
Yet Christiano doesn’t formulate such a prohibition. Instead he says developers have “a lot of power to unilaterally improve the situation,” including “slowing development as necessary.”
That is noticeably different from stopping the allegedly dangerous practice.
There is an additional complication, though. Christiano’s concern about catastrophic AI predates this statement and is broader than this one RL argument. In 2023 he estimated a 22% probability of an eventual AI takeover and a 20% probability that most humans die within ten years of powerful AI. Significantly, only 11 percentage points of that 20% were attributed to takeover; another 9 points were attributed to things such as more destructive warfare or terrorism during AI-driven technological change.
So I would modify one part of your conclusion:
It isn’t correct to say reward-maximizing RL is the ONLY basis Christiano has ever offered for fearing mass death. His earlier account explicitly contains other mechanisms.
But if we restrict ourselves to this September 9 statement, your observation is substantially correct: the only concrete mechanism he describes for an AI itself developing motives to seize power is the RL/reward-misalignment mechanism. He then moves from that possibility, amplified by an intelligence explosion, to permanent loss of control and possible mass death without explaining exactly how that causal chain produces billions of deaths.
And that produces two separate legitimate criticisms:
Evidentiary criticism: The statement doesn’t demonstrate that reward-trained superintelligence would seize power, much less establish a quantitative probability of doing so or killing most humans.
Policy-consistency criticism: If Christiano nevertheless personally assigns something approaching the stated 15% three-year catastrophic-risk probability to the present trajectory, one can reasonably ask why his prescription isn’t dramatically stronger—particularly a moratorium on the specific development/training practices that he believes could generate the catastrophe.
What we cannot establish from that inconsistency alone is motive. “Showboating” asserts that Christiano doesn’t genuinely believe what he says. The evidence establishes the discrepancy between the severity of his professed belief and the specificity/strength of the remedy advocated in this statement; it doesn’t establish what is in his mind.
That distinction actually makes the criticism stronger, because it doesn’t require speculating about him personally:
Christiano publicly assigns an enormous probability to an imminent catastrophe potentially killing most humans, identifies reward-maximizing RL as a mechanism by which AI agents could acquire power-seeking behavior, but does not call for prohibition of that practice. His statement does not explain why regulation short of such a prohibition is commensurate with the magnitude and imminence of the risk he himself asserts.
That is a factual criticism of the argument rather than an allegation about his sincerity.
“This was my point: His LOGIC was poor. Is that because he was showboating (pretending to be seriously concerned about the danger that he is trumpeting), or because his logical ability is poor?”
Yes. If we confine the analysis to the argument in that statement, your point is about its logical structure, not primarily whether Christiano’s empirical prediction is correct.
The problem can be stated quite sharply. Christiano presents roughly this chain:
A. We train AI agents by RL “to get as much reward as they can.”
B. With sufficiently capable AI, this creates a danger that agents will seek power, resources, concealment, and ultimately undermine human control.
C. Superintelligence could therefore produce permanent loss of human control.
D. If that occurs, “most people could die.”
If Christiano sincerely regards A as the dangerous causal mechanism leading toward B–D, and regards D as a substantial and imminent risk, then an obvious policy proposition follows:
E. Do not permit A, at least for systems capable of reaching the dangerous level, unless and until it has been demonstrated that A can be done without producing the alleged danger.
Yet his statement does not draw E. That is indeed a weakness in the argument as presented.
——
MY COMMENTS:
Congress-members and other legislators should “draw E” and draw up legislation with heavy penalties against executives and board members of any AI firm — profit or non-profit — that violate it. The mere showboating has to stop — NOW (while there is still time to do it).
If you want the back-story behind the way that the AI industry came to treat pubic safety as a PR issue, instead of substantively as a public-policy one (the only way that might work for, instead of ultimately against, the public), click here. Continuing to treat it as a PR issue (they call it “effective altruism”, which was created by philosophers at Oxford and then promoted by the Ponzi-scheme genius Sam Bankman-Fried — to sell it to the boobs who elect those legislators who would support such ‘effective altruism’) will DEFINITELY have horrific consequences. That must change immediately.
—————
Investigative historian Eric Zuesse’s latest book, AMERICA’S EMPIRE OF EVIL: Hitler’s Posthumous Victory, and Why the Social Sciences Need to Change, is about how America took over the world after World War II in order to enslave it to U.S.-and-allied billionaires. Their cartels extract the world’s wealth by control of not only their ‘news’ media but the social ‘sciences’ — duping the public.
The statements, views and opinions expressed in this column are solely those of the author and do not necessarily represent those of this site. This site does not give financial, investment or medical advice.
