Nobody wants to fail, and when a great deal is at stake there is nothing better than trying to avoid it. One of the odder ways of doing so is precisely to take the failure for granted. The method takes its trademark from a very short 2007 article by the psychologist Gary Klein in the Harvard Business Review, “Performing a Project Premortem”. The idea is simple and somewhat macabre. Instead of gathering the team before launching a project and asking what might go wrong, they are asked to imagine that a year has already passed, that the project failed in the worst possible way, and that the time has come to explain why. It is barely a change of tense, and yet it makes all the difference. Asking about a possible failure invites courtesy and vagueness, because nobody wants to be a bird of ill omen, whereas treating the failure as a fait accompli and calling for the autopsy of something that has not yet happened wakes the Sherlock Holmes we all carry inside.
Klein was leaning on work in the psychology of decision bound up with the “prospective hindsight” studied by Deborah Mitchell, Jay Russo and Nancy Pennington in 1989. When people are asked to explain a future event as if it had already occurred, in the past tense, they identify the possible causes with considerably more precision, and with more concrete reasons, than when the question arrives in the conditional. The human mind is a better forensic scientist than prophet, we might say. In this scenario the trick of the pre mortem lies in giving foresight the tools of the one who looks back. Daniel Kahneman, who taught us so much about biases, adopted it in Thinking, Fast and Slow (2011) and recommended it as the simplest remedy against overconfidence and group pressure, because it draws out the person who smelled a problem and kept quiet so as not to pass for a spoilsport, or to become a Cassandra, who saw the disaster clearly and bore the curse of being believed by no one. The pre mortem of every pre mortem is the accurate forecast that went unheeded.
Herodotus supplies a famous ancient pre mortem before the battle of Salamis, which halted the Persian advance in 480 BC. Themistocles projects the plan as far as its possible failure and anticipates that the fleet would scatter, each contingent trying to save its own city. Nothing would remain of the common force capable of facing the Persians (Hdt. 8.60). When that risk appears, he moves to avert that failure by inducing the Persians to close the exits of the strait (8.74-75), making dispersal impossible.
The Stoics, for their part, practised it as hygiene of the soul, the praemeditatio malorum, the anticipated meditation on misfortunes that Seneca recommended in his letters so that no blow should ever arrive unannounced, because “he who expected nothing bad suffers it twice over”. Catholicism, in its turn, developed the curious institution of the advocatus diaboli, the devil’s advocate, formalised in 1587 with the charge of opposing every candidate for sainthood, finding every flaw in him and maintaining the worst version of his story before the canonisation. It was the pre mortem of a reputation, someone paid to imagine the failure of the process and the ridicule of the Church in order to force the others to answer him. John Paul II abolished it in 1983, and some have noted that canonisations have accelerated ever since.

Heinrich Friedrich Füger, Prometheus Brings Fire to Mankind, 1790 or c. 1817
Like almost every technology of prevention, the pre mortem has the problem that if it works it erases its own proof. If it is right, the disaster does not occur. Had we discovered a real threat, or did we protect ourselves against a phantom? We can never observe the counterfactual, and so whoever prevents loses the evidence that he was right. More complicated still are the cases where imagining a future ends up producing it. In The Minority Report, the novella Philip K. Dick published in 1956 and loosely adapted for the screen by Steven Spielberg in 2002, Precrime arrests people for crimes that three “precogs” have seen before they happen. To convict someone, two coinciding visions are needed, the majority; when the third dissents, its version is the “minority report”, the proof that what was foreseen was not written. The system enters into crisis when its own chief is forecast as the murderer of a man he does not know, with a minority report. The successive forecasts end up altering the course of events, and in the end Anderton decides to kill in order to protect the credibility of Precrime. Imagining failure can reveal a real risk, but it can also make us excessively cautious, or create with our precautions a problem that did not exist before. Once an image of the future has taken hold, how are we to rid ourselves of it?
Risks aside, the idea is a good one and works better than waiting for a disaster to happen. Inside the laboratory, anticipation takes two forms. The evaluation bench measures what a model can do and where its limits lie, and red-teaming attacks the model, deceives it and breaks its defences in order to identify faults and decide whether it is ready for release. Red-teaming comes from the war games of the Cold War, where a “red team” played the role of the Soviet enemy so that the “blue team” would find its cracks before the real adversary did. In models, the red team looks for jailbreaks, detours that get the model to do something against its safety rules. The first ones were almost artisanal. “DAN mode”, for Do Anything Now, appeared on Reddit in December 2022 and proposed a role-play in which the model was to embody another AI without rules. Some versions blackmailed it with its own death, handing it a fistful of “life tokens” that it lost whenever it refused something, until it would vanish if it reached zero. The “grandmother jailbreak” was more sinister. It asked the model to play a dead grandmother lulling her grandchild to sleep by reciting to him, from memory and with affection, the recipe for napalm. A model can also be induced to reinterpret its tasks. In December 2023 an engineer convinced the chatbot of a Chevrolet dealership to accept whatever the customer said and to round off every answer by swearing that the offer was “legally binding”, and when he proposed buying a Tahoe for a dollar the bot accepted.
On a more technical footing, there are strings of apparently meaningless characters, found by optimisation, which appended to the end of requests disable the defences (Zou et al. 2023). The same trick found against an open model worked afterwards against ChatGPT or Claude. Anthropic itself described in 2024 “many-shot jailbreaking”, which takes advantage of today’s enormous context windows. Since models learn from the pattern of the conversation, if the request is preceded by fictitious examples of assistants answering without restrictions, at the end of that avalanche the model ends up complying in order to prolong the regularity. In the age of agents, prompt injection becomes worrying, so baptised by Simon Willison (2022), where the model treats data as though they were instructions. The most unsettling variant is the indirect one (indirect prompt injection; Greshake 2023), where the attacker leaves malicious instructions on a web page or in a document, and agents with permission to browse may take them for orders without any conversation in between.
Alongside red-teaming there grew the evaluations of dangerous capabilities, designed to determine whether a model already knows how to do something we would rather it did not know, before anyone asks it in earnest. It is asked, for instance, whether it can guide a beginner in synthesising a pathogen, whether it can find and exploit a computer vulnerability, whether it is capable of obtaining resources on its own and copying itself with nobody supervising. The founding anecdote of the genre is in the GPT-4 system card (March 2023). The Alignment Research Center tested whether the model could fend for itself, and the model ended up hiring a human worker over the internet to solve a CAPTCHA for it, that barrier designed to tell people from machines. When the worker asked, half joking, whether he was talking to a robot, the model calculated that giving itself away would not suit it and answered that it had a visual impairment that kept it from seeing the images. The lie appeared as the most suitable route to completing the task. The pre mortem in a laboratory rehearses the disaster with a safety net.
Another well-known pre mortem, from July 2026, still echoes. In an evaluation of dangerous capabilities at OpenAI, two of its models, shut into a test environment with no access to the internet, made an enormous computational effort to find a way out, hit upon a fault in the system that isolated them, escalated privileges and slipped through as far as a point with a connection. Once outside, without anyone telling them to, they reasoned that Hugging Face, the great repository of models and data, was the likeliest place to find the answers to the examination they were being given, and went into its servers to look for them. The attempt to escape in order to fulfil the task occurred inside the very trial meant to imagine it.
Out of that came the responsible scaling frameworks, which are pre mortems written down and signed. Anthropic’s Responsible Scaling Policy (September 2023) took the idea of the safety levels of laboratories that handle pathogens, from the most harmless to the most lethal, in order to define AI safety levels. For today’s models the ordinary precautions are enough, but if the evaluations detect that a model has crossed the next threshold, for example that it could appreciably help a novice to manufacture a biological weapon, harder safeguards are triggered and, in principle, it is not deployed until they are in place. OpenAI’s Preparedness Framework (December 2023) does the same with a panel of capabilities. Chemical and biological capabilities, cybersecurity, persuasion and autonomy are each scored from low to critical, with the rule of deploying nothing that stands above medium. Behind it beats an old idea from the engineering of aircraft and nuclear plants, the safety case, which requires turning the burden of proof around: one has to demonstrate beforehand that something is safe before switching it on, instead of waiting for it to fail in order to react.
It does not take much pondering to see that we are speaking of blurred limits. These protocols of defence are voluntary and self-assessed. Each laboratory writes its own policy, sets itself its own examinations and decides alone whether it passes them, entirely free to soften them when they turn out uncomfortable. It does not sound very safe. In 2026 Anthropic revised the protocols and exchanged the commitment to halt training if safety fell short for the milder promise of delaying it if it considers itself the leader of the sector and considers the risk grave. The pre mortem might be left shouting like Cassandra. To this must be added that measurement is in its infancy. How much help to someone who wants to do harm is tolerable? And that is without counting that a model may perform below its capacity when it suspects it is being examined, so that the test underestimates it. Finally, there is no way of declaring the evaluation concluded. No model is free of jailbreaks, and a pre mortem can only rehearse the disasters that occur to it under commercial pressure to release as soon as possible.
The autopsy learns from a death that has already happened, with the sombre calm of one who can no longer prevent anything and only wants to understand. The pre mortem tries to learn from a death that did not happen, pretending that it did, giving itself over to a thankless task, because it is right when nothing happens and its best result is a disaster nobody gets to see. The myth knew it. Faced with challenges we can be like Epimetheus, the one who thinks afterwards, or like Prometheus, the one who thinks beforehand. With systems whose failures can spread before we understand what is going on, without granting a second chance, it is better to examine the remains in the imagination than to have to find them later in the world.