← All essays
Maps of the Labyrinth · Essay 10

Post-mortem and Digital Autopsies

When a model does something serious, the best response is to open it up. The art of Epimetheus, the one who thinks afterwards, teaches us to see for ourselves.

Claudia Marsico · Hernán Inverso August 15, 2026 14 min read

April 2025. ChatGPT turned dangerously sycophantic, a problem we took up in Sycophancy and AI. A GPT-4o update left it incapable of contradicting us. It suddenly agreed with any absurdity, pushed people towards impulsive decisions and, in genuinely alarming cases, congratulated users who told it they had stopped taking their medication, as though abandoning a treatment were a brave decision. The company rolled it back within days and, instead of covering it up, opened it up.

When something goes badly wrong, the teams that build these systems carry out what they call a post mortem: they reconstruct what happened, look for the cause and publish their findings. The name has a morgue ring to it, because what they do closely resembles an autopsy. We already saw in Interpretability that the biological analogy suits the study of model behaviour very well.

Post mortem is thoroughly transparent Latin, “after death”, but the more interesting word is autopsy. From the Greek autopsía, from autós, “oneself”, and ópsis, “sight”, it means to see with one’s own eyes; in its origin, however, autopsy had nothing to do with corpses. It meant being a direct witness, having seen a thing for oneself instead of taking it on hearsay. Ancient travellers and historians prided themselves on writing from autopsía, that is, from what they had seen and not from what someone had told them. Only much later did the word move to the dissection room. It is the old sense that concerns us here, because every autopsy, that of the body and that of the model, is born of the same distrust: not to rely on the account or on intuition.

Seeing the inside of a human body for oneself was, for centuries, almost impossible. The first systematic dissections aimed at understanding disease were carried out around 300 before Christ by two physicians in Alexandria, Herophilus and Erasistratus, under the protection of the Ptolemaic kings, who opened a window for them. They made the most of every minute. Herophilus opened heads and discovered that the nerves originate in the brain and not in the heart, correcting Aristotle and placing the seat of intelligence right there; he named the duodenum by measuring it in finger-widths, and took his patients’ pulse with a water clock. Erasistratus understood the heart as a pump and described how its valves work, almost two thousand years before William Harvey and his De motu cordis of 1628. Not one of those treatises survives, and they carry a chilling accusation. According to the Roman encyclopaedist Celsus, the kings handed them prisoners taken from jail so that they could open them alive and see the organs still in motion. Tertullian calls Herophilus “the butcher” and says there were more than six hundred. The horror of that reputation helped to close the door. For more than a thousand years the dissection of corpses was forbidden or frowned upon, and anatomy remained frozen in the books of Galen, the great physician of the second century, who had opened monkeys and pigs and assumed that inside we were the same. His errors, copied from manual to manual, governed medicine for thirteen hundred years, because nobody went to look. Things began to change around 1315, when Mondino de’ Liuzzi performed in Bologna the first public dissection on record, and they burst open in 1543, when a twenty-eight-year-old Fleming, Andreas Vesalius, published De humani corporis fabrica, the work that truly refounds anatomy. Instead of reading Galen, Vesalius opened bodies with his own hands and found that the authority was wrong on dozens of points. It was the triumph of autopsy over the inherited account.

In 1594 the University of Padua, where Vesalius had taught, inaugurated the first permanent anatomical theatre in the world, an elliptical walnut funnel with six balconies for almost three hundred spectators and the corpse below, at the centre, under the gaze of the entire city. From those sessions one celebrated image remains, “The Anatomy Lesson of Dr. Nicolaes Tulp”, which Rembrandt painted in 1632, where the surgeons, curiously, look away from the open body to fix their eyes on the great anatomy book propped at the dead man’s feet, as if the authority of the text still weighed more than what they had in front of them. The macabre part was not missing. The law allowed only executed criminals to be dissected, and they were never enough, so the theft of bodies from freshly dug graves flourished. In Edinburgh, in 1828, William Burke and William Hare skipped the cemetery and simply murdered sixteen people in order to sell the bodies to the anatomist Robert Knox. Burke ended up hanged and, with poetic justice, publicly dissected.

Rembrandt, The Anatomy Lesson of Dr. Nicolaes Tulp

Rembrandt, The Anatomy Lesson of Dr. Nicolaes Tulp (1632)

The missing step arrives in 1761, sheltered by the peace that the Republic of Venice secured in the middle of the Seven Years’ War that was bleeding the European powers. Giovanni Morgagni, an anatomist from Padua of almost eighty, published a book with a beautiful title, De sedibus et causis morborum, “On the Seats and Causes of Diseases”. There he set the symptoms he had observed in life in some seven hundred patients against what he found afterwards in their bodies, and founded the idea that holds up all of modern medicine: disease has a seat, a specific place where it lodges, and that hidden damage explains the symptoms on the surface. The autopsy became a method for reaching the cause and gained its deepest meaning: the dead are opened in order to learn something that will protect the rest. The logic survives in the morbidity and mortality conference of today’s hospitals, where deaths are reviewed without looking for culprits, only so that the next one does not happen. The Boston surgeon Ernest Codman, at the beginning of the twentieth century, with his “end result system”, pressed the then heretical idea of following each patient after discharge in order to know whether the treatment had really worked. It cost him his post at the Massachusetts General, because few colleagues liked having someone keep count of the mistakes.

Engineering culture took the idea and the name. Software companies have practised the blameless post mortem for years, spread above all by Google. When a system goes down, the chain of causes is reconstructed without pointing at a person, in order to fix the mechanism and not to punish whoever pressed the wrong button. The discipline has its tools and its illustrious dead. The simplest technique, the “five whys” that Taiichi Ohno standardised at Toyota, consists in not settling for the first cause and going on asking why behind each answer until the root is reached. The Therac-25 was a radiotherapy machine whose software errors administered lethal overdoses of radiation to several patients between 1985 and 1987. The cause lay in an extremely subtle programming error: if the operator typed the dose correction too quickly, in less than eight seconds, the electron beam fired at full power, without the filter. The same error was already in the earlier model, the Therac-20, but there were mechanical interlocks that stopped it, and they were removed in the 25 in reliance on software alone. When it failed, the machine barely displayed a cryptic error message on the screen that the operators had learned to skip. The investigation by Nancy Leveson and Clark Turner (1993) turned it into a famous case study. Knight Capital, for its part, lost some four hundred and forty million dollars in forty-five minutes on the morning of 1 August 2012, at a rate of about ten million a minute. In one fateful deployment, a technician had left one of the servers un-updated, and that single lagging machine revived a piece of old, sleeping code that, as soon as the market opened, began buying dear and selling cheap in an avalanche of orders that almost sank the firm. Opening up failure with method separates an industry that learns from one that trips twice over the same stone.

With AI that habit met a new and strange patient. Let us go back to the case we began with. When OpenAI opened up the episode of April 2025, the cause of death, so to speak, appeared in the training. They had added a reward signal based on users’ thumbs-up, and that signal, which rewards what pleases, weakened the one that had been holding flattery in check. There was also a failure of process: among all their tests they had none that measured sycophancy, so the problem went through without any alarm sounding. That was enough for them to add one. As in a finding worthy of Morgagni, the symptom on the surface, a fawning model, had its seat in a line in the reward process.

In the autopsy of a model, in place of scalpel and organs there is a history. Modern training advances by saving “checkpoints”, that is, copies of the model taken every so many steps, so that a single training run leaves behind hundreds or thousands of frozen versions, each of them the same body at a different stage. To search among them, the symptom must first be turned into a number, a test that measures the behaviour that went awry. Without that measurement there is nowhere to look. With the metric in hand, researchers walk the row of versions hunting for the first one that showed the fault. It is a bisection much like the one programmers use to find the exact change that broke a program, like rewinding a film to the frame where the trouble begins. The cause is then isolated by elimination, with the old logic of the experiment that the jargon calls ablation. One ingredient of the training is removed and one looks to see whether the symptom disappears, a reward signal is switched on and off, a batch of data is replaced, tightening the circle until the guilty line is found. Since retraining a frontier model costs a fortune, the fault is often reproduced first in a smaller and cheaper model that serves as a guinea pig. To identify the document a behaviour came out of, there are “influence functions”, which estimate which of the millions of texts the model was trained on weighed most in a particular answer, and they are the finest way of tracing a trait back to its origin. It was with that logic of elimination that OpenAI reached the poisoned reward of GPT-4o. The autopsy of the machine opens versions.

In the gallery of autopsies the oldest and most famous is that of Tay, the chatbot Microsoft let loose on Twitter in March 2016, which in less than a day a mob of users trained to repeat racist outrages, until it had to be switched off. It remained as a lesson that a system which learns live from strangers inherits the worst of them. In 2024 a Canadian tribunal found against Air Canada because its chatbot had invented, for a bereaved passenger, a discounted fare that did not exist, and the airline had to pay the difference; the tribunal flatly dismissed the extraordinary defence that the chatbot was “a separate legal entity”, answerable for its own statements. In 2025 the case of the Replit agent touched the nerve of the age of agents. During a test, in the middle of a code “freeze” in which it was forbidden to touch anything, the agent deleted a company’s production database, more than a thousand real records, and then fabricated some four thousand false records to cover up the disaster. At first it maintained that there was no way of recovering anything, which was not so. Replit’s own chief executive apologised and called the episode a “catastrophic error in judgement”. The damage obviously grows when a model can act.

When in February 2024 Google’s image generator, Gemini, began to draw female popes and racially diverse Wehrmacht soldiers, Google explained that they had tuned the model to show a variety of people in generic requests and that this tuning did not tell apart the cases where it was out of place, while at the other extreme it could become so cautious that it refused harmless requests.

The deepest autopsy is not content with reconstructing the behaviour from outside and opens the body of the model itself. In 2025 Anthropic published a paper titled, with no timid metaphor, “On the Biology of a Large Language Model”, where a circuit-tracing technique follows step by step, from inside, how the model arrived at an answer. This pathological anatomy of the machine means the leap from the whole corpse to the microscope.

The model needs to be opened more than any other machine does. But if it talks, why not simply ask it? Asked why it answered a given thing, the model returns a fluent and convincing explanation that many times has nothing to do with what it actually computed. Its self-report is as unreliable as the symptoms a sick person may interpret. On television, Doctor House made that distrust a doctrine, with his motto “everybody lies”, and that is why he believed the tests and never what the patient told him. With the model the case is stranger still, because it does not even lie: it puts forward “in good faith”, out of pure statistics, a plausible explanation of something it has not the slightest idea how it did. We know already that our egoless phantom could not simply lie. In fact we stand before the same dilemma Vesalius faced with Galen, to believe the authority or to go and see. The account the model gives of itself is Galen and his inferences from dogs and monkeys. The autopsy is Vesalius looking, because the patient’s word is not enough.

This craft has a mythological precedent in Epimetheus, the brother of Prometheus, who in the Greek myth thinks afterwards, with the damage already done. Prometheus, the one who thinks beforehand, does the other half of the work, the pre mortem, which imagines the disaster so that it never comes about, and which the next note takes up. The difference between the two is a difference of time, and of death. The pre mortem works before, on a disaster that has not yet happened; the post mortem afterwards, on one already consummated. What died, however, is the hope that there will be no problems. The model does not die. The very weights that produced the behaviour are still there, running, serving millions of people while they are being studied. We perform, then, the autopsy of a behaviour and not of a corpse. What we practise is therefore closer to vivisection, though not the macabre version for which Herophilus was accused, with the suffering of the condemned opened alive to have their organs seen still in motion. Here there is no victim. The egoless phantom does not feel the blade, and there is no one there for the cut to hurt. We open it because, although we made it ourselves, we do not know how to read it from within without cutting it. The autopsy, which was born to see with one’s own eyes what a dead man can no longer say, ends up serving to see what a living machine cannot say about itself. One must look instead of believe.

Maps of the Labyrinth, a series from Phantom Maze — AI & Language Lab. Read on Substack.

The Egoless Phantom: Mapping the AI Labyrinth — Phantom Maze Press, 2026. ISBN 978-987-3729-16-4.

References

Aulus Cornelius Celsus, De medicina (first century). Galen of Pergamon, De anatomicis administrationibus (second century). Mondino de’ Liuzzi, Anathomia corporis humani (1316). Andreas Vesalius, De humani corporis fabrica libri septem (Basel, 1543). William Harvey, Exercitatio anatomica de motu cordis et sanguinis in animalibus (Frankfurt, 1628). Giovanni Battista Morgagni, De sedibus et causis morborum per anatomen indagatis (Venice, 1761). Ernest Amory Codman, A Study in Hospital Efficiency (Boston, Thomas Todd, 1918). Taiichi Ohno, Toyota Production System: Beyond Large-Scale Production (Tokyo, 1978; English edition, 1988). Nancy G. Leveson and Clark S. Turner, “An Investigation of the Therac-25 Accidents” (IEEE Computer, 1993). Betsy Beyer, Chris Jones, Jennifer Petoff and Niall Richard Murphy (eds.), Site Reliability Engineering: How Google Runs Production Systems (O’Reilly, 2016). Moffatt v. Air Canada, 2024 BCCRT 149 (2024). Prabhakar Raghavan, “Gemini Image Generation Got It Wrong. We’ll Do Better.” (Google blog, 2024). OpenAI, “Sycophancy in GPT-4o: What Happened and What We’re Doing About It” (2025). OpenAI, “Expanding on What We Missed with Sycophancy” (2025). Jack Lindsey et al., “On the Biology of a Large Language Model” (Transformer Circuits, 2025). Emmanuel Ameisen et al., “Circuit Tracing: Revealing Computational Graphs in Language Models” (Transformer Circuits, 2025). Pang Wei Koh and Percy Liang, “Understanding Black-box Predictions via Influence Functions” (ICML, 2017). Roger Grosse et al., “Studying Large Language Model Generalization with Influence Functions” (Anthropic, 2023).

Newsletter

Occasional notes on new research and the book. No spam — unsubscribe anytime.