← All essays
Maps of the Labyrinth · Essay 13

Osiris in the Chat and the Body Without Eyes

Where does what an AI knows come from? A journey through the world of the corpus.

Claudia Marsico · Hernán Inverso September 5, 2026 13 min read

To train GPT-3, OpenAI gathered forty-five terabytes of Common Crawl, a gigantic snapshot of the web. After filtering and cleaning it, some five hundred and seventy gigabytes remained, equivalent to around four hundred and ten billion tokens. To that block were added WebText2, two large collections of books and the English Wikipedia, until some five hundred billion tokens had been assembled. During training, the model processed some three hundred billion tokens drawn from that mixture. Later corpora multiplied that scale by fifty, into a quantity of text impossible to imagine in human terms, more than could be read by someone who began today and stayed awake for several lifetimes in a row.

No human being, nor any team of human beings, however large it might be, has gone through the whole of the material on which a language model is educated. Nor have there ever been collections of words so scrutinised, counted, filtered and weighed, though neither can we say in the strict sense that anyone has read them.

In artificial intelligence, a corpus is a delimited collection of texts turned into training data, and to that end it was cleaned, segmented, mixed with other sources and presented to the model as linguistic experience so that it may learn regularities. Strictly speaking, what it contains matters, and also what was left out, how it was divided, how many times it was shown and with what weight each source entered, that is, how the corpus was formed, corpus being the name of that collection. From Latin it passed into the Romance languages with the sense of a living body, and into English as corpse, a dead one. We find it in “corporation”, which is a body made of persons, in the habeas corpus of the law, literally “that you may have the body”, the writ that demands that the body of the detainee be produced before the judge, in corpus delicti, the body of the crime, and also in the corpus mysticum of the theologians. In every case it has parts and forms a unity that can be touched, counted and sometimes dismembered.

The term has also been used for centuries to name the complete set of the works of an author, of a tradition or of a language, provided that a unity made up of parts is delimited. Thus we have the corpus of Aristotle or the corpus of Roman law that Justinian ordered to be compiled. When linguistics took a more empirical turn, corpus was the name given to a sample of real language prepared for analysis. One of the first great machine-readable corpora was assembled in 1961, at Brown University, and for that reason it is called the Brown Corpus. The linguists Henry Kučera and W. Nelson Francis designed it with the utmost attention to balance, setting in it a little over a million words of prose printed in the United States that year, distributed across five hundred samples of some two thousand words each. They balanced it carefully among fifteen genres that ran from reportage to the romantic novel and the religious essay. They chose each fragment by hand, cataloguing it and defining its place in the whole until they had formed a body that was small, balanced and still wholly inspectable by its builders.

From the million words selected with obsessive care to the fifteen trillion that are today hunted on the web there is an evident leap in size, but also a change of logic. The Brown Corpus was built as a small, balanced, human-readable sample; the corpora of language models are assembled out of massive crawls, automatic filters and statistical decisions about what to keep, repeat or discard. The word corpus stayed where it was, but it went from naming a curated collection to naming an industrial body, too large to be seen whole.

The Egyptian myth tells that Osiris reigned with justice until his brother Set, the Typhon of Greek mythology, murdered him out of envy, tricking him into entering a chest that he sealed and threw into the Nile. Isis, sister and wife of Osiris, recovered the body, but Set seized it again and this time dismembered it into fourteen pieces, which he scattered across the whole of Egypt so that she could not reassemble it. Plutarch tells that Isis went over the entire country gathering the pieces one by one and sewed them together to give the body back a breath of life. She found all of them but one, the phallus, which had vanished, eaten by a fish. In its place Isis fashioned a prosthesis to complete what was missing.

Strangely, this tale from the sands of Egypt also tells the story of a training corpus. Human writing lies dismembered and scattered across the whole web. A paragraph here, a comment there, half an encyclopaedia on one server, a repair manual on another, forums, recipes, court rulings, poems, insults, technical documentation, anything we have ever cared to write. The crawlers that feed a model perform the task of Isis: they go over that immense territory, gather the pieces wherever they have fallen and sew them into a single body, from which a voice will later come. As in the case of Osiris, a piece is always missing, because it was left outside the filter, because it was never digitised or because it was never written. Then the prosthesis appears: text manufactured to fill the gap. The body of the machine is a reassembled Osiris.

Relief from the temple of Seti I at Abydos: Egyptian figures carved in limestone, in a row

Relief from the temple of Seti I at Abydos, Egypt, thirteenth century BC. Photograph: Kyera Giannini / Institute for the Study of the Ancient World (ISAW)

What is that body made of? It is not the internet plain and simple, nor even the whole web, but a pile of different sources that are selected, cleaned and mixed. At the base there is usually Common Crawl, an archive that a non-profit foundation has been taking from the open web since 2008 and that accumulates petabytes of pages. On that base are added sources chosen for their value, such as Wikipedia, large collections of books, code from repositories such as GitHub, questions and answers from forums such as Reddit or Stack Exchange, scientific articles from arXiv or PubMed, legal texts, manuals and technical documentation. When those mixtures are packaged and published, they are given a proper name. The Pile, which the EleutherAI collective assembled in 2020, brings together eight hundred and twenty-five gigabytes in twenty-two different sources, from medical articles to film subtitles and legal texts. C4, prepared by Google in 2019 to train its T5 model, is a clean version of Common Crawl. Dolma, from the Allen Institute, comes to around three trillion tokens. FineWeb, published by Hugging Face in 2024, reaches fifteen trillion, drawn from ninety-six successive crawls of the web.

This matter is worked with the utmost care. The filters of C4 prune the poor lines, keeping only those that end in a punctuation mark, leave out those with fewer than three words and discard the documents that do not reach five sentences. They also remove pages with filler text such as “lorem ipsum”, with programming braces or with terms from the celebrated List of Dirty, Naughty, Obscene, and Otherwise Bad Words, Shutterstock’s repository for avoiding undesirable words, with files by language and the odd curiosity such as a list in Klingon. The C4 filter leaves out everything the language detector does not recognise as English with a confidence of at least ninety-nine per cent. Other curation systems add, on top of those filters, quality classifiers, which score documents according to how much they resemble the good examples and mark down or discard the rest. Beneath it all runs deduplication, which removes the repeated pieces, a natural plague of the web. It is no cosmetic detail, since in experiments on C4 deduplicating made the models reproduce memorised text far less and allowed them to learn in fewer steps. The well-sewn body of our Osiris works better than a bloated one. At this point people stopped believing that more data was always better and began to look at how they were organised.

In the assembly it is also decided how much each part weighs, which is why the sources do not enter training in the same proportion in which they appear in the corpus. They are rebalanced. In GPT-3, Common Crawl was by far the largest block, with more than eighty per cent of the tokens available. During training, however, it was sampled less than its size would indicate, while smaller and more reliable sources received more turns of reading. Wikipedia and the books were read several times over during training, well above what their size would justify. As in a diet, the choice is made not by availability but by place in the whole. The model may come across a good encyclopaedia article several times and barely leaf through other regions of the ocean of Common Crawl. These choices condition the tone and the gaps of the model.

For a few years the race concentrated on the size of the model, always seeking more parameters. In 2022, a team at DeepMind published the paper that came to be known by the name of its experimental model, Chinchilla, and showed that the industry had been doing the sums wrong. For a given budget of computation, the best thing was not to make the model as large as possible, but to balance size and quantity of text in a proportion of around twenty training tokens for every parameter. To prove it they trained Chinchilla, a model of seventy billion parameters fed with one trillion four hundred billion tokens, which beat Gopher, another DeepMind model four times larger but trained on far less text. From then on the corpus gained centrality, and models even began to be overtrained, giving enormous quantities of text to relatively small models so that they would perform better.

That appetite runs up against a limit that came to be spoken of as “the data wall”. The Epoch AI group estimated that the stock of public human text usable for training comes to around three hundred trillion tokens, and that, if the trend continues, models will finish consuming it between 2028 and 2032, or sooner if overtraining accelerates. Put another way: the industry could be approaching the limit of the stock of usable public human text. One of the ways out most explored is to manufacture what is missing with synthetic text from the models themselves in order to keep feeding them, like the prosthesis of Isis, but it is not so simple. Model collapse lies in wait. When one generation of models learns indiscriminately from texts produced by earlier models, it inherits their errors and gradually loses contact with the original distribution (Shumailov 2024). The rarities disappear, the infrequent cases, the lateral ways of saying things, and the system becomes ever more repetitive and ever more certain of its impoverished version of the world. Like a photocopy of a photocopy, each pass keeps what is most common and washes out the improbable, which is why synthetic text is admissible as a prosthesis but fails as a whole body, and the fish goes from eating the missing piece to swallowing the memory of the other pieces.

The model knows the world through this body without eyes, or ears, or childhood, which is at once everyone and no one. It holds a gigantic portion of written human conversation, but not all of it, because there are languages and regions that produce less indexable text or were left out of the selection. And even of what it read, it had to be made to forget, by dint of filtering and deduplication, an enormous part of what it had found.

Strictly speaking, a corpus is a collection of materials, but it presupposes an implicit theory of what language is and how it can be learned. The Brown Corpus assumed that a language could be represented by a manageable and balanced sample. For the great corpora of the web, a useful representation can emerge from an immense crawl provided that it is afterwards filtered, deduplicated and organised by automatic procedures. The case of GPT-3 suggests that the frequency with which something appears in the world does not determine the frequency with which it is advisable to show it to the learner, so that some sources are repeated and others thinned. With synthetic data, the learner begins to manufacture part of the world from which the next learner will learn. This means that before the model acquires a conception of the world, someone built its learnable world, which constitutes a new cultural object that we still understand poorly.

A training corpus has something of the archive, the canon and the library, but it is above all a selection of culture transformed into possible experience for a machine. Studying it ought to become a task for the Humanities, not because they should replace the engineers in the design of the filters or adopt reactive functions, but in order to understand what those filters do at a cultural scale and perhaps to provide, in the future, better criteria for building them. For example, today’s filters have to distinguish at enormous speed between things that are statistically similar and culturally distinct, such as mechanical duplication and meaningful repetition, noise and popular register, anomaly and informative rarity, bad prose and linguistic variety, redundancy and tradition, statistical marginality and historical importance, and a long list of subtleties that make up the richness of culture and demand that we understand what kind of cultural object we are processing, that body assembled out of fragments, like an industrial Osiris sewn together with pieces of language that suddenly began to speak.

Maps of the Labyrinth, a series from Phantom Maze — AI & Language Lab. Read on Substack.

The Egoless Phantom: Mapping the AI Labyrinth — Phantom Maze Press, 2026. ISBN 978-987-3729-16-4.

References

Brown, Tom B., et al. “Language Models Are Few-Shot Learners.” Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 1877–1901; Kučera, Henry, and W. Nelson Francis. Computational Analysis of Present-Day American English. Brown University Press, 1967; Gao, Leo, et al. “The Pile: An 800GB Dataset of Diverse Text for Language Modeling.” arXiv, 2020, arXiv:2101.00027; Raffel, Colin, et al. “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.” Journal of Machine Learning Research, vol. 21, no. 140, 2020, pp. 1–67; Dodge, Jesse, et al. “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, 2021, pp. 1286–1305; Soldaini, Luca, et al. “Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, 2024, pp. 15725–15788; Penedo, Guilherme, et al. “The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.” Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 30811–30849; Lee, Katherine, et al. “Deduplicating Training Data Makes Language Models Better.” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, 2022, pp. 8424–8445; Hoffmann, Jordan, et al. “Training Compute-Optimal Large Language Models.” Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 30016–30030; Villalobos, Pablo, et al. “Position: Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data.” Proceedings of the 41st International Conference on Machine Learning, vol. 235, Proceedings of Machine Learning Research, 2024, pp. 49523–49544; Shumailov, Ilia, et al. “AI Models Collapse When Trained on Recursively Generated Data.” Nature, vol. 631, 2024, pp. 755–759.

Newsletter

Occasional notes on new research and the book. No spam — unsubscribe anytime.