Most AI reading lists are assembled by people who have read one half of the field, and you can tell which half within two entries.
A list made only of philosophy produces someone who argues about machine intelligence fluently and cannot tell you what a gradient is. They hold opinions about something they cannot describe. The engineering-only list produces the reverse, people who can fine-tune a model and have never asked what understanding would consist of.
Eighteen works below, and the order is the point. Mechanism first, because none of the philosophy means much until you know what the machine is doing.
Papers link straight to a free copy where one exists. Four of the books went out of print decades ago, so those carry an ISBN for the secondhand market instead of a shop link that would only fail.
Stage one, build both
The field spent 40 years arguing between two ways of making a machine intelligent, so build one of each before reading a word of the argument.
Michael Nielsen, Neural Networks and Deep Learning, and Andrej Karpathy’s Neural Networks Zero to Hero at karpathy.ai/zero-to-hero.html. Nielsen’s book is free online and should be worked through with a pen. Karpathy’s series has you write a small GPT from scratch and watch it learn. That is the connectionist side.
Peter Norvig, Paradigms of Artificial Intelligence Programming, 1992, free under MIT licence at github.com/norvig/paip-lisp. Chapter 4 builds GPS, Newell and Simon’s own General Problem Solver, and chapter 5 builds Weizenbaum’s ELIZA. That is the symbolic side, and you will meet both authors again further down this list, this time as combatants.
Two weekends for the network, one afternoon for ELIZA, and nothing else on this page pays back as much.
Vaswani and colleagues, Attention Is All You Need, 2017. arxiv.org/abs/1706.03762
The architecture behind every model you have used. Read cold it looks like magic. Read after you have built the pieces it looks like engineering, which is what it is. Everything below depends on having made that switch, and almost everybody skips this stage, which is why so much commentary on AI is confidently wrong about mechanism.
Stage two, where the idea came from
Turing, Computing Machinery and Intelligence, 1950. Mind LIX, 236, pages 433 to 460. courses.cs.umbc.edu/471/papers/turing.pdf
Everyone cites it and few have read it. The imitation game occupies a small part of the paper and the rest answers objections people still raise as though they were new.
McCarthy, Minsky, Rochester and Shannon, the Dartmouth proposal, 1955. www-formal.stanford.edu/jmc/history/dartmouth/dartmouth.html
The document that named the field. Read it for the confidence. Four men proposed that 10 people working for two months might make significant progress on language, abstraction and self-improvement. Anybody publishing an AGI roadmap should be made to read it first.
Stage three, the three research programmes arguing
In this sequence, because each rebuts the one before. Out of order they are three opinions.
Newell and Simon, Computer Science as Empirical Inquiry: Symbols and Search, 1976. Communications of the ACM 19(3), pages 113 to 126. Their lecture for the 1975 Turing Award, stating the physical symbol system hypothesis: such a system has the necessary and sufficient means for general intelligent action. The boldest claim anybody has made about intelligence, stated by the two men who believed it most.
Rumelhart, Hinton and Williams, Learning representations by back-propagating errors, Nature 323, 1986, pages 533 to 536. Four pages. Every idea in it will be familiar because you implemented them in stage one. Worth knowing that the method has contested parentage, discovered by Werbos, rediscovered by Parker, and made famous by these three.
Brooks, Intelligence Without Representation, Artificial Intelligence 47, 1991, pages 139 to 159. people.csail.mit.edu/brooks/papers/representation.pdf
Brooks declares that AI research had foundered on representation, then rejects both preceding programmes in favour of creatures that act in the world. He opens by asking you to imagine the 1890s, when artificial flight was the glamour subject for science and venture money. Worth reading for that opening alone, which is 400 words on why imitating a phenomenon tells you little about the mechanism.
Boden, Artificial Intelligence and Natural Man, 1977. Out of print since well before most of its readers were working, ISBN 9780855277000. archive.org/details/artificialintell0000bode_k9b4
537 pages, published out of Hassocks in Sussex, and still the best book on what AI research reveals about people. Boden explained the programme rather than attacking it, and went on to found computational creativity as a subfield. Her distinction in The Creative Mind between combinational, exploratory and transformational creativity does more work on whether generative models create anything than three years of commentary has managed.
Stage four, the critics, and what each objected to
Grouping these four as people who said it would not work is wrong, and the distinction is worth making. Dreyfus argued from phenomenology that symbolic AI would stall, which is a claim about feasibility. Winograd and Flores argued the foundations were mistaken. Searle argued that a running program is not sufficient for understanding, which leaves feasibility untouched. Weizenbaum’s question was which judgements humans should never delegate at all.
Weizenbaum, Computer Power and Human Reason, 1976. Freeman original ISBN 9780716704645, Penguin edition ISBN 9780140179118, both secondhand. Borrow at the Internet Archive
He built ELIZA in 1966, watched people confide in 200 lines of pattern matching, and spent his career appalled at what that revealed. Anybody deploying a conversational system where users are vulnerable is working on ground he mapped 50 years ago.
Searle, Minds, Brains and Programs, Behavioral and Brain Sciences 3, 1980. DOI 10.1017/S0140525X00005756.
The Chinese Room, and the most misread argument in the field. Searle does not claim machines cannot think. He claims running the right program is not sufficient for understanding. Read the replies printed alongside it, which are half the value.
Winograd and Flores, Understanding Computers and Cognition, 1986. Out of print, ISBN 9780201112979. archive.org/details/understandingcom00wino
The strongest entry here. Winograd built SHRDLU, the flagship success of symbolic AI, then wrote a book explaining why the whole approach was mistaken. He had built it, which is why the rejection reads differently from an outsider’s. Weizenbaum called it groundbreaking, and it draws on Heidegger and Gadamer, so read it with Dreyfus.
Dreyfus, What Computers Still Can’t Do, 1992. MIT Press paperback, still in print, ISBN 9780262540674.
The revised edition of his 1972 book, which the field treated as an attack and refused to engage with. He argued that expertise is embodied and not rule-governed, so symbolic AI would stall. It stalled.
Stage five, causality, which is the actual gap
Pearl and Mackenzie, The Book of Why, 2018. Penguin paperback, ISBN 9780141982410. Waterstones | Blackwell’s | Author’s page
Everything in production models correlation. You cannot see the shape of that limit without the vocabulary to name it. Ask what a system would need to know to answer a question starting “what would have happened if”, then notice that nothing on the market can. Read Pearl before anybody tries to sell you an AI that explains why something occurred.
Stage six, the present, read against itself
Never a single source here. Pair every claim with its correction.
The scaling laws sequence. Kaplan and colleagues 2020, arxiv.org/abs/2001.08361. Then Hoffmann and colleagues 2022, arxiv.org/abs/2203.15556, the Chinchilla paper that corrected it. Then Besiroglu and colleagues 2024, arxiv.org/abs/2404.10102, who tried to replicate Chinchilla. A claim, its correction and a replication attempt in sequence teaches more about how this field works than any commentary.
Bender, Gebru, McMillan-Major and Shmitchell, On the Dangers of Stochastic Parrots, 2021. Read the paper rather than the argument about it, because the argument has almost nothing to do with what it says.
Melanie Mitchell, Artificial Intelligence: A Guide for Thinking Humans, 2019. Pelican paperback, ISBN 9780241404836. Mitchell was Hofstadter’s doctoral student, which makes this the right companion to the bonus read below. The best single account of the gap between benchmark performance and understanding.
Narayanan and Kapoor, AI Snake Oil, 2024. Paperback, ISBN 9780691249148. Waterstones | Blackwell’s | Princeton University Press They do the empirical work rather than posting about it.
Then current mechanistic interpretability work, published mostly at transformer-circuits.pub. It moves too fast to name papers in something meant to last. If a piece of writing about AI capability does not engage with interpretability, treat it as marketing.
Then start again
Agre, Toward a Critical Technical Practice: Lessons Learned in Trying to Reform AI, 1997. In Bowker, Star, Turner and Gasser, editors, Social Science, Technical Systems and Cooperative Work, Erlbaum. On how AI researchers mislead themselves through the vocabulary they choose. Words like planning and reasoning get borrowed from human life, wired to a mechanism, then quietly read back as though the mechanism had earned the word. Read it last, because it sends you through everything above with better questions.
The bonus read
Hofstadter, Gödel, Escher, Bach, 1979.
It sits outside the sequence because it supplies no mechanism, and that is not a criticism of it. Read it for what it does to your thinking about self-reference and formal systems, which nothing else on this list will do for you. Worth reading now for a second reason. In a 2023 interview Hofstadter said the pace of progress had caught him off guard and that he found it terrifying, and named one thing that surprised him completely, which was that deep thinking could come out of a feed-forward network running in one direction only. His unease is older than the current models, though. Melanie Mitchell opens her 2019 book, also on this list, with Hofstadter telling a room of Google engineers in 2014 that he was terrified.
What the sequence is for
By stage six you should be able to meet any vendor claim with three unwelcome questions. What mechanism is doing the work. What was measured, and whether the measure is the thing. Whether the system knows why anything happened, or only what tends to follow what.
Nobody who has worked through stages one to five gets much use out of the phrase artificial general intelligence.