---
title: "Ilya Sutskever Recommended Reading"
slug: ilya-sutskever-recommended-reading
type: idea
stime: 2026-08-13 @ 6:08PM
status: growing
certainty: likely
importance: 6
tags: [ai, deep-learning, reading]
---

[Ilya Sutskever](ilya-sutskever.md) gave John Carmack about 30 papers and said that if you really learn all of them, you'll know 90% of what matters. Carmack had asked how to get current on AI.

I have not worked through these yet. This note is the list, with links filled in where the circulating dump left titles bare.

## List

- **The Annotated Transformer.** Sasha Rush, et al. [blog](https://nlp.seas.harvard.edu/annotated-transformer/) · [code](https://github.com/harvardnlp/annotated-transformer/)
- **The First Law of Complexodynamics.** Scott Aaronson. [blog](https://scottaaronson.blog/?p=762)
- **The Unreasonable Effectiveness of Recurrent Neural Networks.** Andrej Karpathy. [blog](https://karpathy.github.io/2015/05/21/rnn-effectiveness/) · [code](https://github.com/karpathy/char-rnn)
- **Understanding LSTM Networks.** Christopher Olah. [blog](https://colah.github.io/posts/2015-08-Understanding-LSTMs/)
- **Recurrent Neural Network Regularization.** Wojciech Zaremba, et al. [arxiv](https://arxiv.org/abs/1409.2329) · [code](https://github.com/wojzaremba/lstm)
- **Keeping Neural Networks Simple by Minimizing the Description Length of the Weights.** Geoffrey E. Hinton and Drew van Camp. [acm](https://dl.acm.org/doi/10.1145/168304.168306) · [pdf](https://www.cs.toronto.edu/~hinton/absps/colt93.pdf)
- **Pointer Networks.** Oriol Vinyals, et al. [arxiv](https://arxiv.org/abs/1506.03134)
- **ImageNet Classification with Deep Convolutional Neural Networks.** Alex Krizhevsky, et al. [nips](https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks)
- **Order Matters: Sequence to sequence for sets.** Oriol Vinyals, et al. [arxiv](https://arxiv.org/abs/1511.06391)
- **GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism.** Yanping Huang, et al. [arxiv](https://arxiv.org/abs/1811.06965)
- **Deep Residual Learning for Image Recognition.** Kaiming He, et al. [arxiv](https://arxiv.org/abs/1512.03385)
- **Multi-Scale Context Aggregation by Dilated Convolutions.** Fisher Yu and Vladlen Koltun. [arxiv](https://arxiv.org/abs/1511.07122)
- **Neural Message Passing for Quantum Chemistry.** Justin Gilmer, et al. [arxiv](https://arxiv.org/abs/1704.01212)
- **Attention Is All You Need.** Ashish Vaswani, et al. [arxiv](https://arxiv.org/abs/1706.03762)
- **Neural Machine Translation by Jointly Learning to Align and Translate.** Dzmitry Bahdanau, et al. [arxiv](https://arxiv.org/abs/1409.0473)
- **Identity Mappings in Deep Residual Networks.** Kaiming He, et al. [arxiv](https://arxiv.org/abs/1603.05027)
- **A simple neural network module for relational reasoning.** Adam Santoro, et al. [arxiv](https://arxiv.org/abs/1706.01427)
- **Variational Lossy Autoencoder.** Xi Chen, et al. [arxiv](https://arxiv.org/abs/1611.02731)
- **Relational recurrent neural networks.** Adam Santoro, et al. [arxiv](https://arxiv.org/abs/1806.01822)
- **Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton.** Scott Aaronson, et al. [arxiv](https://arxiv.org/abs/1405.6903)
- **Neural Turing Machines.** Alex Graves, et al. [arxiv](https://arxiv.org/abs/1410.5401)
- **Deep Speech 2: End-to-End Speech Recognition in English and Mandarin.** Dario Amodei, et al. [arxiv](https://arxiv.org/abs/1512.02595)
- **Scaling Laws for Neural Language Models.** Jared Kaplan, et al. [arxiv](https://arxiv.org/abs/2001.08361)
- **A Tutorial Introduction to the Minimum Description Length Principle.** Peter Grunwald. [arxiv](https://arxiv.org/abs/math/0406077)
- **Machine Super Intelligence.** Shane Legg. [pdf](https://www.vetta.org/documents/Machine_Super_Intelligence.pdf)
- **Kolmogorov Complexity and Algorithmic Randomness.** A. Shen, V. A. Uspensky, and N. Vereshchagin. [ams](https://www.ams.org/books/mmono/220/)
- **CS231n: Convolutional Neural Networks for Visual Recognition.** [course](https://cs231n.stanford.edu/)

[AI Alignment](ai-alignment.md)
[LLM Ethics](llm-ethics.md)

## Related

- [AI Agent Resources](ai-agent-resources.md)
- [cheap intelligence enables retroactive idea mining](cheap-intelligence-enables-retroactive-idea-mining.md)
- [Agentic Port Of Sophia to Rust](agentic-port-of-sophia-to-rust.md)
- [claudecode](claudecode.md)

## Bibliography

- [Keshavchan's tweet circulating the list](https://twitter.com/keshavchan/status/1787861946173186062)
- [Arc folder of the same list](https://arc.net/folder/D0472A20-9C20-4D3F-B145-D2865C0A9FEE)
- [dzyim/ilya-sutskever-recommended-reading](https://github.com/dzyim/ilya-sutskever-recommended-reading), the dump this note was copied from
