Ilya Sutskever Recommended Reading

Ilya Sutskever gave John Carmack about 30 papers and said that if you really learn all of them, you'll know 90% of what matters. Carmack had asked how to get current on AI.

I have not worked through these yet. This note is the list, with links filled in where the circulating dump left titles bare.

List

  • The Annotated Transformer. Sasha Rush, et al. blog · code
  • The First Law of Complexodynamics. Scott Aaronson. blog
  • The Unreasonable Effectiveness of Recurrent Neural Networks. Andrej Karpathy. blog · code
  • Understanding LSTM Networks. Christopher Olah. blog
  • Recurrent Neural Network Regularization. Wojciech Zaremba, et al. arxiv · code
  • Keeping Neural Networks Simple by Minimizing the Description Length of the Weights. Geoffrey E. Hinton and Drew van Camp. acm · pdf
  • Pointer Networks. Oriol Vinyals, et al. arxiv
  • ImageNet Classification with Deep Convolutional Neural Networks. Alex Krizhevsky, et al. nips
  • Order Matters: Sequence to sequence for sets. Oriol Vinyals, et al. arxiv
  • GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism. Yanping Huang, et al. arxiv
  • Deep Residual Learning for Image Recognition. Kaiming He, et al. arxiv
  • Multi-Scale Context Aggregation by Dilated Convolutions. Fisher Yu and Vladlen Koltun. arxiv
  • Neural Message Passing for Quantum Chemistry. Justin Gilmer, et al. arxiv
  • Attention Is All You Need. Ashish Vaswani, et al. arxiv
  • Neural Machine Translation by Jointly Learning to Align and Translate. Dzmitry Bahdanau, et al. arxiv
  • Identity Mappings in Deep Residual Networks. Kaiming He, et al. arxiv
  • A simple neural network module for relational reasoning. Adam Santoro, et al. arxiv
  • Variational Lossy Autoencoder. Xi Chen, et al. arxiv
  • Relational recurrent neural networks. Adam Santoro, et al. arxiv
  • Quantifying the Rise and Fall of Complexity in Closed Systems: The Coffee Automaton. Scott Aaronson, et al. arxiv
  • Neural Turing Machines. Alex Graves, et al. arxiv
  • Deep Speech 2: End-to-End Speech Recognition in English and Mandarin. Dario Amodei, et al. arxiv
  • Scaling Laws for Neural Language Models. Jared Kaplan, et al. arxiv
  • A Tutorial Introduction to the Minimum Description Length Principle. Peter Grunwald. arxiv
  • Machine Super Intelligence. Shane Legg. pdf
  • Kolmogorov Complexity and Algorithmic Randomness. A. Shen, V. A. Uspensky, and N. Vereshchagin. ams
  • CS231n: Convolutional Neural Networks for Visual Recognition. course

AI Alignment LLM Ethics

Bibliography

Last updated