Lesson 24 of 29 · Language Models
Recurrent Neural Networks, Transformers and Attention — MIT 6.S191
The academic version of the transformer chapters above, tracing the path from recurrent networks to attention and explaining why the older approach was abandoned. Good for seeing the same idea derived rather than illustrated.
Video not playing? Some channels disable embedding, and there is no way around it.Open on YouTubeinstead — then come back and mark it complete.