Efficient Transformers - Lecture 20

Overview

Explore efficient transformers in this lecture from MIT's TinyML and Efficient Deep Learning Computing course. Dive into techniques for optimizing transformer models to run on resource-constrained devices like mobile phones and IoT hardware. Learn about model compression, pruning, quantization, neural architecture search, and knowledge distillation approaches to reduce the computational and memory requirements of transformer architectures. Discover how to apply these methods to enable powerful natural language processing capabilities on edge devices. Gain practical insights for deploying transformer-based AI applications in mobile and embedded systems. Access accompanying slides and resources to reinforce key concepts covered in the 1 hour 18 minute video lecture led by Professor Song Han of the MIT HAN Lab.

Syllabus

Lecture 20 - Efficient Transformers | MIT 6.S965

Taught by

MIT HAN Lab

Reviews

Start your review of Efficient Transformers - Lecture 20

Taught by

Efficient Transformers - Lecture 20

Neural Architecture Search for Efficient Deep Learning - Lecture 9

Efficient Video Understanding and Generative Models - Lecture 19

Efficient Video Understanding and Generative Models - Lecture 19

TinyML and Efficient Deep Learning Computing - Lecture 24: Course Summary

TinyEngine - Efficient Training and Inference on Microcontrollers - Lecture 17

Never Stop Learning.