Recently, I published an article about optimization approaches to speed up training and reduce memory consumption during training large-scale (parameters scale) NLP models such as Transformers.
There are explanations of each approach and additionally, most of the proposed approaches were implemented via PyTorch and HuggingFace API.
Firstly, the article describes basic methods such as Gradient Accumulation, Automatic Mixed Precision, Freezing, etc., and then goes throw more NLP-specific approaches such as Dynamic Padding, Uniform Dynamic Padding, and Fast Tokenizers.
Thereby, there is no need to have powerful GPUs or buy expensive clusters to train Transformers!
Enjoy!
https://www.kaggle.com/code/vad13irt/optimization-approaches-for-transformers
Post #6
1.13K