It should be really useful as according to this paper
https://arxiv.org/abs/1905.05583, the unsupervised finetuning and layer wise LR , and one-cycle are crucial for BERT performance. They mange to beat ULMFiT on IMDB with BERT-Base
Post #2137
1.31K