U.S. Patent Phrase to Phrase Matching ended 1 week ago. We (Me, Lukasz Borecki, and Remek Kinas) we're participating there for almost 2 months, and we were able to get 14th place (first position on the silver medal zone) on Public Leaderboard, but unfortunately, then we took only 26th place (after disqualifying some teams) on Private Leaderboard of 1850+ teams! This is my first ever competition medal in Kaggle!
The task was to predict semantic similarity between two sequences but based on some CPC contexts (e.g. Math, Engineering, Chemistry, Science, Physics, etc.). For example, "car" and "washing machine" in the context of "Engineering" are quite semantic similar sequences, but in the context of "Driving," they are different.
We tried various approaches: Pseudo-Labeling (creating new pairs), Pretraining (since most of the BERT-like architectures weren't pretrained on scientific texts), Adversarial Training (to improve model generalization ability), Siamese Networks (to process two input sequences independently), Contrastive Learning (to increase/decrease distances between two embeddings), however, most of the approaches worked worse, so we just ended with an ensemble of different model architectures (bert-for-patents, deberta-v3-large, roberta-large, deberta-v3-small, etc.), some of the models had custom heads, and additionally we did strong post-processing. Also, I would like to emphasize that we did a custom warmup, we freeze some model's layers for a certain number of epochs, and then unfreeze them and change the learning rate. Such training setup gave significant improvement.
Despite this, even such a pipeline was very far from top places, so now I am going to tell you about some interesting approaches, that were used by top participants in the competition:
1. Adversarial Training. We tried it, but Adversarial Training methods require careful tuning of its parameters, e.g epsilon.
2. Prompt Learning. This is a new methodology in NLP, here model trains on Language Modeling (LM) tasks to predict some tokens, which are labels.
Generally, I am very happy with our final results. My advice, try everything that you thought! Competitions are experience, experience is more expensive than any money!
Good luck!
Post #7
1.34K
- 👍 2