To create high-quality embeddings that work effectively in an RAG pipeline, one needs to utilize pre-trained embedding models, like OpenAI’s embedding models. There are various types of embeddings one can create and corresponding models available. For instance:
Word Embeddings: In word embeddings, each word has a fixed vector regardless of context. Popular models for creating this type of embedding are Word2Vec and GloVe.
Contextual Embeddings: Contextual embeddings take into account that the meaning of a word can change based on context. Take, for instance, the bank of a river and opening a bank account. Some models that can be used for producing contextual embeddings are BERT and OpenAI’s embedding models (like text-embedding-ada-002)
Sentence Embeddings: These are embeddings capturing the meaning of full sentences. A popular embedding model for creating sentence embeddings is Sentence-BERT.
Post #797
48
AI Scope https://towardsdatascience.com/rag-explained-understanding-embeddings-similarity-and-retrieval/?utm_source=roadmap&utm_medium=Referral&utm_campaign=TDS+roadmap+integration