As we discussed embeddings can be different depending on the task, but one common task is predicting the context of a word. This is the foundation of how large language models (LLMs) work. LLMs don’t actually understand human language; instead, they understand the numerical relationships between words.
Key properties of a good word embedding:
✔️ Similar words should have similar vector scores.
✔️ Different words should have vector scores with values that are far away from each other.
When each word or data point has a single embedding vector, this is called a static embedding. Static embeddings, once trained, contain some semantic information, especially in how words relate to each other. It means that word similarity depends on the data the model was trained on.
To address the limitations of static embeddings, contextual embeddings were developed. Contextual embeddings allow a word to have multiple representations based on its surrounding context. For example, the word
orange can mean a fruit or a color, so it would have a different embedding for each meaning depending on the context.Tensorflow offers a great tool called Projector where you can play with `word2vec` static embedding model and see how trained data impacts words relationships.
Embeddings can be used for words, sentences and even documents. Interesting feature that is built on top of that - context search. Instead of searching by a particular word, you can search by the meaning of that word and find a related document even if there are no exact matches.
#aibasics