In the Data Vectorization post, we learned that ML algorithms work with feature vectors. For example, we had a vector for car colors, where a blue car was represented as
(0.0, 1.0, 0.0, 0.0, 0.0).Imagine we create feature vectors for meal items in a dataset with 5,000 different elements. Each vector would have 5,000 elements, all set to
0.0 except for one element set to 1.0. This approach would require a high number of weights, a lot of memory and computational resources, making the model inefficient and hard to maintain.To optimize this embedding techniques are used. Embedding is a projection of high-dimensional space of initial data vectors into a lower-dimensional space.
For example, in the meal dataset, we could introduce a feature like "sandwichness" and evaluate how likely an item is a sandwich. A sandwich might have a score
0.99, shawarma 0.9, and soup 0.0. But one feature isn’t enough, so we could add other dimensions, like "dessertness", "vegannes" or "liquidness," to better describe each item.With features like sandwichness, dessertness, and liquidness, the vector for a hotdog might look like
(0.95, 0.8, 0.0). Real-world models use many dimensions, but these vectors are much shorter and more efficient than the original 5,000-element vectors with only 0.0 and 1.0 values.ML practitioners select dimensions based on the task they want to solve. This means that embeddings for the same items can be different depending on the context, task and provided data.
References:
- What are Embeddings in ML?
- ML Google Crash Course: Embeddings
#aibasics