🤔TOP 4 risks of embeddings in ML
Embedded models are widely used in machine learning to translate raw input data into a low-dimensional vector that captures its semantic meaning and can be used for various subsequent models. Pretrained embeddings as feature extractors are used to develop features of various input types (text, image, audio, video, multimodal) or categorical features with high cardinality.
The main risks of embeddings are:
• High obfuscation - Changing the output of the upstream embedding model affects the performance of the downstream model. Relying on the output of the embedding model, downstream models should be retrained.
• Hidden feedback loops. Pre-trained embedding features are often used as a black box. But knowing what the raw input embedding was trained on is very important for the quality of the model and its interpretability.
• High costs of real-time output (storage and maintenance). This directly affects the return on investment in ML. It is important to ensure quality during embedding service outages, cost of training, and cost of service per request.
• High cost of debugging: - debugging and monitoring built-in features or root causes of failures is very expensive. Therefore, built-in features that are not of great importance for the model should be abandoned.
https://medium.com/better-ml/embeddings-the-high-interest-credit-card-of-feature-engineering-414c00cb82e1
Post #527
770