🥁Data engineering for Data Science: SynapseML by Microsoft
Microsoft has adapted Apache Spark for the tasks of data engineers and DS specialists, released SynapseML - a framework for creating scalable ML pipelines. This open source library was previously called MMLSpark. SynapseML powered by SparkML to the Spark ecosystem of deep learning and data analysis, seamless ML pipeline transformation tools with Open Neural Network Exchange (ONNX), LightGBM, Cognitive Services, Vowpal Wabbit, and OpenCV. This allows you to create powerful and highly scalable predictive and analytical models for a variety of data sources.
It is noteworthy that SynapseML is able to work with untagged datasets thanks to API methods of ready-made AI services for quickly solving typical ML tasks. SynapseML requires Scala 2.12, Spark 3.0+, and Python 3.6+. The framework allows you to write code in any Spark-compatible language: Python, Scala, R, Java, .NET, and C #. Over the HTTP protocol, users can embed any web service into their SparkML models, and the clustered nature of Spark allows ML projects to scale.
https://microsoft.github.io/SynapseML/
https://github.com/microsoft/SynapseML
Post #354
462