What exactly is Apache Spark, especially for those who have never worked with it?
Let's simplify the concept. Spark is a powerful tool for in-memory processing. And it one of the best data engineering tool that might cover 80% of data engineering use cases.
It reads raw data from a source, and optionally, transforms and manipulates it before writing it to a target. When writing data to a destination, we can define a TABLE that points to the data files in storage. This table is not for storing data; it's just metadata, containing information about the data's path, partitions, indexes, etc.
An important thing to note about Apache Spark is that it's neither a data warehouse nor a data lake. It's a distributed computing engine that runs Spark software.
The choice of where and in what format to store data is up to us.
In terms of how Spark handles data, it uses an abstraction called a dataframe, similar to a traditional database table with columns and rows. The key difference is that a dataframe is an abstraction of the data, whereas a table usually contains the data. This distinction might not always hold true for systems like Trino, Presto, Athena, and various serverless offerings from vendors.
So, in a nutshell, remember that Spark is a computing engine that facilitates data READ, WRITE, and TRANSFORMATION operations. It's a widely used solution for building Data Lakes or Lakehouses, often utilizing the Delta format.
Furthermore, at Surfalytics, we're gearing up to focus on Apache Spark, Databricks, and Delta Lake starting in January 2024. We'll start with basics like running Spark on a local laptop and using Docker. After gaining a thorough understanding of Spark, we'll advance to Databricks and Delta Lake, exploring more complex topics like Unity Catalog, Structured Streaming, MLOps, and LLMs.
Are you interested in learning in a cohort-based setting and advancing your career? Join us for the cost of a Netflix subscription ->
https://surfalytics.com/