Let's take a high-level look at the history of #dataengineering and #analytics. We can examine at least the past 40 years to identify several key moments that shaped the industry.
In the past, organizations heavily relied on relational databases. For any kind of data job, they needed programmers to write custom software for data integration or visualization. A simple report with a few charts required significant effort. The first release of Microsoft Excel, therefore, was a game-changer.
Smart individuals saw an opportunity to develop products focusing on these pain points. They introduced tools like BusinessObjects, Cognos, and Crystal Reports in the field of BI. Similar developments happened in data integration and mining (yes, before the 'sexiest job' of data scientist emerged).
Later, these companies were acquired by major vendors like SAP, Oracle, and IBM. This era was marked by everything "Enterprise" grade, such as BI, ETL, DW. At the same time, Informatica was launched, targeting the "enterprise" market.
The volume of relational data grew, and companies like Teradata leveraged horizontal scaling and "shared nothing" architecture, popularizing the MPP approach. Others introduced their MPPs like Exadata, Netezza, and Vertica.
However, this was quite expensive, and only large "enterprise" companies could afford these technologies. Many didn't realize they were falling into a trap of vendor lock-in, with many S&P companies still using Teradata.
After the 2000s, the rise of Hadoop introduced the key idea of decoupling storage and compute, allowing companies to process and store various data types on a larger scale and more affordably than with "enterprise" DWs. These tools often complemented each other. I almost learned Java to write MapReduce jobs – I'm glad I didn't. Due to Hadoop's imperfections, we saw the emergence of many great products, including Apache Spark.
Simultaneously, the field of Data Science grew, and there was confusion about the term 'data model', which could refer to a data science model or an ER diagram. There was also a surge in Python and R, with numerous free courses on Coursera about data science and R. I almost learned R – again, I'm glad I didn't, as I still see some data engineering projects struggling with R in their pipelines.
Finally, with the rise of Cloud Computing, a new way of building data solutions emerged, starting small and cheap. There are three major players in the public cloud – AWS, Azure, GCP – each with their own set of analytics tools. The first notable one was Redshift. I would erect a monument to the Redshift data warehouse for changing our perception of traditional data warehouses and leading us to consider cloud data warehouses.
We have witnessed many new concepts and products that have transformed the industry, encouraging vendors and companies to move forward with a faster pace and a better engineering experience, exemplified by Snowflake and Databricks.
PS please like here: https://www.linkedin.com/posts/dmitryanoshin_dataengineering-analytics-activity-7134984496461844481-vTGI
PPS It is very important understand this history to see that every new tool is 50% old tool and due to new techmologies like Hadoop, Cloud it has new features and close some pain points.
Post #85
285