👀TOP 5 Open Source Data lineage Tools
Data lineage allows a company to comply with regulatory requirements, better understand and trust its data, and save time on manually analyzing the effects of data changes. To monitor data, there are many tools on the market today, open source and proprietary. Of the source code tools, the most important are:
• Tokern allows users to get column-level data passing data from databases and data stores hosted by Google BigQuery, AWS Redshift, and Snowflake. Tokern can be integrated with other bugs as well. it works well with a large number of data directories with large source code and ETL frameworks. Tokern also provides users with the ability to generate data from external sources or its frequency ETL, making an overview for high end BI and ETL tools. Tokern uses PostgreSQL as its local storage, and NetworkX for quick analysis of network graphs. These available users can be manipulated, visualized, and analyzed in column-level library origin data. Additionally, users can also interact with walkthrough data using the Tokern SDK or API. Tokern also discovered PII (personal information) and PHI (personal health information) using PIICatcher, combining regular expressions with NLP libraries Spacy and Stanford NER.
• Egeria - a metadata standard with original source code, enabling seamless integration of data processing tools for a robust and unified representation of metadata. Egeria allows you to build better solutions for data production, data quality checks, PII identification, and more. in addition to cataloging and metadata retrieval. Egeria Law on the OpenLineage data collection and storage standard. This allows users to gain a more complete view of the data by providing horizontal and vertical assignment and tracing. To get information about the origin of the data, Egeria listens to the Kafka events sent by the original messages.
• Pachyderm is a data collection tool that empowers developers to create machine learning pipelines regardless of language and framework, instead of focusing on cloud storage. It uses a version control system like LakeFS or Git saves and saves changes like commits keeping a complete and unaltered audit trail. Pachyderm also has a full audit trail and uses a central repository based on object storage in a custom Pachyderm File System to track data generation and version control. Pachyderm collects user data sources, uses global origin event and data object identifiers. Pachyderm allows you to create an immutable data production graph as a DAG in the user interface, which is especially useful when working with ML pipelines. Thick Skin integrates well with many databases, repositories and data lakes. Therefore, many companies use it for MLOps operations, ETL unstructured data, and NLP workflows.
• OpenLineage is a Linux Foundation project based on the highest ETL platforms, data orchestration mechanisms, metadata catalogs, data quality control mechanisms, and data inheritance tools. OpenLineage uses JSONSchema as the API definition and supports multiple languages and platforms. The previously mentioned Egeria has a core metadata layer built on top of OpenLineage. WeWork's Marquez also underpins the OpenLineage architecture, a preferred UI and metadata repository, as well as an API for collecting metadata, validating via GraphQL, and a REST API.
• TrueDat is a comprehensive data management solution that allows you to efficiently classify, search and evaluate data, as well as visualize the entire data lifecycle. The tool was created in 2017 by BlueTab, which is part of IBM, and is still actively developed.
https://blog.devgenius.io/5-best-open-source-data-lineage-tools-in-2022-f8ef39a7d5f6
Post #512
870
- 👍 1