Statistical Data Validation for Pandas.
A data validation library for scientists, engineers, and analysts seeking correctness.
pandera provides a flexible and expressive API for performing data validation on tidy (long-form) and wide data to make data processing pipelines more readable and robust.
pandas data structures contain information that pandera explicitly validates at runtime. This is useful in production-critical data pipelines or reproducible research settings. With pandera, you can:
- Check the types and properties of columns in a pd.DataFrame or values in a pd.Series.
- Perform more complex statistical validation like hypothesis testing.
- Seamlessly integrate with existing data analysis/processing pipelines via function decorators.
- Define schema models with a class-based API with pydantic-style syntax and validate dataframes using the typing syntax.
https://pandera.readthedocs.io/en/stable/index.html
#python #ds
Post #734
3.05K