TGViewer
Data science/ML/AI Data science/ML/AI @datascience_bds · 14K subscribers
Post #1229 2.03K
▎Common Data Cleaning Terms

1. Data Cleaning: The process of identifying and correcting inaccuracies, inconsistencies, and errors in a dataset to improve its quality and reliability for analysis.

2. Missing Values: Data points that are absent or not recorded in a dataset; handling missing values is crucial for accurate analysis.

3. Outliers: Data points that deviate significantly from the rest of the dataset; identifying and addressing outliers is important to prevent skewed results.

4. Data Imputation: The method of replacing missing values with substituted values, which can be based on statistical methods, such as mean, median, or mode, or predictive models.

5. Normalization: The process of adjusting values in a dataset to a common scale, often to eliminate units of measurement or to reduce skewness.

6. Standardization: A technique used to center and scale data by transforming it to have a mean of zero and a standard deviation of one, making it suitable for comparison.

7. Deduplication: The process of identifying and removing duplicate records from a dataset to ensure each entry is unique.

8. Data Transformation: The process of converting data from one format or structure into another, often to improve compatibility with analytical tools or models.

9. Data Validation: The process of checking data for accuracy and quality before it is processed or analyzed, ensuring it meets predefined criteria.

10. Data Type Conversion: Changing the data type of a variable (e.g., from string to integer) to ensure consistency and compatibility in analysis.

11. String Manipulation: Techniques used to modify or extract information from text data, including trimming, concatenation, and pattern matching.

12. Categorical Encoding: The process of converting categorical variables into numerical format, such as one-hot encoding or label encoding, to facilitate analysis.

13. Data Profiling: The examination of data sources to understand their structure, content, relationships, and quality; often used to identify issues that need cleaning.

14. Anomaly Detection: The identification of unusual patterns or deviations in data that may indicate errors or significant events requiring further investigation.

15. Data Aggregation: The process of summarizing data points into a single value, such as calculating averages or totals, often used for reporting purposes.

16. Data Filtering: The process of removing unwanted or irrelevant data points from a dataset based on specific criteria or conditions.

17. Data Enrichment: The process of enhancing existing data by adding additional information from external sources to provide more context or insights.

18. Schema Validation: Ensuring that the structure of the dataset adheres to a predefined schema, including the correct data types and relationships between entities.

19. Data Sampling: The selection of a subset of data points from a larger dataset for analysis, often used when working with large datasets to reduce processing time.

20. Data Pipeline: A series of processes through which raw data is collected, cleaned, transformed, and made ready for analysis or storage in a database.
  • ❤ 4
  • 🔥 3
More from @datascience_bds
  1. Oct 9, 20268 RAG architectures for AI Engineers, visually explained:
  2. Oct 8, 2026document post
  3. Oct 7, 2026🧮 NumPy: Why axis=0 and axis=1 Feel Backwards You've probably seen: np.mean(X, axis=0) an…
  4. Oct 6, 2026document post
  5. Oct 5, 2026📊 Pandas Cheatsheet Every Data Analyst Should Save Pandas is one of the most important to…
  6. Oct 4, 2026SQLBolt: Interactive SQL You can learn SQL by writing real queries directly in the browser…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →