TGViewer
Data science/ML/AI Data science/ML/AI @datascience_bds · 14K subscribers
Post #1238 1.76K
▎Common Data Analysis Terms

1. Data Cleaning: The process of correcting or removing inaccurate records from a dataset.

2. Exploratory Data Analysis (EDA): Analyzing datasets to summarize their main characteristics, often using visual methods.

3. Statistical Analysis: The application of statistical methods to collect, review, and draw conclusions from data.

4. Data Visualization: The graphical representation of information and data to communicate insights effectively.

5. Machine Learning: A subset of AI that enables systems to learn from data and improve performance without explicit programming.

6. Predictive Analytics: Techniques that use statistical algorithms and machine learning to identify the likelihood of future outcomes based on historical data.

7. Data Mining: The practice of examining large datasets to uncover patterns and relationships.

8. Feature Engineering: The process of selecting, modifying, or creating new features from raw data to improve model performance.

9. Outlier Detection: Identifying and handling anomalies in data that do not conform to expected patterns.

10. Clustering: A method of grouping similar data points together based on their characteristics.

11. Natural Language Processing (NLP): A field of AI that enables computers to understand and interpret human language.

12. Data Ethics: The study of moral issues related to data collection, analysis, and usage, including privacy concerns.

13. Data Sampling: Selecting a subset of individuals from a population to estimate characteristics of the whole population.

14. SQL (Structured Query Language): A programming language used for managing and querying relational databases.

15. NoSQL Databases: Non-relational databases designed to handle large volumes of unstructured or semi-structured data.

16. Data Integration: Combining data from different sources into a unified view for analysis.

17. Hyperparameter Tuning: The process of optimizing the parameters that govern the training of machine learning models.

18. Cross-Validation: A technique for assessing how the results of a statistical analysis will generalize to an independent dataset.

19. Ensemble Methods: Techniques that combine multiple models to improve prediction accuracy, such as bagging and boosting.

20. KPI (Key Performance Indicator): A measurable value that demonstrates how effectively a company is achieving key business objectives.
  • ❤ 5
More from @datascience_bds
  1. Oct 9, 20268 RAG architectures for AI Engineers, visually explained:
  2. Oct 8, 2026document post
  3. Oct 7, 2026🧮 NumPy: Why axis=0 and axis=1 Feel Backwards You've probably seen: np.mean(X, axis=0) an…
  4. Oct 6, 2026document post
  5. Oct 5, 2026📊 Pandas Cheatsheet Every Data Analyst Should Save Pandas is one of the most important to…
  6. Oct 4, 2026SQLBolt: Interactive SQL You can learn SQL by writing real queries directly in the browser…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →