TGViewer
Channel Public Channel
Data science/ML/AI

Data science/ML/AI

@datascience_bds

Data science and machine learning hub

Python, SQL, stats, ML, deep learning, projects, PDFs, roadmaps and AI resources.

For beginners, data scientists and ML engineers
πŸ‘‰ https://rebrand.ly/bigdatachannels

DMCA: @disclosure_bds
Contact: @mldatascientist
Subscribers
14K
Photos
608
Videos
7
Links
341

Showing posts older than #1288 Β· Back to latest

Older Posts 20 shown
Post #1287 1.76K
SQL Query Execution Order
  • ❀ 8
Post #1286 1.77K
πŸš€ Build products faster with ready-to-use Web Data APIs

Introducing CoreClaw β€” a platform that helps developers collect structured data without building and maintaining scrapers.

✨ What you can do:
βœ… Google Maps data extraction
βœ… Instagram posts & comments scraping
βœ… YouTube data collection
βœ… Amazon & LinkedIn data APIs
βœ… JSON / CSV structured output
βœ… API & no-code workflows

Stop spending time maintaining scrapers.
Start building with reliable data.

πŸ”— Try CoreClaw Freeβ€”>https://coreclaw.com
  • ❀ 6
  • πŸ‘ 2
Post #1285 1.73K
πŸ” 5 resources most ML people never stumble on

These are the things practitioners quietly rely on but rarely share.

1. Google's "Rules of Machine Learning"
43 numbered rules from Google engineers on when to add complexity, how to catch training/serving skew, and when a heuristic beats a model. Written from real production postmortems.

2. Chip Huyen's ML Systems Design notes
Free breakdown of how companies actually design ML systems: data pipelines, feature stores, serving latency, model monitoring. The stuff no ML course teaches.

3. Full Stack Deep Learning
Free course built on one premise: training the model is the easy 20%. Covers deployment, cost tradeoffs, data labeling, and how models fail in production.

4. alphaXiv
Same papers as arXiv, but with inline comment threads under each section, sometimes answered by the paper's own authors. Turns a static PDF into an ongoing discussion.

5. Sebastian Raschka's "Ahead of AI"
Newsletter that dissects specific architecture and training decisions (why this optimizer, why this attention variant) at a depth most blogs skip.
  • ❀ 5
Post #1283 2.03K
Statistics for Data Science
  • ❀ 9
Post #1282 1.95K
Machine Learning Roadmap

Machine Learning Roadmap
|
|-- Fundamentals
| |-- Mathematics
| | |-- Linear Algebra
| | |-- Calculus (Gradients, Optimization)
| | |-- Probability and Statistics
| | |-- Matrix Operations
| |
| |-- Programming
| | |-- Python (NumPy, Pandas, Scikit-learn)
| | |-- R (Optional for Statistical Modeling)
| | |-- SQL (For Data Extraction)
|
|-- Data Preprocessing
| |-- Data Cleaning
| |-- Feature Engineering
| | |-- Encoding Categorical Data
| | |-- Feature Scaling (Standardization, Normalization)
| | |-- Handling Missing Values
| |-- Dimensionality Reduction (PCA, LDA)
|
|-- Supervised Learning
| |-- Regression
| | |-- Linear Regression
| | |-- Polynomial Regression
| | |-- Ridge and Lasso Regression
| |-- Classification
| | |-- Logistic Regression
| | |-- Decision Trees
| | |-- Support Vector Machines (SVM)
| | |-- Ensemble Methods (Random Forest, Gradient Boosting, XGBoost)
|
|-- Unsupervised Learning
| |-- Clustering
| | |-- K-Means
| | |-- Hierarchical Clustering
| | |-- DBSCAN
| |-- Dimensionality Reduction
| | |-- Principal Component Analysis (PCA)
| | |-- t-SNE
| |-- Association Rules (Apriori, FP-Growth)
|
|-- Reinforcement Learning
| |-- Markov Decision Processes
| |-- Q-Learning
| |-- Deep Q-Learning
| |-- Policy Gradient Methods
|
|-- Model Evaluation and Optimization
| |-- Train-Test Split and Cross-Validation
| |-- Performance Metrics
| | |-- Accuracy, Precision, Recall, F1-Score
| | |-- ROC-AUC
| | |-- Mean Squared Error (MSE), R-squared
| |-- Hyperparameter Tuning
| | |-- Grid Search
| | |-- Random Search
| | |-- Bayesian Optimization
|
|-- Deep Learning
| |-- Neural Networks
| | |-- Perceptrons
| | |-- Backpropagation
| |-- Convolutional Neural Networks (CNN)
| | |-- Image Classification
| | |-- Object Detection (YOLO, SSD)
| |-- Recurrent Neural Networks (RNN)
| | |-- LSTM
| | |-- GRU
| |-- Transformers (Attention Mechanisms, BERT, GPT)
| |-- Tools and Frameworks (TensorFlow, PyTorch)
|
|-- Advanced Topics
| |-- Transfer Learning
| |-- Generative Adversarial Networks (GANs)
| |-- Reinforcement Learning with Neural Networks
| |-- Explainable AI (SHAP, LIME)
|
|-- Applications of Machine Learning
| |-- Recommender Systems (Collaborative Filtering, Content-Based)
| |-- Fraud Detection
| |-- Sentiment Analysis
| |-- Predictive Maintenance
| |-- Autonomous Vehicles
|
|-- Deployment of Models
| |-- Flask, FastAPI
| |-- Cloud Deployment (AWS SageMaker, Azure ML)
| |-- Containerization (Docker, Kubernetes)
| |-- Model Monitoring and Retraining



Enjoy Learning πŸ™‚
  • ❀ 18
  • πŸ‘ 1
  • 😁 1
Post #1281 1.68K
Check List to Become DataScientist
  • ❀ 8
Post #1279 2.14K
LLM Pipeline (How Large Language Models Generate Responses)
  • ❀ 5
  • πŸ‘ 4
Post #1278 1.91K
β–ŽUnderstanding Overfitting in Machine Learning

Overfitting is a common challenge in machine learning that can confuse both beginners and experienced practitioners. In this lesson, we'll break down what overfitting is, why it occurs, how to identify it, and strategies to prevent it.

β–Ž1. What is Overfitting?

Overfitting occurs when a machine learning model learns not only the underlying patterns in the training data but also the noise and outliers. As a result, the model performs exceptionally well on the training dataset but poorly on unseen data (test dataset). Essentially, the model becomes too complex and tailored to the training data, losing its ability to generalize.

Key Characteristics of Overfitting:

β€’ High accuracy on the training set.
β€’ Poor accuracy on the validation/test set.
β€’ The model captures noise rather than the actual signal.


β–Ž2. Why Does Overfitting Happen?

Overfitting can happen due to several reasons:

β€’ Complex Models: Using highly complex algorithms (e.g., deep neural networks) with many parameters can lead to overfitting, especially if the dataset is small.
β€’ Insufficient Data: When there isn’t enough data to represent the underlying distribution, models can latch onto random noise.
β€’ Too Many Features: Including too many irrelevant features can confuse the model and lead to overfitting.


β–Ž3. Identifying Overfitting

To identify overfitting, you can use the following techniques:

A. Train/Test Split

Divide your dataset into a training set and a test set (often a 70/30 or 80/20 split). Train your model on the training set and evaluate it on the test set. If you see a significant difference in performance (high training accuracy vs. low test accuracy), your model may be overfitting.

B. Cross-Validation

Use k-fold cross-validation to assess model performance across different subsets of your data. This method provides a more reliable estimate of how well your model will perform on unseen data.

C. Learning Curves

Plot learning curves that show training and validation error as a function of the number of training examples. If the training error continues to decrease while validation error increases, it indicates overfitting.


β–Ž4. Preventing Overfitting

There are several strategies to mitigate overfitting:

A. Simplifying the Model

Choose a simpler model that is less likely to overfit. For example, if you’re using a polynomial regression model, consider reducing the degree of the polynomial.

B. Regularization

Apply regularization techniques like L1 (Lasso) or L2 (Ridge) regularization, which add a penalty for large coefficients in the model. This discourages complexity and helps improve generalization.

C. Pruning (for Decision Trees)

If you’re using decision trees, consider pruning them by removing branches that have little importance. This reduces complexity while retaining essential patterns.

D. Data Augmentation

If you have limited data, consider augmenting your dataset through techniques like rotation, scaling, or flipping images. This increases the diversity of your training data without requiring additional data collection.

E. Early Stopping

In iterative algorithms like gradient descent, monitor validation performance and stop training when performance begins to degrade.
  • ❀ 6
Post #1277 1.71K
RAG Design Patterns
  • ❀ 6
  • πŸ”₯ 2
Post #1275 2.04K
Hey, I guess you already saw it in our other channels, but it's my 31st birthday today πŸŽ‰. so I want to give you a gift 🎁 from my side. You can read more about it here.

Also, I've worked really hard over the last 30 days to prepare the data science course I promised a long time ago. It's finally starting to take shape. I'll keep you informed.

Sincerely, yours
@bigdataspecialist
  • πŸ”₯ 5
  • ❀ 1
Post #1274 1.77K
β–ŽCommon Agentic AI Terms

1. Agentic AI: AI systems designed to act autonomously, perceive their environment, make decisions, and take actions to achieve specific goals with minimal human intervention.

2. Autonomous Agent: An AI that can operate independently, make choices, and execute tasks without direct human command for each step.

3. Perception: The ability of an agent to interpret sensory input from its environment (e.g., text from a user, data from a system, visual information).

4. Action: The output or execution performed by an agent based on its perception and decision-making process (e.g., writing text, calling an API, performing a calculation).

5. Goal-Oriented: Agents designed with specific objectives or tasks they are programmed to achieve.

6. Planning: The process by which an agent determines a sequence of actions to achieve its goal, often involving breaking down complex tasks into smaller sub-tasks.

7. Reasoning: The cognitive process an agent uses to process information, draw inferences, and make logical deductions to inform its actions.

8. Memory: The agent's ability to store and recall information from past perceptions, actions, or conversations to inform future decisions.

9. Tools: External functions, APIs, or services that an agent can leverage to perform actions beyond its core capabilities (e.g., a calculator, a web search API, a database query tool).

10. Tool Use: The capability of an agent to identify, select, and invoke appropriate tools to gather information or execute tasks required to achieve its goals.

11. ReAct (Reasoning and Acting): A framework that combines thought processes (reasoning) with actions, allowing agents to iteratively plan, act, and observe the environment's response.

12. Self-Reflection / Self-Correction: The agent's ability to evaluate its own past actions and reasoning, identify errors or suboptimal steps, and adjust its strategy.

13. Multi-Agent Systems: Systems composed of multiple AI agents that can interact with each other to collaborate, compete, or achieve complex, distributed goals.

14. Task Decomposition: The process of breaking down a large, complex goal into a series of smaller, manageable sub-tasks that an agent can execute sequentially or in parallel.

15. State Management: Keeping track of the current situation, context, and progress of an agent throughout its execution of a task or interaction.

16. Goal Setting: The ability of an agent to define, refine, or adapt its own goals based on context or external feedback.

17. Environment Interaction: The agent's ability to perceive changes in its operating environment and react accordingly.

18. Human-in-the-Loop (HITL): A system design where human feedback or intervention is incorporated at specific points in the agent's decision-making or execution process.

19. Prompt Chaining: A technique where the output of one prompt or agent's action becomes the input for the next, creating a workflow of sequential tasks.

20. Agent Orchestration: The management and coordination of multiple agents or multiple steps within a single agent's workflow to ensure tasks are completed effectively and efficiently.
  • ❀ 5
  • πŸ‘ 2
Post #1273 1.6K
6 Data Warehouse Design Patterns
  • ❀ 9
Post #1272 1.81K

Forwarded from Cool GitHub repositories

Summer2026-Internships

Use this repo to share and keep track of Summer 2026 tech internships across software engineering, data science, quant, hardware engineering, AI/ML and more.
The list is updated and maintained daily!

Creator: SimplifyJobs
Stars ⭐️: 45,169
Forked by: 3,186

Github Repo:
https://github.com/SimplifyJobs/Summer2026-Internships

βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–
Join @github_repositories_bds for more cool repositories. This channel belongs to @bigdataspecialist group
GitHub GitHub - SimplifyJobs/Summer2027-Internships: Summer 2027 software engineering, data science, AI, quant, product management, and… Summer 2027 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC. - SimplifyJobs/Summer2027-Internships
  • πŸ”₯ 3
  • ❀ 2
Post #1271 1.65K
βœ… 8-Week Beginner Roadmap to Learn Data Analysis πŸ“Š

πŸ—“ Week 1: Excel & Data Basics 
Goal: Master data organization and analysis basics 
Topics: Excel formulas, functions, PivotTables, data cleaning 
Tools: Microsoft Excel, Google Sheets 
Mini Project: Analyze sales or survey data with PivotTables

πŸ—“ Week 2: SQL Fundamentals 
Goal: Learn to query databases efficiently 
Topics: SELECT, WHERE, JOIN, GROUP BY, subqueries 
Tools: MySQL, PostgreSQL, SQLite 
Mini Project: Query sample customer or sales database

πŸ—“ Week 3: Data Visualization Basics 
Goal: Create meaningful charts and graphs 
Topics: Bar charts, line charts, scatter plots, dashboards 
Tools: Tableau, Power BI, Excel charts 
Mini Project: Build dashboard to analyze sales trends

πŸ—“ Week 4: Data Cleaning & Preparation 
Goal: Handle messy data for analysis 
Topics: Handling missing values, duplicates, data types 
Tools: Excel, Python (Pandas) basics 
Mini Project: Clean and prepare real-world dataset for analysis

πŸ—“ Week 5: Statistics for Data Analysis 
Goal: Understand key statistical concepts 
Topics: Descriptive stats, distributions, correlation, hypothesis testing 
Tools: Excel, Python (SciPy, NumPy) 
Mini Project: Analyze survey data & draw insights

πŸ—“ Week 6: Advanced SQL & Database Concepts 
Goal: Optimize queries & explore database design basics 
Topics: Window functions, indexes, normalization 
Tools: SQL Server, MySQL 
Mini Project: Complex query for sales and customer analysis

πŸ—“ Week 7: Automating Analysis with Python 
Goal: Use Python for repetitive data tasks 
Topics: Pandas automation, data aggregation, visualization scripting 
Tools: Jupyter Notebook, Pandas, Matplotlib 
Mini Project: Automate monthly sales report generation

πŸ—“ Week 8: Capstone Project + Reporting 
Goal: End-to-end analysis and presentation 
Project Ideas: Customer segmentation, sales forecasting, churn analysis 
Tools: Tableau/Power BI for visualization + Python/SQL for backend 
Bonus: Present findings in a polished report or dashboard

⦁  Practice querying and analysis on public datasets (Kaggle, data.gov)
⦁  Join data challenges and community projects
  • ❀ 7
Post #1270 1.53K
Explanatory Data Analysis Process
  • ❀ 8
Post #1269 1.82K
πŸ“Š Data Analytics Basics Cheatsheet

1. What is Data Analytics?
Analyzing raw data to find patterns, trends, and insights to support decision-making.

2. Types of Data Analytics:
⦁ Descriptive: What happened?
⦁ Diagnostic: Why did it happen?
⦁ Predictive: What might happen next?
⦁ Prescriptive: What should be done?

3. Key Tools & Languages:
⦁ Excel – Quick analysis & charts
⦁ SQL – Query and manage databases
⦁ Python (Pandas, NumPy, Matplotlib)
⦁ Power BI / Tableau – Dashboards & visualization

4. Data Cleaning Basics:
⦁ Handle missing values
⦁ Remove duplicates
⦁ Convert data types
⦁ Standardize formats

5. Exploratory Data Analysis (EDA):
⦁ Summary stats (mean, median, mode)
⦁ Data distribution
⦁ Correlation matrix
⦁ Visual tools: bar charts, boxplots, scatter plots

6. Data Visualization:
⦁ Use charts to simplify insights
⦁ Choose chart types based on data (line for trends, bar for comparisons, pie for proportions)

7. SQL Essentials:
⦁ SELECT, WHERE, JOIN, GROUP BY, HAVING, ORDER BY
⦁ Aggregate functions: COUNT, SUM, AVG, MAX, MIN

8. Python for Analysis:
⦁ Pandas for dataframes
⦁ Matplotlib/Seaborn for plotting
⦁ Scikit-learn for basic ML models

9. Metrics to Know:
⦁ Growth %, Conversion rate, Retention rate
⦁ KPIs specific to domain (finance, marketing, etc.)

10. Real-World Use Cases:
⦁ Customer segmentation
⦁ Sales trend analysis
⦁ A/B testing
⦁ Forecasting demand
  • ❀ 4
  • πŸ”₯ 1
Older posts β†’
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook β†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 β†’