TGViewer
Channel Public Channel
Data science/ML/AI

Data science/ML/AI

@datascience_bds

Data science and machine learning hub

Python, SQL, stats, ML, deep learning, projects, PDFs, roadmaps and AI resources.

For beginners, data scientists and ML engineers
πŸ‘‰ https://rebrand.ly/bigdatachannels

DMCA: @disclosure_bds
Contact: @mldatascientist
Subscribers
14K
Photos
608
Videos
7
Links
341

Showing posts older than #1350 Β· Back to latest

Older Posts 20 shown
Post #1349 1.2K
SQL Roadmap
  • ❀ 5
  • πŸ‘ 4
Post #1348 1.32K
πŸ”„ Why Cross Validation Is Better Than One Train/Test Split

Imagine flipping a coin 10 times. You might get 8 heads.
Does that mean the coin is biased?
Not necessarily.

A single train/test split can also give a misleading performance estimate.

Cross Validation repeats the process multiple times using different splits.
Instead of trusting one lucky result... You measure average performance across several experiments.

It's a much better estimate of how your model will perform on unseen data.
  • ❀ 5
Post #1345 1.42K
πŸ€– RAG: How AI Can Answer Questions Using Your Own Data

Large Language Models are powerful, but they don't automatically know everything inside your private documents, databases, or company knowledge base.

That's where Retrieval-Augmented Generation (RAG) comes in.
πŸ”Ή 1. User asks a question
The system receives the user's query.
πŸ”Ή 2. Relevant information is retrieved
The query is converted into an embedding and compared against stored documents in a vector database.
πŸ”Ή 3. Context is added
The most relevant information is provided to the language model as context.
πŸ”Ή 4. The AI generates an answer
The model uses the retrieved information to produce a more relevant response.

A simple RAG pipeline looks like:
Documents β†’ Chunking β†’ Embeddings β†’ Vector Database β†’ Retrieval β†’ LLM β†’ Answer


#RAG
  • ❀ 4
  • πŸ‘ 1
Post #1343 1.3K

Forwarded from Programming, data science, ML - free courses by Big Data Specialist

πŸ“š What I’m learning for 2027

I’ve been working in software and data science for over 8 years, but lately I’d be lying if I said I wasn’t a little worried about where our jobs are heading. πŸ˜…

The future feels more uncertain than ever, so I’ve been thinking seriously about what’s actually worth learning to stay relevant in 2027 and beyond.

I searched around for resources I’d personally want to invest my time in, and i figured why not sharing with you guys as well. This is my shortlist πŸ‘‡

🧠 1. Let’s Build GPT from Scratch, Andrej Karpathy
Build a GPT yourself and finally understand what’s happening behind the API.
⏱️ ~2h
πŸ”— https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ

πŸ”₯ 2. Neural Networks: Zero to Hero, Andrej Karpathy
A deeper dive into neural networks, backpropagation, language models, GPT and tokenization.
⏱️ ~19h
πŸ”— https://karpathy.ai/zero-to-hero.html

πŸ€– 3. Hugging Face AI Agents Course
Learn how AI agents actually work: tools, actions, reasoning and agentic workflows.
πŸ’° Free
πŸ”— https://huggingface.co/learn/agents-course/unit0/introduction

πŸ— 4. Designing Data-Intensive Applications, Martin Kleppmann
The classic for understanding databases, distributed systems, replication, partitioning, streams and designing systems that scale.
πŸ“– ~600 pages
πŸ”— https://github.com/aasthas2022/SDE-Interview-and-Prep-Roadmap/blob/main/System%20Design/Resources/Designing%20Data%20Intensive%20Applications%20by%20Martin%20Kleppmann.pdf

βš™οΈ 5. Made With ML
The production side of ML: deployment, testing, monitoring, data pipelines and MLOps.
πŸ’° Free
πŸ”— https://madewithml.com/#course

🎯 Why these?
My bet for 2027 is that writing code itself will become easier, while understanding AI + production systems + architecture will become even more valuable.
So that’s what I’m focusing on.

If you know a resource that belongs on this list please share it so everybody can find it valuable.

Hope this helps ❀️
  • ❀ 5
Post #1342 1.13K
Difference Between AI Systems: A Human Analogy
  • ❀ 7
Post #1341 1.14K
✍️ SQL JOIN Explained Visually

#SQL
  • ❀ 4
Post #1339 1.18K
Machine Learning Visualized

This is an interactive curriculum with animations and exercises that show how machine learning algorithms actually work. You can watch gradient descent, decision boundaries, neural networks, clustering, and more evolve step by step. It is great for building intuition instead of treating models as black boxes.

🎬 Free Interactive + Animation Course
⏰ Duration: Self-paced
πŸƒβ€β™‚οΈ Self Paced
πŸ‘¨β€πŸ« Created by: Daniel Sobrado / community project
πŸ”— Link

#MachineLearning #Interactive #Visualization #Course
βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–
πŸ‘‰ Join @bigdataspecialist for more πŸ‘ˆ
Machine Learning Visualized Machine Learning Visualized - Interactive AI and ML Visualizations | Machine Learning Visualized Interactive visualizations and animations for machine learning concepts including transformers, attention mechanisms, neural networks, and more.
  • πŸ”₯ 3
Post #1338 1.2K
5 LLM quantization techniques, clearly explained:

1. RTN: ignores them. Rounds every weight to the nearest grid level with no calibration data. Cheapest option, weakest at low bit widths.

2. GPTQ: repairs after rounding. Quantizes a layer column by column and adjusts the remaining weights to absorb the error before moving on.

3. AWQ: protects before rounding. Finds the ~1% of weight channels that matter most and scales them up so they survive quantization. Everything still ends up in plain INT4.

4. LLM. int8(): isolates at inference. Outlier dimensions run in FP16, the other 99.9% run in INT8, and the results are merged.

5. QAT: solves it during training. The model is fine-tuned with rounding baked into every forward pass, so it adapts to the damage before quantization is actually applied.

All five produce the same artifact, a model at a fraction of its trained precision. They differ only in where the outlier problem gets addressed.

The visual above nicely summarises these techniques.

#LLM
  • ❀ 8
Post #1336 1.42K
πŸ—‚ 10 Websites for Finding Real-World Datasets

Finding good datasets is often harder than building the model. These websites cover almost every domain imaginable.

1. Kaggle Datasets
2. Hugging Face Datasets
3. Google Dataset Search
4. UCI Machine Learning Repository
5. OpenML
6. Our World in Data
7. World Bank Open Data
8. data.gov
9. FiveThirtyEight Data
10. AWS Registry of Open Data

You'll rarely run out of project ideas with these bookmarked.

#Datasets
  • ❀ 3
  • πŸ”₯ 2
Post #1334 1.58K
🎲 What Makes Random Forest "Random"?

A Random Forest isn't just "many decision trees."
Each tree sees a different random sample of the data.

Then... At every split... It only considers a random subset of features.

So instead of producing 100 identical trees... You get 100 different opinions.

The final prediction is the majority vote (classification) or average (regression).

The randomness is exactly what makes the forest stronger.
  • ❀ 6
Post #1332 1.7K
πŸ“¦ Your CSV Might Be Using Twice the Memory It Needs

Open a CSV in Pandas.

Run:
df.info()


You'll often notice many text columns have the type:
object


If a column contains repeated values like:
London
London
London
Paris
Paris
Berlin

convert it to:
category


Instead of storing the full text every time, Pandas stores each unique value once and references it internally.

On large datasets, memory usage can drop dramatically.
  • ❀ 4
  • πŸ‘ 2
Post #1331 1.43K
NumPy Cheat Sheet for Beginners
  • ❀ 8
Post #1330 1.46K
From Zero to Data Scientist

This is a free, open-source curriculum from Microsoft's Azure Cloud Advocates team that breaks data science down into 20 digestible lessons spread across 10 weeks.

πŸ‘‰ Free curriculum with quizzes and assignments
πŸ‘‰ No prior experience needed to start
πŸ‘‰ It's Project-based so you're building a portfolio as you learn
πŸ‘‰ Created by Microsoft experts and students
πŸ‘‰ Has a strong discord community to back you up

Explore it here: https://github.com/microsoft/data-science-for-beginners
GitHub GitHub - microsoft/Data-Science-For-Beginners: 10 Weeks, 20 Lessons, Data Science for All! 10 Weeks, 20 Lessons, Data Science for All! Contribute to microsoft/Data-Science-For-Beginners development by creating an account on GitHub.
  • ❀ 8
Older posts β†’
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook β†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 β†’