TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1810 · Back to latest

Older Posts 20 shown
Post #1809 231
We have too much infrastructure tied to pandas, So how do you actually get off the pandas needle?

Narwhals is a library that provides a Polars-like API and serves as a compatibility layer between different DataFrame libraries.

🟢 How Narwhals helps?
It allows writing one set of logic that runs natively on input data backend and returns same type of DataFrame that was input.

It doesn't constantly convert tables into one central format, it wraps/translates calls into the native backend API to keep computations "native" (i.e., avoid expensive conversions). This reduces overhead and preserves performance.

- Supports pandas index, even though other backends do not have it at all.

- Has lazy computations for those who love optimization and deferred execution.

As a result, we get the ability to write libraries and utilities without worrying about which tabular backend team uses. This advantage is already used by popular projects such as Plotly, Bokeh and Darts.


🤖 Data Science, ML & Big Data with @DataXplore
Post #1808 259
How to make ultra-large models work on dozens of AWS GPUs simultaneously?

This is impossible : AWS network (EFA) does not support GPUDirect Async, so GPUs on different machines cannot exchange data fast enough.

🟢 How Perplexity find solution?
They built new software that transfers coordination to the CPU, allowing GPUs to still synchronize almost directly.
This makes inference of models with *1 trillion parameters* efficient on regular AWS clusters, not just on specialized supercomputers.

They prepared expert-parallel kernels for fast MoE inference on AWS EFA:
1T MoE works practically without degradation, and the multi-node mode is comparable to or faster than single-node on 671B DeepSeek V3 with medium batch sizes, opening the way to serving Kimi K2.

PROBLEM: EFA does not support GPUDirect Async, and the standard NVSHMEM-proxy provides MoE routing with latencies above 1 ms.

SOLUTION: kernels pack tokens into single RDMA writes directly from the GPU, and a special CPU thread launches the transfer and overlaps it with GEMM computations.
The result is EFA suddenly becomes a viable option for massive MoE inference.


This is solid engineering and a reasonable balance of accuracy and memory for teams needing portability between clouds.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1807 272
Many companies serve LLMs.

Some use ready-made tools that and some find the OpenAI-compatible protocol insufficient.

Today, a key skill for ML infrastructure is building extensible systems rather than tying yourself to a single API.

How Tochka bank built own LLM infrastructure? : Recommend reading the detailed analysis

🤖 Data Science, ML & Big Data with @DataXplore
Post #1806 341
12 Python libraries for free market data that everyone should know:

yfinance : Stock data: history, intraday quotes, fundamentals. Plus FX, crypto, and options. Uses Yahoo Finance, so all data from there is available through yfinance.

pandas-datareader : Used to be part of pandas, now a separate project. Data on stocks, currencies, economic indicators, Fama-French factors, and much more.

IBApi : Official Interactive Brokers API with access to all their data. Replaced IBPy.

Alpha Vantage : Free API with real quotes and popular financial indicators. JSON or CSV format.

Nasdaq Data Link (formerly Quandl) : Millions of financial and economic datasets from hundreds of sources directly in Python.

Twelve Data : Access to 100,000+ tickers for stocks, forex, indices, and fundamental data worldwide.

Polygon.io : Real-time and historical data on stocks, currencies, and cryptocurrencies.

Tradier : Python libraries for working with the Tradier API.

alpaca-py : Anything you want: from streaming market data to developing your own investment apps.

Finnhub : Real-time REST API and websockets for stocks, currencies, and crypto.

marketstack : Intraday and historical data for 30+ years, 170,000+ tickers.

Tiingo : API with end-of-day quotes. Focus on reliability, transparency, and completeness.


🤖 Data Science, ML & Big Data with @DataXplore
Post #1805 533
Create AI Agents from Scratch

A practical guide to building AI agents without using frameworks.

🟢 What you learn here?
- Step-by-step examples of creating AI agents
- Learning the basics of LLMs and their architectures and their interaction with tools
- Applying system prompts and tools
- Developing agents with memory and strategic thinking
- Practical understanding of working without frameworks
- deeper understanding of how modern AI systems operate.


GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Post #1804 326
Introduction to Machine Learning Systems

Created by Harvard professor Vijay Janapa Reddi. This is an open textbook that teaches you how to build real, working AI systems: from edge devices to the cloud.

It takes learning beyond just "training a model" and shows how to make the model actually work - stably, efficiently, and with high performance.

The PDF and online version are available here, the repository is here

🤖 Data Science, ML & Big Data with @DataXplore
Post #1803 314
How to train models to THINK rather than just MEMORIZE?

Why ordinary transformers are almost incapable of multi-digit multiplication and how to fix it?

MIT + Harvard + Google DeepMind Trained 2 small Transformers to perform 4-digit × 4-digit multiplication.

1️⃣ The First used Implicit chain-of-thought (ICoT) method:
the model first sees all the intermediate calculation steps, and then these steps are gradually removed.

In other words, the model is forced to “think internally” rather than rely on visible hints.

Result: 100% accuracy on all examples.


2️⃣ The second used regular training:
input → answer, without intermediate steps.
Result: about 1% correct answers.

Why is that?

- Multi-digit multiplication requires long-range dependencies
- It is necessary to remember and carry over the “sum + carry” between different positions
- The model must store intermediate partial products and return to them later
- A working model forms a “running sum” and carry, like a human
- Inside attention, a structure resembling a small binary tree appears
- Digit representations form a special space (five-pointed prism + Fourier code)

Regular training captures the “edge” digits and gets stuck — it cannot connect the middle.
ICoT provides the correct inductive bias: it forces the model to build an internal algorithm rather than guess a pattern.


AI needs a computational process to do arithmetic and logic, not just more data.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1802 465
Me, who tries each chatbot in turn to fix my bug...
Post #1801 297
ThinkMorph - a new leap in multimodal thinking, trained on 24K high-quality interleaved reasoning traces and can generate progressive text-visual reasoning steps where text and image enhance each other.

🟢 What it provides?
- a sharp increase in tasks with visual context
- gradual multimodal reasoning step by step
- unexpected abilities: adaptive logic and unprecedented visual manipulations

The model not only sees and describes the image - it evolves during reasoning, correcting and supplementing its conclusions with each new text-graphic step.


This is no longer just a VLM - it is a mechanism that learns to think using both image and text simultaneously, enhancing one with the other.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1800 515
Michael Burry legendary investor who predicted the 2008 mortgage crisis bets $1.1 billion on AI bubble bursting.

Put options against two major companies in the AI sector. (right to sell shares at predetermined price)

If market falls, owner of such contracts profits.

Part of the bet may be a hedge, not a pure bet on the AI market crash but he is betting on the collapse of the overheated AI boom.

#investing #finance #AI #stocks #MichaelBurry

🤖 Data Science, ML & Big Data with @DataXplore
Post #1797 189
Cache-to-Cache

A method (by microsoft) is proposed that works for any pair of models, even from different families, different companies and different architectures.

🟢 How models can communicate in their "own language"?
When two agents communicate in a multimodal system, they usually do so via text. This is quite inefficient because each model actually has a Key-Value Cache, an internal attention states that essentially store all the information about the model's thoughts. And if agents could learn to communicate not with tokens but with the KV cache itself, it would be much faster and the information would be more complete.

Thus appears Cache-to-Cache (C2C) which is paradigm of direct exchange of meaning, not words. The source (Sharer) sends its cache, and the receiver (Receiver) embeds this cache into its own space through a neural network projector.

Directly, without a projector, this would not be possible because different models have different hidden spaces. Therefore, the authors trained a Projection module that connects the caches of the Sharer and Receiver into a single embedding understandable to both models. Besides the Projection module, the protocol also includes a weighting module that decides what information from the Sharer is worth transmitting.

THIS GIVES…
1️⃣ Speed, obviously. Compared to Text-to-Text, everything happens 2-3 times faster.
2️⃣ Accuracy improvement : If two models are combined this way and tasked with solving one problem, the metric improves on average by 5% compared to the case when models are combined but communicate via text.


A big practical downside is that the approach is not universal. For each pair of models, you have to train your own "bridge." There are only a few MLP layers, but still. Also, if the models have completely different tokenizers, it’s a hassle and you’ll have to do Token alignment.


By exchanging caches, models really understand each other better than when exchanging tokens. This is a cool result.

GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Post #1795 244
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning.

The model learns not to guess the final answer, but to plan and verify each step of reasoning.

🟢 How it works?
- Expert solutions are cut into small steps : The model learns to think step-by-step, not just copy the solution

- SRL gives a reward for each step in the chain so model takes a step → receives a score of closeness to the expert

- Small models receive a real training signal and also start planning

- Uses text-matcher + a small format penalty
- Updates in GRPO style with dynamic batch selection to avoid empty signals

The model gains Early planning, Correction on the go, Self-checking of the result
- Also answers don't get longer - Quality grows due to thinking, not rambling


SRL looks like a natural bridge between supervised training and classic RL: controlled stability + depth of reasoning.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1794 239
Evaluation enhances the momentum strategy

The model doesn't "predict the market."
It simply *reads the news* and refines the classic factor by adding a filter of the real information background.

🟢 How does this idea work?
Classic momentum buys recent "winners" but doesn't consider what the news says.

This work added a semantic filtering layer: the model reads fresh headlines and gives each company a score between 0 and 1.

Then the portfolio is reshuffled: a higher score means more weight.

Sharpe ratio increases from 0.79 to 1.06, lower volatility and drawdowns, higher return per unit of risk


Momentum + current headlines → smarter, more stable, safer.

Also... why does GPT-5 trade so poorly then? Trading≠momentum

🤖 Data Science, ML & Big Data with @DataXplore
Post #1792 244
Microsoft is back in its style : Building a solution on an AI agent almost never works on the first try.

Agent Lightning solves : Days spent tweaking prompts, adding examples, hoping for improvement.

No system, just constant trial and error.

🟢 How it works?
The agent works as usual with any framework. You just add a simple call to agl.emit() or let the tracer collect data itself.

Agent Lightning collects every prompt, tool call, and reward, saves everything as structured events.

You choose an algorithm (RL, prompt optimization, fine-tuning). It reads events, finds patterns, generates improved prompts or policy weights.

Trainer loads updates back into the agent. The agent gets smarter without rewriting code.

You can optimize each agent in a system of multiple agents.

Suitable for LangChain, AutoGen, CrewAI, OpenAI SDK, or just Python.

It is open-source framework that trains ANY AI agent using reinforcement learning.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1791 251
🌍 Awesome-World-Models is a large curated repository.

It's an approach in AI where the system builds an internal model of world to understand the environment and predict future actions within it.

🟢 What you will find?
- embodied AI and robotics
- autonomous driving
- NLP models with long-term context and planning
- other fields where AI needs to build a representation of the world and act within it


If the topic of world models interests you, this is a great starting point for study.

#worldmodels

🤖 Data Science, ML & Big Data with @DataXplore
Post #1790 236
Glyph: scaling context through visual-text compression

Instead of feeding model a kilometer-long text, Glyph turns it into an image and processes it through a vision-language model.

🟢 How it works?
A LLM-driven genetic algorithm is used to select the best parameters for the visual representation of the text (font, density, layout), balancing compression and accuracy.

This drastically reduces computational costs while preserving the semantic structure of the text.

At the same time, accuracy hardly drops: on long-context tasks, Glyph performs at the level of modern models like Qwen3-8B.

With extreme compression, a VLM with a 128K context can effectively handle tasks equivalent to 1M+ tokens in traditional LLMs.

In fact, long context becomes a multimodal task rather than purely textual.


HF & Repository

#AI #LLM #Multimodal #Research #DeepLearning

🤖 Data Science, ML & Big Data with @DataXplore
Post #1789 304
Learn by Building: 60 GenAI Projects + AI Engineering Hub

1️⃣ 60 ready projects on generative AI
List of 60 projects on GitHub with open code on generative AI from text models to audio and video.

Each project includes a description and a link to the repository. You can choose an idea, run it locally, and build your own AI portfolio.

Github and Even more useful stuff.


2️⃣ Create own Bash agent with NVIDIA Nemotron in 1 hour
NVIDIA showed how to build an AI agent that understands your natural language requests and executes Bash commands itself.
Based on Nemotron Nano 9B v2 model: compact, fast, perfect for local experiments.

The agent can:
- recognize commands in natural language ("create a folder", "show files"),
- turn these commands into working Bash scripts
- ask for confirmation before execution.

Entire code is about ~200 lines of Python, works via FastAPI and LangGraph.
It can be extended for DevOps, Git operations, log analysis, or server management.
👉 Guide


3️⃣ AI Engineering Hub for learning & developing AI-based solutions
- 93+ production-ready projects for any level
- detailed tutorials on LLM, RAG, agents, and much more
- real examples of AI agent applications
- ready-made examples for implementation, adaptation, and scaling in your projects

Grab it on GitHub


#GenAI #AI #Projects

🤖 Data Science, ML & Big Data with @DataXplore
Post #1788 292
Build Agents, Scale Systems, Ship Projects

1️⃣ Fresh series on Python and AI by Microsoft
• Content: Includes 9 lectures supplemented with videos, detailed presentations, and code examples. series - training in AI agent development - accessible, clearly written, even for programming beginners.
• Topics: cover topics such as RAG (Retrieval-Augmented Generation), embeddings, agents, and the MCP protocol.
👉 Course


2️⃣ Ready-made guide to train & host LLM from scratch by HuggingFace
– Architectures, their features, and hyperparameter optimization
– Working with data
– Pretraining and the pitfalls involved
– Post-training: all modern approaches and how to apply them
– Infrastructure, how to build and optimize it properly
Guide with 200+ pages, 7 big chapters, read + lots of diagrams and examples with Simple English.


3️⃣ Database sharding guide from PlanetScale
Learn how to scale databases through sharding - splitting data across servers to increase performance and fault tolerance.

• Sharding is needed when a single database can no longer handle the load.
• There are two popular approaches — range-based and hash-based.
• It is important to choose a stable key (e.g., user_id) and avoid cross-shard queries.
• A proxy layer slightly increases latency but provides scalability.

Excellent material if you want to understand how systems at YouTube scale. And here is a lot of SQL basics Read


#Python #AI #DataScience #ML #freecourses

🤖 Data Science, ML & Big Data with @DataXplore
Post #1787 337
Start Learning AI & ML - From Python To Harvard Level

1️⃣ FREE mini-courses on Python, DS and ML

What's inside:
• Completely and highly practical.
• Python, Pandas, visualization
• Basics of machine learning and feature engineering
• Data preparation and working with models

Practice without unnecessary theory: you learn and immediately apply.
👉 Course


2️⃣ A huge collection of the 17 best GitHub repositories for learning Python
☞ 30-Days-Of-Python
— covers python basics

☞ Python Basics — simple-clear Python basics for beginners

☞ Learn Python — guide with examples and code

☞ Python Guide — best practices, tools, advanced topics

☞ Learn Python 3 — An easy-to-understand guide to Python 3 with practice

☞ Python Programming Exercises — 100+ Python problems

☞ Coding Problems — algorithmic problems, perfect for interview prep

☞ Project-Based-Learning — learn Python through real projects

☞ Projects — ideas for practical skill improvement

☞ 100-Days-Of-ML-Code — a step-by-step guide to ML in Python

☞ TheAlgorithms/Python — a huge collection of algorithms in Python

☞ Amazing-Python-Scripts — useful scripts from automation to advanced utilities

☞ Geekcomputers/Python — a collection of practical scripts: networking, files, automation

☞ Materials — code, exercises, projects from Real Python

☞ Awesome Python — a top list of best frameworks and libraries

☞ 30-Seconds-of-Python — short snippets for quick solutions

☞ Python Reference — life hacks, tutorials, and useful scripts


3️⃣ Harvard Machine Learning Course
Iconic CS 249 track has been turned into an interactive textbook - arguably one of best starting points for engineers who want to build real ML systems, not just play with models.

• Complete ML foundation: explains fundamentals from scratch, only Python knowledge is required
• System design and data engineering
• Dataset preparation, MLOps, monitoring
• AI deployment in IoT and production

Its a practical course: not about formulas, but about how to implement ML so that it brings business profit.
If you want to understand how models operate in production - an ideal start.
👉 Course


#Python #AI #DataScience #ML #freecourses

🤖 Data Science, ML & Big Data with @DataXplore
Post #1786 251
Air is a Python framework designed with an AI-first focus

It is still in the alpha stage, but it is already clear: this is not just a framework - it is an attempt to rethink how web applications are built in the AI era.

🟢 What makes Air special?
- Compatibility with FastAPI / Starlette: routes, middleware, OpenAPI — all in place.
- Integration with databases through air.ext.sqlmodel (SQLModel / SQLAlchemy).
- Basic authorization ready "out of the box" — OAuth, login via GitHub.
- Approach to interfaces: templates + declarative tags, reactivity without heavy JS — inspired by HTMX.
- Every component and API aims to be understandable, simple, like in Django, but with added AI orientation.

But it is important to remember

Air is currently an experiment.
APIs may change, functionality is not fully implemented.
The authors ask for understanding and participation in the framework's development.


If you are tired of “ordinary” web frameworks and are thinking about how to embed AI into the architecture from the very beginning, Air might be the very start of a new path.

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →