TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1949 · Back to latest

Older Posts 20 shown
Post #1948 228
An interview with a 23-year-old Gabriel Pettersson, OpenAI employee who learned DL without attending university.

An interesting story that makes you think about education and career.

Dropped out of school in a remote Swedish town, Didn't attend university, but is currently working as a Research Scientist at OpenAI in Sora team.

🟢 How he landed in silicon valley?
We live in a time when the monopoly of universities on fundamental knowledge has been shaken.

Traditional education is a "bottom-up" approach. Want to do machine learning? First, learn linear algebra, then calculus, then topology. It's a long process, and often motivation and understanding of why you need it right now are lost.

Companies, which also don't want to wait, are adding fuel to the demotivation fire. Palantir, for example, is already hiring high school students, bypassing universities. And Gabriel's story is a telling example of this trend.

He didn't follow the classic path of "school - bachelor's degree - master's degree". Instead, he used ChatGPT as a personal mentor. And it's not about asking the chatbot to "write the code for me". Gabriel used a method he calls "recursive gap filling".

The essence is to go "top-down". He takes a complex project: for example, he wants to understand how diffusion models work. He asks ChatGPT to write the code. Naturally, at first, he doesn't understand anything.

And here he starts asking questions about each incomprehensible module. "What does this block do?". Let's say it's a ResNet block. He asks: "Why does this help the model learn?". And he digs deeper. If an unfamiliar concept pops up - he asks to explain the mathematical basis underlying it.

This is recursion: layer by layer, until all knowledge gaps are filled. He doesn't learn mathematics for the sake of it, he learns the mathematics he needs right now to work on the code.

How did a foreigner without a diploma get a visa to the USA and a job in Silicon Valley?

For the O1 talent visa, he used his reputation on Stack Overflow and recommendations, which were viewed by millions of people, as proof of his contribution to the industry.

Gabriel advises: forget about HR. Resumes and diplomas don't matter if you can show results. His strategy is MVP or a product demo and writing directly to the company's top management with an offer of free work for a week. This removes the risks for employer and gives you a chance to show yourself.

His main message: if you're ready to actively ask questions and aren't afraid to look stupid in front of AI while learning the basics, you're already in the top 1%. Because most people just go with the flow.


YouTube • #AI #ML #Interview #OpenAI

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1947 211
Instability of training in complex architectures by DeepSeek

Dedicated to one of the most painful problems of modern neural networks.

🟢 What solution proposed?
An approach called mHC (Manifold-Constrained Hyper-Connections).

The idea is that the researchers took a powerful but unstable architecture of Hyper-Connections and imposed restrictions on internal connections.

1️⃣ Projection onto a manifold
Instead of leaving Hyper-Connections free, mHC imposes a restriction on them, they are projected onto a special manifold (matrices with special properties).
This restores identity-mapping, thanks to which the signal remains stable even after tens or hundreds of layers.

2️⃣ Stability & scalability
Thanks to this restriction, the network no longer "explodes" or "attenuates" the signal during deep learning, and it can be effectively used in large models without degrading quality and without complex tricks.

3️⃣ Infrastructural optimizations
The authors also added engineering improvements:
- kernel fusion
- reducing memory overhead
- mixed-precision effects
This makes mHC fast and effective in real tasks even during large-scale training.

The result is impressive:

• training becomes more stable on large scales
• models scale better
• productivity increases
• memory consumption decreases
• mHC outperforms classic Hyper-Connections

DeepSeek shows that the path to the future is not only large models, but also architectures that are stable from within.

Article • #AI #DeepSeek #MachineLearning #NeuralNetworks #Research

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1946 192
GPU Glossary: A comprehensive database on GPUs.

Modal Labs compiled a detailed glossary to solve the problem they themselves encountered when working with graphics processors in the Modal service.

The documentation is fragmented and it's often very difficult to compare concepts at different levels of the stack.

Modal Labs (the Modal brand) read the PDF documentation from NVIDIA, delved into thematic Discord communities, and even bought paper textbooks to compile a knowledge base that covers the entire stack in one place:

☞ CUDA cores, SM, tensor cores, warp schedulers;
☞ Streams, PTX, memory hierarchy;
☞ Roofline, divergence;
☞ Nvcc, nvidia-smi, cuBLAS, Nsight, libcuda.


In the guide, all pages are linked together, so you can go to the section on Warp Scheduler to better understand the streams you read about in the article on the CUDA programming model.

The project itself is open and available on Source and GitHub

#AI #ML #GPU #Glossary #Modal

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1945 377
Free course on autonomous AI agents

Teaches not just text generation, but the creation of systems that understand the task, plan steps, and execute actions.

What's inside?
- how AI agents are structured and how they differ from regular LLMs
- tools and functions that the agent controls
- planning and reasoning
- memory and context in agents
- RAG and agent architectures
- multi-agent systems
- practical cases and production patterns

Who it's suitable for:
- developers who want to build autonomous AI systems
- product managers and analysts who need to understand the architecture
- anyone who wants to quickly get started with agentic AI

Why it's useful:
- agents can make decisions, call APIs, collect data, and automate complex tasks
- the course is offered for free, although it used to be paid


GitHub

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1944 378
Vision-First Agentic Document AI

Turn thousands of PDFs into data ready for LLM.

🟢 What LandingAI Done?

LandingAI introduced Agentic Document Extraction (ADE) DPT-2 Mini - a lightweight version of Document Pretrained Transformer 2, specifically for streamlined document processing.

Ideal for "clean" digital PDFs, where visual context is still crucial for accurate extraction.

Suitable for:
• invoices
• contracts
• letters
• memos
• any neatly formatted PDFs

Key features:

• Structured extraction from digital documents
• Precise understanding of simple PDF layouts
• Support for various block types: paragraphs, images, logos, cards, etc.
• Reliable transcription of English text
• Scale optimization - fast, stable, cost-effective


DPT-2 Mini is focused on speed, reliability, and low cost - when documents are simple and mass-scale, clean structured extraction is needed.

GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1943 171
META proposed new way of training agents

Modern AI is still directly dependent on human labeling and human data in general. And there are a lot of problems with this: it's expensive, time-consuming, "data runs out", etc.

🟢 What META done?
They are also convinced that this is essentially a hard ceiling on the path to AGI: if you only train agents on human traces, then the learning boils down to refining human experience. So can we be 100% sure that such systems can learn something outside the distribution and become smarter than us? This is especially true for areas such as coding, which will be discussed further.

The researchers proposed Self-Play SWE-RL - a way to train agents so that they can self-improve on their own data.

Self-Play SWE-RL consists of two entities: Bug-injector and Bug-solver. The system receives a repository of code, and the Bug-injector studies it, breaks the code, and weakens the tests so that the bug can hide.

The task of Bug-solver is obvious: to fix the code, without issue-text, without hints, without ready-made test runners. And if he breaks something in the process, this case also becomes part of the dataset and expands the sample.

It's important to understand that these are not just synthetic bugs. Here, the same policy breaks and fixes the code (that is, these are just different roles of one agent). In this sense, the approach somewhat resembles GAN: the solver learns at the expense of the injector becoming smarter, and vice versa.

The results are as follows:

- Code World Model (CWM) on 32B, which has already passed the sft stage and was trained in this way, achieved +10.4% on SWE-bench Verified and +7.8% on SWE-bench Pro

- Compared to conventional RL, this approach gives +2.4% on SWE-bench Verified and +3.6% on SWE-bench Pro


Not a breakthrough, but few pipelines today give such significant increases, so it's quite interesting (but code wasn't provided).

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1942 182
Matrix Exponential Attention (MEA)

An experimental attention mechanism for transformers

MEA offers an alternative to classic softmax-attention. Instead of normalization via softmax, a matrix exponential is used, which allows modeling more complex, high-order interactions between tokens.

🟢 How it works?
IDEA:
Attention is formulated as exp(QKᵀ), and the calculation of the exponential is approximated by a truncated series. This makes it possible to calculate attention linearly along the length of the sequence, without creating huge n×n matrices.

What does this provide
- More expressive attention compared to softmax
- Higher-order interactions between tokens
- Linear complexity in memory and time
- Suitable for long contexts and research architectures

The project is at the intersection of Linear Attention and Higher-order Attention and is of a research nature. This is not a ready-made replacement for standard attention, but an attempt to expand its mathematical form.


For ML researchers and engineers who are studying new forms of attention, alternatives to softmax, and architectures for long sequences.

GitHub Not for production yet

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1941 191
PyTorch is pushing the boundaries of ML

Neural Operator officially becomes part of the PyTorch ecosystem - Neural Operators have officially joined the ecosystem.

🟢 What and Why?
Neural Operators are a class of models that learn not to approximate data, but to approximate the operators themselves. Simply put, they learn to solve entire classes of problems, not individual examples.

Why is this needed:
- Solving differential equations
- Physical modeling
- Climate and weather
- CFD, materials, biology
- Scientific and engineering simulations

Unlike conventional neural networks:
- Neural Operators generalize to different grid resolutions
- Work with continuous functions
- Are better suited for tasks where data describe physical processes

What does integration into PyTorch bring:
- A single standard and API
- Compatibility with autograd, GPU, and distributed training
- Easier to implement in real ML and scientific pipelines
- Fewer barriers between research and production


PyTorch is increasingly becoming not just a framework for DL, but a basic platform for scientific computing and physically meaningful AI.

ML and scientific computing continue to converge - and this is one of the strongest signals in recent times.

Source

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1940 207
Two powerful open-source tools to master Local AI efficiently

1️⃣ LEANN: Extreme Compression for RAG

This open-source repo compresses 60 million text chunks from approximately 201 GB to about 6 GB 🤯

That's about 97% less, while the quality of the retrieval remains very close to standard setups.

• No cloud
• No GPU
• Runs locally on a regular laptop
• Full privacy
• 100% open source

LEANN achieves this by not storing embeddings permanently.

Instead, it uses a compact graph and recalculates embeddings only when they are actually needed.
GitHub

2️⃣ Transformer Lab: All-in-One set of tools for working with LLMs.

☞ Allows you to train, fine-tune, and communicate with any LLM locally.

☞ One-click model loading,
Simple drag-and-drop interface for RAG.

☞ Completely open sourced.

GitHub

•••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1939 201
Why Microsoft bought LinkedIn for $26 billion? for users?

No they needed Economic Graph. Google relies on Knowledge Graph for search. Amazon controls retail thanks to Product Graph.

World's most powerful companies aren't just looking for data. They're connecting it and letting systems build on a common semantic layer.

And 99% of AI stacks still perceive memory as a bunch of embeddings in a vector database. That's not understanding, it's an approximate match.

In building agents that read but don't understand. Cognee, brings Big Tech-level semantic infrastructure to open-source stack.

So AI can have memory, a semantic layer, and proper context out of the box.


🟢 Why Cognee Works?
1. The advantage of "cognify"
The usual scheme is ETL (Extract, Transform, Load). Cognee is ECL (Extract, Cognify, Load). It doesn't just dump text into a database, but calculates embeddings, builds connections between entities, and stores them as a semantic data layer. It turns unstructured chaos into something resembling a brain.

2. Understanding time
This is rare. Most RAGs are static snapshots. Cognee understands data dynamics. If a project's status changed today, it remembers the history, not just the latest value.

3. It actually learns
Built-in feedback mechanisms allow the graph to improve over time. It doesn't just give answers, it becomes more accurate.

Big Tech has poured billions into this. Here, you can implement the first memory with just two-three lines of code.

Vectors find similarities.
Graphs find meaning.
Put them together and start building brains.


GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1938 179
Agent Skills for Context Engineering - teaching agents to "think contextually"

This repository shows how to enhance LLM agents so that they better understand the task, the dialogue history, and the conditions, rather than just generating responses.

Useful for:
• skills in managing long contexts
• neat structuring of data and instructions
• templates for searching, filtering, and making decisions
• examples of real scenarios (chats, memory tasks, integrations)

Often, agents lose important details, confuse steps, and "forget" the goal. This library teaches them to keep the context under control and act more consistently.


Github

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1937 201
Scaling ML Systems from JAX-ML

Scaling Book project is a freely available interactive online resource dedicated to scaling machine learning.

💡 What's inside?

Covers key methods, practices, and architectural approaches that help build scalable, high-performance ML systems.

— Basics of model scaling and training
— Data parallelism, parameter parallelism, and mixed strategies
— Distributed training technologies (TPUs/GPUs)
— Computation and memory optimization
— Practical examples on JAX and other stack tools
— Schemes, codes, and visualizations for specific training patterns

📍 Why it's useful:
— Suitable for both experienced ML engineers and those who want to move from prototypes to industrial ML systems
— Combines the theory and practice of distributed training
— Discusses the real limitations of architectures and ways to address them
— Shows how to think systematically about scaling, rather than just copying hacks


Read Here

🤖 Data Science, ML & Big Data with @DataXplore
Post #1936 220
What if the nested definitions of StructType could be replaced with a single line?

When parsing nested JSON in PySpark, you usually have to describe StructType within StructType within StructType. In the end, you get cumbersome, inflexible code that easily breaks down with any changes to the JSON structure.

In PySpark 4.0, the Variant type was introduced, which allows you to completely avoid describing the schema. Simply use parse_json() to load the data and variant_get() to extract values via JSONPath.

Key advantages:
• no need to describe the schema in advance
• any depth of nesting through simple syntax $.path
• changes to the schema don't break the code
• you extract only the necessary fields and only when they are really needed

Update your pipelines to PySpark 4.0:
pip install pyspark>=4.0


Article about PySpark 4.0, Run the code]

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1934 869
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Live stream scheduled for Jan 3 at 16:30
Launching a weekly Saturday series with industry pros.

First up: Q/A with AI Engineer.

📅 Time: TODAY, Jan 3, 2026 | 4:30 PM UTC (10 PM IST)
🎧 Listen Only: Join Livestream

🎤 Want to Speak or Ask?
Comment for Speaker Link

#DataXplore #AI #ML #VoiceChat

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1933
Live stream scheduled for Jan 3 at 16:30
Post #1931 249
Train FunctionGemma on your own.

LM Studio in collaboration with Unsloth have published a detailed tutorial on fine-tuning the recently released Google model: FunctionGemma.

A Reduced version of Gemma (only 270 parameters) for agent scripts and working as an application backend, which can be run on almost any device.

The guide consists of a detailed description of the entire process from training the model to calling tools to converting it to GGUF format and subsequently running it in LM Studio.

The tutorial is suitable for local training (Unsloth works on NVIDIA, AMD, and Intel), but there is also a ready-made Collab Notebook for training in the cloud.


⚠️ FunctionGemma is not intended for use as a direct dialogue model.

#AI #ML #LLM #Tutorial #Unsloth #LMStudio

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1930 211
Limited by GPU and don't want to set up a local cluster for RL?

OpenTinker an open-source RL-as-a-Service is your solution:

You design agent locally and training and inference can easily be offloaded to remote GPUs. No hassle with infrastructure, no rigid coupling of agent logic and execution.

🟢 Why it matters & How?
- you can prototype RL tasks locally without worrying about hardware
- all the heavy lifting - training and inference - is done on cloud GPUs
- supports single-turn and multi-turn tasks
- the trained model can be immediately deployed for inference, without additional code

How it works?

- a disaggregated architecture
- a lightweight client runs locally
- experiments are sent to a cloud scheduler
- the scheduler matches with available GPUs and orchestrates tasks based on resources
- the task is launched remotely, and metrics are streamed in real time to the dashboard

API for developers:

- wrap the environment, reward, and policy once
- OpenTinker handles data loading, rollouts, training, and inference itself

Familiar interfaces:

- for the environment: env.reset() and env.step()
- for training: high-level fit() - a complete end-to-end training loop
- under the hood, fit() is composed of train_step(), validate(), and save_checkpoint()
- want it fast - use fit()
- need control - customize the steps manually

The agent's runtime looks like a state machine:

PENDING - preprocessing and tokenization

GENERATING - the model generates a response

INTERACTING - the agent acts in the environment

TERMINATED - the task is completed

Pure, scalable agent-based RL.
Repository

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with
@DataXplore
Post #1929 534
Boost Your Knowledge

Ahead of New Year holidays, a free-gift of educational materials on the main areas of AI for you.

👨‍🎓 Courses from HF Learn + IBM on AI

☞ LLM Course - introduces you to large language models and natural language processing using libraries from the HF ecosystem: Transformers, Datasets, Tokenizers, and Accelerate.

☞ Robotics Course - takes you from classical robotics to modern approaches based on ML.

☞ Model Context Protocol Course - a course created in partnership with Anthropic, which teaches you to understand, use, and create applications using MCP.

☞ Smol-course - the most comprehensive (and shortest) track on fine-tuning language models.

☞ AI Agents Course - teaches you to understand and use the most topical topic of today: creating and applying AI agents.

☞ Deep RL Course - a course on the most interesting topic in the field of AI: deep reinforcement learning.

☞ Computer Vision Course - a detailed analysis of computer vision, created by the HF community, consisting of theory, practical exercises, and fascinating tasks.

☞ Audio Course - teaches you to use Transformers for audio processing. You will get an idea of the specifics of working with audio data, study various Transformers architectures, and train your own models.

☞ ML for Games Course - learn how to integrate AI models into game development processes and create unique gaming experiences.

☞ Diffusion Course - a full-scale source of knowledge and skills on diffusion. Theory and practice: from studying the Diffusers library to creating data processing pipelines.

☞ ML for 3D Course - an author's set of educational materials on the use of machine learning in 3D from Dylan Ebert (IndividualKex) - a 3D graphics developer at HuggingFace.

☞ Free courses and projects on AI, DS and clouds on Cognitive Class from IBM: over 100 materials on modern technologies.


👨‍💻 28 Ready-made AI Projects: Not just code files, end-to-end working apps that can be launched, tested, used in production/portfolio.

Save it for holidays, this year they're long.

#AI #ML #HuggingFace #LearnDays

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1928 176
Gemma Scope 2

This is a new version of "LLM microscope", or more precisely a set of tools (interpretability tools), designed for interpreting the behavior of LLMs. Specifically, from the Gemma 3 family.

🟢 Why it matters?
Scope works on the basis of SAEs - sparse autoencoders. These are models that untangle the activations of LLMs and extract interpretable concepts from them.These are called "features": they can be things from the real world (bridges, cows) or abstractions (lying, responsiveness).

In essence, by analyzing these features, we can see what the model was actually thinking when generating a particular output. For example, it generates seemingly harmless code, but "thinks" about the concept of a "cyberattack". And this tells us something.

SAEs, by the way, were proposed for use by Anthropic in 2023 (here's our analysis of their article that made the approach popular). But it was Google that brought autoencoders to the production level. Now, this is actually the first and only open tool for such detailed interpretation of LLMs.

The first version of Scope was released in 2024. Then it only worked for small models and simple queries. Now, the approach has been scaled up even for a 27B model.

Plus, the tool has now become more versatile. If the original Scope only existed for a limited number of layers, now it's possible to analyze complex dialogue mechanisms in their entirety.

According to the article, this was mainly achieved by adding Skip-transcoders and Cross-layer transcoders to the model. These are modules that help to see the connections between distant layers and facilitate the analysis of distributed computations. And, by the way, SAEs were trained using the matryoshka method, like Gemma 3n (we wrote about this method here).


If you want to try and delve into the thoughts of models: Huggingface, Colab notebook, technical report, documentation

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1927 194
Context as an infrastructure for AI applications

MemoDB created Acontext, an open-source project that solves one of the most painful problems of AI systems: managing context, memory, and state between requests.

🟢 What Acontext does?
- Extracts context from prompts into a separate layer
- Provides structured "memory" instead of chaotic text
- Allows storing, updating, and reusing context between model calls
- Simplifies building stateful AI applications
- Reduces token overage and the cost of inference

The key idea:
context is not a string, but a manageable object.

Why this is important:
- Prompts stop growing uncontrollably
- The model's behavior becomes more stable
- It's easier to debug and scale the system
- It's easier to add new knowledge sources

Acontext is particularly useful for:
- AI agents
- chatbots with memory
- multi-step reasoning
- instrumental LLM pipelines


Aimed at developers who are building:
- LLM applications
- agent systems
- RAG pipelines
- long-running AI processes

Without a context management layer, it will only get worse from here.

Repository

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →