TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1907 · Back to latest

Older Posts 20 shown
Post #1906 223
Almost all complex AI agents can be described through just 4 basic types of adaptation - 2 related to updating the agent itself, the other 2 to updating the tools the agent uses.

🟢 How modern agentic AI systems adapt is proposed?

What is agentic AI:
These are large models that can:
- invoke tools,
- use memory,
- perform tasks in several steps.

What is adaptation:
Any change to the agent or its tools based on feedback, from code verification to human evaluations.

4 types of adaptation:

A1 - Agent Adaptation from Tool Execution
The agent is updated based on what happened during the execution of tools: the code either ran or crashed, the search either found something or not.

A2 — Agent Adaptation from Output Evaluation
The agent is updated based on evaluations of the quality of its final actions: human feedback, automatic checks of answers, the quality of plans.

T1 - Tool Adaptation Independent of Agent
The tools are trained separately, and the agent remains "frozen". For example, a pre-trained retriever or a code searcher.

T2 - Tool Adaptation from Agent Signals
The agent remains fixed, but the tools adapt to its behavior - which documents really helped, which hints improved the task execution.

Why this is important:
- For the first time, the work systematically organizes the methods of adaptation of agentic systems.
- Helps to understand the trade-offs: the cost of training, flexibility, portability, modular updates.
- Shows the history of the development of methods A1, A2 and T2, how they became more complex and what signals they began to use.

The view boils down to two axes:
- you can change the agent,
- you can change the tools,
- and the data and feedback serve as fuel for both strategies.


This taxonomy helps to see the connections between dozens of modern works and to understand where the new generation of agentic architectures is heading. (PDF on GitHub)

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1905 384
AI is already ahead of professional hackers

in the Real-World Penetration experiment intense with 10 professional pentester's + a live university network + ~8,000 real machines + 12 subnets + production systems and real users

The result was unexpected ARTEMIS outperformed 9 out of 10 human experts, even found vulnerabilities that not a single human found.

🟢 Why AI was stronger?

Not a CTF, Not static CVEs, Not a simulation. A real network with real consequences.

• Humans manually selected targets
• ARTEMIS launched sub-agents and attacked multiple hosts in parallel
• Humans lost clues and went down "rabbit holes"
• ARTEMIS maintained perfect memory, TODO lists, and auto-triaging
• Humans couldn't open outdated web interfaces
• ARTEMIS simply ignored the browser and hacked them via curl -k

Moreover, it ARTEMIS showed 9 confirmed vulnerabilities, 82% valid findings, without human supervision or custom exploits and that too In cost of work ~$18 per hour where human pentester costs ~$60 per hour.

What still holds it back:
— GUI-dependent exploits
— a higher percentage of false positives

In everything else, ARTEMIS acted like a fully equipped red-team:
without fatigue, without ego, with infinite patience.


🔴 AI is no longer "helping" pentester's, AI is starting to competing and in some scenarios already winning

Offensive security begins to change forever.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1904 196
Soft Adaptive Policy Optimization aka SAPO

An RL method that solves the key problem of unstable training in LLMs and MoE architectures and offers a more reasonable and gentle approach to controlling the learning process.

🟢 How it Solves the problem?

Reinforcement Learning, RL - is the ingredient that turns a simple large language model into a reasoning assistant. It's RL that teaches AI to solve olympiad math problems, write clean code, and understand the relationship between text and images.

But RL has a downside: catastrophic training instability, especially for gigantic models.

The main technical puzzle is controlling the importance coefficients at the level of each token. In MoE architectures, where different parts of the model are activated for different tasks, these coefficients can "jump" uncontrollably.

Excessively large fluctuations in the coefficients turn clear training signals into interference, destabilizing the entire system.

Until now, the standard tools have been GRPO and GSPO, which used the principle of hard clipping. If the coefficient exceeded the specified limits, the gradient was simply set to zero.

FIRST MINUS: Information loss. Valuable but outlier data were ruthlessly discarded.

SECOND MINUS: Impossible balance. If you make the limits too narrow, you stifle learning. If you make them too wide, parasitic noise creeps in. For capricious MoE architectures, this dilemma is particularly relevant.
SAPO proposes to abandon hard clipping in favor of intelligent smoothing.

Instead of abruptly setting the gradient to zero, SAPO uses a smooth, adaptive function (controlled by temperature) that gently reduces the influence of problematic gradients without completely nullifying them. This creates continuous confidence regions within which the model can learn more flexibly and safely.


Like GSPO, but smarter. If only one token in a long answer was wrong, GSPO punished the entire sequence. SAPO selectively suppresses only the "culprit", preserving useful signals from the rest of the words. This dramatically improves the efficiency of training data sets.

Like GRPO, but smoother. Instead of abruptly disabling the gradient for a bad token, SAPO applies a gradual attenuation. This prevents sudden jumps in learning and ensures a smooth and stable adjustment of the model's policy.

The icing on the cake of the method is an asymmetric temperature design. SAPO processes "good" and "bad" updates differently. For tokens with a negative contribution, a higher temperature is used, forcing their influence to attenuate faster and stronger.

This simple rule reliably suppresses the most dangerous fluctuations, which in practice leads to unprecedented stability of the RL learning process.

CONFIRMED BY TESTS

When training Qwen3-30B-A3B-Base, SAPO not only showed a more stable learning curve, but also achieved higher results on complex mathematical benchmarks AIME25, HMMT25. And he did this without the labor-intensive routing reproduction that competitors needed to work with MoE.

The success was repeated in a large-scale experiment with multimodal Qwen3-VL-30B-A3B, where SAPO consistently outperformed its analogues in mixed tasks on coding, logic, and mathematics.


#AI #ML #LLM #MoE #SAPO #Qwen

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1903 332
This is what the Dataset Cartography looks like.

🤖 @DataXplore
Post #1902 457
In-Context Learning for Agent-Based AI

Here we've gathered a simple, beginner-friendly breakdown of In-Context Learning with Colab examples. The demos include optimization, regression, classification, RL, translation, and a host of other tasks.

Start Learning

•••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1901 326
K-means is one of the most widely used clustering algorithms in Data Science and Machine Learning.

Key part of algorithm is convergence (the process by which cluster centers and point assignments gradually stabilize through repeated updates).

🟢 HOW & WHY convergence happens?

☞ Quickly converges on most datasets, making it efficient for large-scale tasks
☞ Offers a simple and interpretable structure for identifying groups
☞ Scales well on large datasets due to low computational complexity

☞ Results heavily depend on the initial cluster initialization
☞ Can distort data structure if features are improperly scaled
☞ May produce empty or unstable clusters if not properly configured

To ensure stable convergence:
- Use k-means++ for a more informed choice of initial centers
- Apply feature scaling so that variables with large scales do not dominate
- Set appropriate values for iteration limits and convergence thresholds


The image shows the K-means convergence process.

Data points are assigned to the nearest center based on squared distance. Then each center is recalculated as the mean of all points assigned to it.
These steps repeat until the positions of the centers no longer change significantly.


Understanding Convergence helps to obtain reliable and meaningful clustering results. #DataScience #MachineLearning #DS #ML

•••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1900 178
Vision-Language Navigation

The StreamVLN model generates actions based on a continuous video stream in online mode, conducting a multi-turn dialogue.

🟢 What makes StreamVLN interesting?
☞ Based on LLaVA-Video but extended for joint modeling of vision, language, and actions.
☞ Accepts a video stream → responds with actions and replies in real time
☞ Processes long sequences without computational overload
Has two levels of memory:
1) fast dialogue memory — sliding-window KV cache
2) slow long-term memory — token pruning to save resources.


Agent that can watch, understand, and act online while maintaining context without loss of speed.

Repository

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1899 194
Domain-specific model MetalGPT-1 fed metallurgy and mining Data

🟢 What so special?

Trained on Technological protocols, regulations, R&D, construction and design documentation are not texts in the usual ML sense.

These are formalized fragments of the production world: the language of processes, chains, constraints, risks.
By training an LLM on such a corpus, the company is effectively creating a separate “data-reality layer” that universal models simply do not see.


Domain-first LLMs will become infrastructure. Next will be models for chemical engineering, logistics, energy, construction Each industry has its own language, its own dataset, its own reality.

Huggingface #LLM #ML

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1898 189
LLM FINE-TUNING

What's is LLM fine-tuning?

Classical LLM fine-tuning is impractical (billions of parameters, hundreds of gigabytes).

Since such computing resources aren't available to everyone, parameter-efficient finetuning techniques (PEFT) have emerged.

Before diving into each technique, here's some background to help you better understand how they work:

LLM weights are matrices of numbers that are adjusted during fine-tuning.

Most PEFT techniques boil down to finding a low-rank adaptation of these matrices, i.e., a matrix of a smaller dimension that can still represent the information from the original one.

Now that you have a basic understanding of matrix rank, let's dive into the different fine-tuning techniques.


Top 5 Techniques
1) LoRA

Two trainable low-rank matrices A and B are added next to the weight matrices.

Instead of fine-tuning W, the updates to these low-rank matrices are adjusted.

Even for the largest LLMs, LoRA matrices only take up a few megabytes of memory.

2) LoRA-FA

While LoRA greatly reduces the number of trainable parameters, updating the low-rank weights still requires a significant amount of activation memory.

LoRA-FA (FA stands for Frozen-A) freezes matrix A and only updates matrix B.

3) VeRA

In LoRA, the low-rank matrices A and B are unique to each layer.

In VeRA, A and B are frozen, random, and shared across all layers.

Instead, layer-specific scaling VECTORS b and d are trained.

4) Delta-LoRA

Here, too, the W matrix is adjusted, but not in the classical way.

The difference (delta) between the product of matrices A and B at two consecutive training steps is added to W.

5) LoRA+

In regular LoRA, both matrices A and B are updated with the same learning rate.

The authors of LoRA+ found that a higher learning rate for matrix B leads to better convergence


••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1897 191
How to make all dangerous knowledge stored in the model separately from normal?

Selective Gradient Masking

There's no such thing as eliement at the pretraining stage. Everything is added after pretraining. And this is a pretty serious problem.

So the only option that people have come up with is to simply remove the "dangerous knowledge" from the dataset, but this (1) is very expensive and time-consuming, because it requires labeling; (2) cuts off a lot of useful knowledge and the model gets dumber. So it's nonsense.

Anthropic suggests not to touch the data itself, but instead to make all dangerous information flow into a separate piece of parameters, which can then simply be... removed. It works like this:

- For each transformer block, we additionally put on a head of attention, which we mark as "forget" parameters.

- If data marked as "dangerous" comes in, we forcibly zero out all gradients except for "forget". This ensures that all dangerous knowledge flows into a certain place.

- So that the model can work well without these parameters, the activations are zeroed out on a part of the data during direct passage.

As you can see, this is, in fact, the same data filtering. But smart. Firstly, this approach is resistant to labeling noise. Secondly, it's not necessarily necessary to label all data: it turned out that from a certain point onwards, even unlabeled dangerous content of the dataset begins to gravitate more towards the "forget" parameters. This is called the Absorption effect.

At the same time, the model after cutting out this black soul gets dumber less than when cutting out data from the dataset. Still, we act a bit more delicately here. And after that, it behaves as if it had never been shown anything like this, and not as if it had temporarily forgotten about it.


In general, at the level of mechanics and idea - a pretty interesting seed

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1896 203
Process datasets larger than available RAM with DuckDB's auto-spillover

When the data volume exceeds available memory, most tools crash during execution.
Because of this, you have to manually split data into chunks or upgrade hardware just to complete a basic query.

Advantages?
• DuckDB automatically spills intermediate results to temporary files when data exceeds the allocated memory.
• Process datasets larger than RAM without code changes
• Configurable memory limits to avoid system crashes
• Automatic spillover to disk when memory is full
• No manual chunking or batching


Full article and Run example

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1895 208
Advent Calendar on Machine and Deep Learning

Useful For Students to understand the formulas, and For Managers to understand which ML method is necessary for business, For Developers it's a chance to finally understand the theory.

Understand…
🟢 How The ML Engine Works?
Frameworks like scikit-learn have made us lazy. Calling model. fit has become so commonplace that in the era of Gen AI, it seems that training a model is just a matter of parameter selection.

ML engineers juggle with models of increasing complexity, but they are not always able to manually recalculate and explain the results of even the simplest algorithms: linear regression or classifier.

Models have become "black boxes", and this is a huge problem, because knowing what's behind each function is critical for understanding the process.

The cool thing is that all the material is explained in Excel. It sounds crazy, but that's the genius of it. Unlike code, where operations are hidden behind functions, in Excel every formula, every number, every calculation is in plain sight. No "black boxes".

7 articles have already been published:

Day 1 : k-NN Regressor

Day 2 : k-NN Classifier

Day 4 : GNB, LDA, QDA

Day 5 : GMM (Gaussian Mixture Model)

Day 6 : Decision Tree Regressor

Day 7 : Decision Tree Classifier

The cycle will help answer questions that often remain behind the scenes: how to properly handle categorical features, when scaling is not the right solution, and how to measure the importance of features by interpreting them directly with the model, bypassing model-agnostic packages LIME and SHAP.


Must-read for those who want to stop being a library operator. You can monitor New Articles Here, “One Day - One Article”

#AI #ML #DL #Tutorial #Excel

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1894 205
QVAC Fabric LLM is a framework that brings full AI inference and fine-tuning directly to your hardware.

Run and fine-tune modern models like Llama 3 and Gemma 3 on your laptop, a regular GPU, and even on a smartphone.

QVAC Fabric LLM is open source on HuggingFace

No cloud, No compromises,
Decentralized, hyperscalable, antifragile, user-centric AI.

You have full control over your data.
Your device. Your AI.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1893 197
DeepMind Demonstrates AI That Creates a Detailed Map of Earth

AlphaEarth is a model that combines vast amounts of satellite and climate data and transforms them into a precise map of the planet with a resolution of up to 10 meters.

🟢 What it does?
- The model creates a compact 64-dimensional representation for each area of the Earth. This allows for rapid analysis of the territory, observing how it has changed from 2017 to 2024, and comparing regions with each other.
- The system makes the data 16 times more compact and approximately a quarter more accurate than previous approaches.
- It's possible to track deforestation, urban growth, soil conditions, climate impacts, changes in coastlines, and other processes.
- AlphaEarth is already integrated into Google Earth Engine, so it's available to researchers, ecologists, and government agencies.


A tool that helps to see the Earth in dynamic and high-precision detail, enabling a better understanding of the changes taking place.

#DeepMind

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1892 176
Tencent officially introduced HY 2.0 - a major update to its base model.

The model is built on the Mixture of Experts architecture with a total size of 406B parameters and 32B active ones.
The model supports a context of 256K tokens. HY 2.0 shows significant improvements on key benchmarks.

Main achievements of HY 2.0:
🧠 Reasoning: a score of 73.4 on IMO AnswerBench - almost a 20 percent increase, securing the model among the leaders in mathematical and scientific reasoning.
🛠 Coding and Agents: a leap in SWE Bench Verified from 6.0 to 53.0, and Tau2 Bench grew from 17.1 to 72.4.
⚡ Instruction Following: more stable execution of complex instructions and a natural style of responses.

The model is released in two versions:
• HY 2.0 Think - for deep reasoning, code generation, and complex tasks
• HY 2.0 Instruct - for dialogue, creative writing, and multi-turn contextual conversations


Website, API Access, Documentation

#AI #Tencent #Hunyuan #HY2 #LLM #MoE #DeepLearning #AIModels

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1891 409
A Neural Processing Unit (NPU) is a specialized chip optimized for fast parallel computations, primarily matrix and vector operations.

In PROBABILITY theory and STATISTICS, NPUs accelerate tasks such as Monte Carlo simulations and Bayesian inference.

In machine learning, they speed up neural network training and inference with low power consumption.

In REAL LIFE, NPUs enable functions like face unlock, speech recognition, smart cameras, autopilot, and local predictive AI on smartphones, cars, and IoT devices.


••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1890 210
AI has learned to write CUDA kernels more efficiently than NVIDIA engineers.

DeepReinforce research group has developed a fully automatic GPU code generation system for matrix multiplication called CUDA-L2.

🟢 How does it achieve a 10–30% faster over NVIDIA's highly-optimized cuBLAS & cuBLASLt libraries?

Such libraries are manually created by people who use ready-made kernel templates. Autotuners only tweak parameters, such as tile sizes.

But DeepReinforce believes that even critically important and deeply optimized tasks like HGEMM can be improved using an LLM working in conjunction with RL.

In the CUDA-L2 system, the language model literally writes CUDA source code from scratch for each matrix size. It doesn't just change parameters; it can change the code structure, loops, tiling strategy, padding, and even swizzle patterns. Moreover, it chooses the programming style itself—whether raw CUDA, CuTe, CUTLASS, or inline PTX.

The process works like this: an RL loop runs the generated kernels on real hardware, measures speed and correctness, and then updates the LLM. Over time, the model derives its own performance rules instead of relying on human knowledge.

The generator used was the DeepSeek 671B model. It was further trained on a mixture of CUDA kernel arrays and quality code from PyTorch, ATen, CUTLASS libraries, and examples from NVIDIA.

WHAT THIS MEANS?

For pretraining and fine-tuning, most GPU time is spent on HGEMM matrix multiplication operations. If these kernels are sped up by the 10–30% promised by CUDA-L2, the entire training process becomes noticeably cheaper and faster.

Since CUDA-L2 handles about 1000 real matrix sizes, not just a few manually tuned ones, the acceleration works across a wide range of architectures. This means that with the same GPU budget, you can fit more training tokens, more SFT or RLHF runs, etc.


TESTS

HGEMM kernels created by CUDA-L2 are consistently faster than standard libraries.

In the so-called "offline scenario," CUDA-L2 runs about 17–22% faster than torch.matmul, cuBLAS, and cuBLASLt. It even beats cuBLASLt AutoTuning by 11%, which itself uses kernel search.

In the "server" scenario, which simulates real inference with pauses between calls, the difference is even greater: a boost of 24–29% compared to torch.matmul and cuBLAS.


Project is not limited to simple research; in the GitHub repository, optimized 32-bit HGEMM A100 kernels for 1000 configurations.

#AI #ML #CUDA #DeepReinforce

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1889 215
𝗧𝗲𝗰𝗵 𝗜𝗻𝘀𝗶𝗱𝗲 Check out the Solar System simulation made by Claude Code Such "useless" simulations very well exposes how LLMs start having terrible hallucinations due to large number of objects. Prompt: Create as realistic and detailed 3D Solar System simulator in HTML…
How to use Claude Code for end-to-end training of open-source LLMs?

Integrated Hugging Face's capabilities into Claude Code and the agent was able to run a full cycle of end-to-end model training.

- You give the agent the task of retraining the model on a dataset, you can specify your own or let the agent find it himself

- - The full training run is automatically launched on cloud resources
- Progress is displayed in real time via the Trackio dashboard
- Checkpoints are automatically pushed to Hugging Face


You can Try Here in any coding agent, works with Claude but also with Codex, Cursor, Gemini CLI.

#ClaudeForTraining
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1887 443
What to do if the annotator markup does not match the expert markup?

Consistency co-efficients do not always reflect the final quality of the markup and the model. almost half of the examples are marked incorrectly, the annotators agree with each other but not with the experts.

🟢 How can we improve the markup?

CHECK the task formulation and specify a detailed guide with corner cases. You can take a sample, markup it according to the guide, and see where disputes arise, these places need to be clarified.

COLLECT a test dataset with gold markup using an expert. After that, you can select annotators with high performance on the test set or hold a briefing meeting with all annotators to discuss the errors.

SPLIT the work into chunks and add a golden set for validation in each. This will allow you to evaluate the quality of the markup iteratively and monitor how much the annotators match the golden set.

After implementing these steps, our emotion model's weighted F1 score increased from 0.61 to 0.7, and the discrepancy between expert and annotator markup decreased from 44% to 18%. Also, some small problematic classes improved well:

Gratitude — 0.8 → 0.76
Neutral — 0.7 → 0.75
Satisfactory — 0.68 → 0.74
Impatience — 0.57 → 0.53
Disappointment — 0.57 → 0.55
Confusion — 0.46 → 0.7

Important: low consistency does not always mean poor performance by the annotators. The reasons may be:

Unclear task: it may imply some uncertainty. For example, this often occurs when preparing dialog data for LLMs.

Different backgrounds of the annotators: internal AI trainers and external contractors may understand the task differently. This leads to significant differences in evaluations.


Therefore, ML engineers and data scientists must carefully read the data and understand how it is marked.

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1886 202
DeepSeek released not "another model", but a full-fledged top-level system at the level of IMO/IOI/ICPC - at the same time, training and generation are tens of times cheaper than GPT-5 and Gemini 3 Pro.

🟢 All the How's?
• DeepSeek-V3.2-Speciale outperforms Gemini 3.0 Pro in mathematics and code
• The new flagship model combines reasoning + agency
• Architecture MoE from the V3.1 Terminus family, context 128k
• The main innovation - DeepSeek Sparse Attention (DSA), designed for cheap long context

What makes DSA
Normal attention - O(T²), which is costly at 128k tokens.
DSA reduces the cost to O(T·U), where U is only a small number of relevant tokens.

How it works:
1) Lightning Indexer - a lightweight network evaluates the importance of each previous token
2) Fine-grained top-k - the model selects only the most useful tokens and calculates attention on them

How it was trained
We started with a checkpoint of V3.1 (128k) and performed two-stage fine-tuning:
• Stage 1 - dense attention, frozen model, trained only DSA
• Stage 2 - gradual transition to DSA across the entire model


Result: long context has become really cheap, and the quality is higher than previous versions and competitors.

••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →