TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1886 · Back to latest

Older Posts 20 shown
Post #1885 173
Understand the mathematics underlying diffusion...

🤖 @DataXplore
Post #1884 181
Your container doesn't contain GPU drivers at all then,

🟢 How does PyTorch inside it use the host's GPU?

You need to understand what is happening on the host side
The NVIDIA driver in the kernel provides GPU access through device files: /dev/nvidia0, /dev/nvidiactl, and so on

Any application communicates with the GPU exactly through these device files

PyTorch does not directly access the driver
It works through CUDA Runtime (libcudart.so) which is a high-level API that handles allocations, kernel launches, and synchronization

This runtime library is inside your container

The entire stack looks like this:
PyTorch → CUDA Runtime → CUDA Driver → /dev/nvidia0 → kernel → GPU

Runtime lives in the container
Driver lives on the host


🔵 How do they connect?

Look at the container launch:
Containerd → containerd-shim → OCI-runtime (runc) → container

But if the driver is on the host and the runtime is in the container, how does the application access everything at once?

Answer: OCI hooks

The OCI specification defines hooks (code) that runs at different stages of the container's lifecycle:

prestart/createRuntime
createContainer
startContainer
poststart
poststop

NVIDIA uses these hooks to add GPU support

Before the container starts, the hook does the following:

1. Mounts GPU devices (/dev/nvidia*)
2. Places driver libraries from the host into the container
3. Sets the necessary environment variables
4. Configures device cgroups

Your application has not even started yet

All of this is handled by NVIDIA Container Toolkit. It intercepts container creation and carefully inserts everything needed for GPU operation

Your image remains normal.
GPU capabilities appear at runtime.


••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1883 183
Great resource if you want to understand in-depth

How parallel execution works on GPU.

NVIDIA PTX Documentation reveals the low-level execution model: instruction devices, hierarchy of threads, blocks, warps, registers, and memory types.

This is fundamental material, without which it's difficult to understand why GPU cores behave the way they do, and how to correctly write high-performance code for CUDA.

Link

#nvidia

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1882 216
A new Python library for agent-based data-processing and ETL using AI has been released.

What is important in DocETL?

1. What is DocETL?

It is a tool for creating and running data pipelines, especially well-suited for complex document processing.

It includes:

- interactive UI playground
- Python package for running pipelines in production


2. DocWrangler

DocWrangler helps gradually build the pipeline:

- trying different prompts and viewing results in real-time
- building the pipeline step by step
- exporting the final configuration for production


3. Python package DocETL

It is used for running pipelines in production. In the example, a pipeline is created that analyzes medical transcripts, identifies drug names, normalizes similar names, and generates summaries of side effects and areas of application.


GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1880 201
How to Evaluate Consistency?

Annotation usually goes through strict stages: data collection → guide creation → annotation launch usually with overlap → label aggregation → model training → quality evaluation.

Let's say we want to get data annotations for a sentiment classifier of support conversations. We have obtained labels from three annotators for some number of real dialogues.

🟢 How can we understand how high-quality the annotation result is? The Agreement

In such a situation, we can calculate the agreement coefficient, The Agreement. It shows the proportion of matching labels between annotators.

In our case, the average pairwise agreement is 83.9%, which is quite good. However, there is a catch: this coefficient, like accuracy in the case of binary classification, can be misleading in the case of class imbalance. In our dataset, more than 70% of the labels fall on two classes — "Neutral" and "Confusion." Let's use other statistical coefficients to ensure high consistency:

📌 Cohen's Kappa (Cohen's Kappa) : pairwise coefficient. It evaluates the normalized agreement between two annotators.

How to interpret the results:

<0.6 : poor consistency.
0.6...0.8 : good consistency, can be used in practical tasks.
>0.8 : very high consistency.

📌 Fleiss's Kappa (Fleiss's Kappa) — suitable for evaluating consistency between several (more than two) annotators.

In our case, Cohen's Kappa varies from 0.7 to 0.84, indicating high consistency. For additional verification, we took several hundred random examples from the dataset and manually assigned labels. It turned out that in 44% of cases, our labels and annotator labels differed that is, annotators on average consistently assign incorrect labels. That is why even high consistency scores do not guarantee high-quality data annotation.


•••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1879 185
Multi-Agent Evolve is now fully open-source 🚀

With its codebase, you can take any LLM checkpoint and allow it to self-evolve without external supervision.
This is an experimental system where agents evolve by creating and evaluating their own improvements.

Code & Models

#AI #LLM #MultiAgent #OpenSource #EvolutionaryAI

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1878 341
Bagging vs Boosting in machine learning, a visual explanation

•••••🤖 @DataXplore
Post #1877 205
Goldmine of AI resources

Building AI prototypes locally is fun. experiment → push code → try different models with almost no environment setup.

But when you start making AI for real users, things get more complicated. You need to consider data storage, efficient retrieval, performance, security, and scalable context management.

🟢 What gap AI Resource Hub from MongoDB fills?

It provides a whole ecosystem of guides, demos, and learning tracks designed for developers who want to build production AI applications on a reliable data infrastructure.

Two especially useful resources to start with:

1. Basics of vector search in MongoDB: understand how semantic search really works and build a working search pipeline.

2. Building memory agents with MongoDB, Fireworks AI, and LangChain: train an agent to remember past interactions and pull context directly from your operational data.

What makes this library even more interesting is that the content is not limited to AI only. It covers all the supporting components needed to run AI in production, for example:

↳ Storage architectures for AI applications
↳ High-throughput indexing and retrieval
↳ Caching to speed up inference
↳ Best security practices for AI data pipelines
↳ End-to-end examples with real datasets


All tutorials are focused on building working systems and real AI engineering tasks, not just explaining concepts.

Try Here

•••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1875 196
Full open source access to the GigaChat AI model lineup has been released online

Sber has published the entire stack of models with permission for commercial use.

The flagship is GigaChat 3 Ultra-Preview a 702B-MoEmodel, fully trained from scratch on a corpus of 14 trillion tokens. This is not an adaptation or fine-tuning of foreign weights: the model has its own dataset, its own synthetic pipeline, and a redesigned architecture. On Russian-language and STEM benchmarks, Ultra-Preview confidently outperforms Russian open source counterparts, as well as surpasses DeepSeek V3.1.
Context memory up to 128k tokens.

Also available in open source is the Lightning version a compact 10B-MoE model that competes in inference speed with Qwen3-1.7B and approaches the quality of dense models around 8B. GigaAM-v3 is also open — a set of five models for working with audio. It recognizes speech excellently — showing a −50% WER compared to Whisper-large-v3.

The open GigaChat lineup effectively forms a new open ecosystem for development, generation, and automation and does so as an independent architecture, not a continuation of someone else’s solutions.


•••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1874 217
INTELLECT-3 - a 106B by Prime Intellect

A powerful open Mixture-of-Experts model trained on GLM-4.5 Air Base with two stages: SFT and large-scale RL fine-tuning.

🟢 The Features:

This is the first model of this scale where Asynchronous RL is not an experiment but the foundation of training. As a result, the model demonstrates strong performance in mathematics, coding, and reasoning.

The model's focus is on long action chains and agent tasks, not just text generation.

- The model shows top results for its size in mathematics, coding, and reasoning.
- Training was conducted on 512×H200 for about ~2 months.
- Used own stack: PRIME-RL, Verifiers, Environments Hub, and sandbox infrastructure.
- Everything is open: code, environments, tools.


Technical Report, HF, PRIME-RL, Verifiers, Environments Hub

#AI #intellect3 #Primeintellect #GLM45
🤖 @DataXplore for Data Science, ML, DL, NN & Big Data Analysis
Post #1873 211
Linear Gradient Matching (LGM)

Synthetic dataset can train linear probes on huge vision models better than real images. Demonstrated by MIT researchers

🟢 Why it matters and usefull?

1️⃣ Take a frozen base model (DINO, CLIP, etc.)
2️⃣ Observe the gradients it produces on real images
3️⃣ Generate synthetic images so that the gradients match
4️⃣ Train a linear classifier - and it works better than training ining ng on thng on the original data

Why this is useful?
— works across models (generated for DINO → works great on CLIP)
— especially strong on fine-grained classifications, where micro-details matter
— helps to see what the model is really looking at: serious correlations, similar clusters, embedding space structure


This changes the understanding of data.
Before:
"You need to collect millions of images."
Now:
"You need to correctly generate dozens."

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1871 224
Mathematical roadmap for ML

Understand core math “Under The Hood”. Material is useful for those who want to delve deeper into theory beyond calling .fit() in scikit-learn.

🟢 How algorithms work?

1️⃣ Linear Algebra:: Language for describing data and models (vectors, matrices, tensors).

2️⃣ Calculus: The toolkit for training and optimization (derivatives, gradients).

3️⃣ Probability Theory: The framework for assessing uncertainty.

From understanding how Backpropagation and SGD work, TO the causes of gradient explosion and choosing the loss function.


APPROACH must be Intuition-based NOT memorizing formulas.

Full Guide

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1870 489
Deepnote has released a new open notebook format that solves the old problems of Jupyter.

Instead of noisy JSON, now YAML with proper git diff. Support for Python and SQL in one file, shared project settings, new blocks for charts and SQL, proper teamwork and collaborative editing. Conversion from .ipynb in one step. Everything is open-source. 💃

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1868 209
Miles is an RL framework for training MoE models from LMSYS ORG team focused on enterprise-level applications.

Remember the slime project? a lightweight tool used in many modern post-training pipelines.

Slime proved that a lightweight design works, and Miles takes the next step like large-scale training of MoE architectures and support for heavy industrial workloads.

🟢 What Features+Stability Miles offers?
Miles offers what is called "True On-Policy." Previously, there was often a divergence between training and inference. Now, thanks to an infrastructural approach, LMSYS has achieved zero divergence. This was made possible by using Flash Attention 3, the DeepGEMM library, and kernels from Thinking Machines Lab, working in conjunction with torch.compile.

The second feature is the use of speculative decoding. Usually, in RL, the draft model is frozen, which prevents it from following the target model's policy. LMSYS added online training of the draft model.

Test results are positive: generation speedup of more than 25%, especially in the later stages of training.

STABILITY.

For enterprise, memory is money. Miles include mechanisms that prevent system crashes due to non-critical OOM errors and fix excessive memory consumption in FSDP.

The project promises support for multimodal training, compatibility with SGLang v2 and extended speculative decoding.


Github #AI #ML #RL #Miles #LMSYS

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1867 300
NVIDIA has released quantized version DeepSeek V3.1 FP4 on Hugging Face

Provides significant memory savings and speeds up performance when using TensorRT LLM.

At the same time, the model maintains high-quality text generation.

HuggingFace

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1865 240
ZAYA1: First MoE fully trained on AMD stack

Zyphyra with AMD & IBM, disproved the popular belief that serious neural network training is only possible on chips from one well-known company.

🟢 How experiment practically proved an alternative?

The project setting was truly "red": AMD Instinct GPUs, AMD Pensando network interfaces, and the ROCm software stack.

ZAYA1 turned out to be quite interesting. It has 8.3 billion total parameters, of which only 800 million are active.

Despite its compactness, it performs well in tests. In reasoning, mathematics, and programming, ZAYA1 outperformed Llama-3-8B and OLMoE. And overall, it stands alongside Qwen3-4B and Google's Gemma3-12B.

Training took place on an IBM Cloud cluster, where the model processed 14 trillion tokens. But it’s not just about the hardware; architectural innovations were used in the pipeline:

A new attention mechanism - Compressed Convolutional Attention. It uses convolutions inside the attention block, reducing computational and memory load.

Redesigned MoE router. Instead of the standard linear router, ZAYA1 uses a complex sequence of operations that make the "experts" inside the neural network specialize much better.

Residual Scaling. Trainable scalar gates were added to the residual stream at the outputs of each block, allowing the model to control the degree of forgetting.


To run inference, zaya branch of the transformers fork from the Zyphra repository is required.

Apache 2.0 License, Article, Model, Arxiv

#AI #ML #LLM #MoE #Zyphra

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1864 214
MODEL-CENTRIC vs DATA-CENTRIC approaches to improving ML models.

Why does it make sense to spend Less Time On Modeling & Experiments and More Time On Data Preparation in situations With Low Metrics?

Suppose we have an emotion classifier and want to boost its metrics.

☞ Change the training approach : experiment with architecture, pretraining, optimizers.

☞ Work with the data : check the dataset, review the labeling, find noise and errors.

☞ Or completely abandon the classic ML model and try feeding everything to an LLM, hoping for the model's zero/few-shot capabilities.

Most engineers choose the first two options. But, as practice shows, it is precisely the search for errors in datasets and improving data quality that gives the most noticeable gain.

Surface defect detection task Example

If you improve model or training approach, you'll not notice quality improvement, or it'll be minimal.
If you work with data, you will see a significant quality increase.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1859 207
Reconstruction of Visual Space from Any Views

Depth Anything 3 (DA3) is a model that predicts spatially consistent geometry from arbitrary visual inputs.

🟢 Key Features:
- The DA3 model outperforms previous versions in depth estimation.
- Supports monocular and multi-view depth estimation.
- High-accuracy pose estimation.
- User-friendly interface and export options to various formats.
- Specialized models for metric depth evaluation.
- Uses a simple transformer and a unique depth representation, enabling high performance in depth and pose estimation.


GitHub #python

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1858 204
Microsoft… Google… AWS… Everyone try to solve same problem for AI agents:

🟢 How can we build knowledge graphs that Run Fast Enough for Real-Time LLMs?

FalkorDB is an open-source graph database that solves this problem by rethinking the very principle of How Graphs Work with using sparse matrices and linear algebra instead of classic graph traversal.

1️⃣ Understand Why its so fast?

Traditional graph databases store connections as linked nodes and traverse them one hop at a time.

But there is a PROBLEM: When you query connections, the database goes through nodes and edges, literally following a map. For huge knowledge graphs powering AI agents, this creates a serious bottleneck.


2️⃣ What if you represent the entire graph as a mathematical structure?

Sparse Matrices: A sparse matrix stores only existing connections. No extra space, no unnecessary data.

And here is the breakthrough:

When your graph is represented as a sparse matrix, you can perform queries using linear algebra instead of traversal. Queries turn into mathematical operations, not step-by-step node transitions.

Mathematics is faster than traversal. Much faster.

Plus, sparse matrices allow incredibly efficient memory usage. You store only what exists, so you can keep huge knowledge graphs in memory without burning resources.


3️⃣ Why not just use Vector Search?

Vector search is fast, but it only captures naive similarity. It can find patterns but does not see structure.

Graphs capture subtle relationships between entities. This ensures the context you bring up for the agent is accurate and relevant, not just similar.


4️⃣ What FalkorDB gives?

↳ Ultra-fast multi-tenant graph database
↳ Efficient storage through sparse matrices
↳ Compatibility with OpenCypher (the same query language as Neo4j)
↳ Specifically designed for LLM applications and agent memory
↳ Runs on top of Redis for easy deployment


If you are building AI agents that need access to connected data in real time, definitely worth trying on GitHub.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1857 1.17K
LLM Council : Vibecode project by Andrey karpati

Instead of asking a single LLM a question, you can combine your queries into a "Council of Models."

Looks Like…
1️⃣ Using OpenRouter, the request is sent to several models (currently GPT-5.1, Gemini 3 Pro, Sonnet 4.5, and Grok 4). Each writes its own version of the answer.

2️⃣ Then all models are shown each other's anonymous answers, they check and rank them, leaving their comments.

3️⃣ All this is eventually sent as context to the "LLM chairman," who then compiles them into single Final Answer.


Interestingly, Quite often models willingly choose another LLM's answer as best.

For example, they constantly praise GPT-5 as the best "council member," while Claude is called the worst.

GitHub #AI #ML #LLM

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →