TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
585
Photos
845
Videos
446
Links
675

Showing posts older than #1742 · Back to latest

Older Posts 20 shown
Post #1741 209
New Way To Fight Cancer

Gemma C2S-Scale 27B helped scientists to analyze gene activity in cells and find for substances to enhance immune response against tumors. (Google-Calico)

🔴 What is the challenge?
Many tumors remain "cold". the immune system "does not notice" them. To reverse this, antigen presentation needs to be triggered, but precisely, only where there is already a weak immune signal, not in all cells indiscriminately.

Gemma was able to predict that a combination of the drug silmitasertib (a CK2 inhibitor) and a low dose of interferon increases MHC-I expression. This makes “cold” tumors more visible to the immune system.


🟢 What Laboratory test results confirmed?
- the model's prediction confirmed by the lab test with combined application indeed enhanced antigen activity by about 50%, and this could become the basis for new types of immunotherapy.

The main achievement: AI not only accelerated data analysis but formulated a new scientific hypothesis that was confirmed in real experiments.


Its an example of how large models go beyond text generation and begin to discover new drugs-mechanisms of action.

More details on Github too

#AI #GoogleDeepMind #BioTech

🤖 Data Science, ML & Big Data with @DataXplore
Post #1740 245
Oh, PyTorch 2.9 is out

🟢 What's new?

⁠☞ FlexAttention has become even more widely available. Added support for Intel GPU, plus flash-decoding optimizations on x86 CPU (speeding up transformer inference without special CUDA cores).

⁠☞ In torch.compile, you can now switch behavior on graph breaks, meaning you can choose to either throw an error or continue.

⁠☞ Official binaries now come immediately with versions for ROCm (AMD), XPU (Intel GPU), and CUDA 13. This greatly simplifies GPU startup out of the box outside NVIDIA, especially on ROCm.

⁠☞ Updated stable libtorch ABI. There should now be fewer surprises when building and updating third-party C++/CUDA extensions.

⁠☞ Rolled out symmetric memory. a new memory model designed to make writing multi-GPU kernels and code easier.


By the way, the release was built from 3216 commits by 452 contributors. Open source is power 💪

🤖 Data Science, ML & Big Data with @DataXplore
Post #1739 162
NVFP4 - a new mathod that trains the 12B Mamba Transformer to store numbers in 4 bits instead of 8 or 16, with almost no loss in training quality and without loss of accuracy

🟢 What is the smart block quantization idea?

- All values are divided into blocks of 16 numbers.
- Each block has its own local scale (8 bits).
- The entire tensor gets a global scale (32 bits).

This preserves high precision of local values and does not lose extremely large or small numbers.


🟣 Results:
- Training the 12B Mamba Transformer on 10T tokens in 4 bits showed accuracy comparable to FP8.
- Computations became 2–3 times faster, and memory usage decreased by 50%.
- Accuracy loss does not exceed 1–1.5% by metrics.
- MMLU Pro: 62.58% (NVFP4) versus 62.62% (FP8).
- MBPP+: 55.91% versus 59.11%.
- Gradients use stochastic rounding to avoid error accumulation.
- Compared to MXFP4, NVFP4 requires 36% less data for the same loss level.


At later training stages, switching to BF16 almost eliminates the quality gap.
NVFP4 is already supported in the Transformer Engine and on Blackwell GPU, including all necessary rounding modes.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1738 116
A rather radical cum categorical joint-work article by OpenAI, DeepMind, and Anthropic

🟢 How any existing methods of protecting LLMs from jailbreaks can be broken?

As an example, they take 12 popular protection mechanisms (Spotlighting, PromptGuard, MELON, Circuit Breakers, etc.) and demonstrate that each can be bypassed with 90–100% success. Even if the original articles claim "0% successful attacks."

The whole point is how we measure the quality of algorithms. In most works, the mechanism is naively tested on a fixed set of known jailbreaks, which do not take the protection itself into account. It's like testing an antivirus only on old viruses. Naturally, nothing works that way.

The authors say a different approach is needed. The model should be challenged not by old templates but by a dynamic algorithm that adapts to the attack and can change strategy. This could be:

➖ An RL agent trained on the model's feedback.
➖ Some kind of search-based attack like beam search and genetic algorithms.
➖ If the model is open, you can optimize the gradient at the token level. That is, gradually change 1-2 tokens, observe the effect, and adapt.
➖ Or simply red-teaming with live people if money is no object. This is still the most effective method.

Currently, any of these methods have up to 95% success in hacking the most popular protection systems. It seems like a simple stress test, but no one has passed it. Funny, of course, but a fact. Essentially, this means that models are a new kind of universal virus that we don't know how to catch at all.


Meanwhile, any system map of any startup says: yes, everything is safe, we swear ☕️

🤖 Data Science, ML & Big Data with @DataXplore
Post #1737 188
And so everyday story.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1736 114
Google has introduced Coral NPU - an open platform for creating smart AI devices on edge devices.

🟠 What is it?
This is a full stack for developing local artificial intelligence that works without the cloud and with virtually no latency.

Coral NPU is a new type of neural processor (Neural Processing Unit) designed for smart gadgets, IoT, and wearable devices.

You can train and run models directly on low-power devices - from sensors and drones to mini-robots and cameras. Coral NPU enables this quickly and securely.

🧩 Inside:
- SDK and tools for TensorFlow Lite and ONNX
- Compiler, quantization, and model optimization
- Support for Python, C++, and microcontrollers

🟢 How it works?
1. The model is trained (in TensorFlow / PyTorch).
2. The Coral NPU compiler converts it into optimized code through MLIR → IREE → NPU binary.
3. The code runs directly on the device using:
- RISC-V (manages tasks)
- Vector units (perform parallel operations)
- MAC matrix accelerators (compute neural networks using milliwatts of energy).

The result is AI inference with performance up to 512 billion operations per second, while the device consumes very few resources and does not send data to the cloud.


Edge AI gets its open architecture from Google.

#EdgeAI #GoogleResearch #CoralNPU #RISC_V #AIHardware

🤖 Data Science, ML & Big Data with @DataXplore
Post #1735 188
🤖 Tongyi DeepResearch: A powerful language model for deep search specially designed for deep information-oriented tasks. It

🟢 Features:
- High performance on complex search tasks with 30.5 billion parameters.

- Demonstrates outstanding results on various benchmarks, including Humanity's Last Exam and WebWalkerQA,

- Fully automated data synthesis process and advanced reinforcement learning methods.

- Compatibility with multiple inference paradigms.

- Efficient training using agent interaction data.


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
Post #1734 321
Finally, in Python 3.14 you can disable the GIL

This is a big event because before, even if you wrote multithreaded code, Python still executed only one thread at a time, without any performance gain.

But now Python can truly run your multithreaded code in parallel.

And uv fully supports this!

🤖 Data Science, ML & Big Data with @DataXplore
Post #1733 114
How Mixture-of-Experts (MoE) LLMs can be made truly cheap by tailoring the architecture to the hardware? (analysis)

🔴 What is the problem?
- MoE activates only a subset of experts per token → saving compute.
- But with large batch sizes, communication and memory increase:
- more experts are loaded,
- KV cache grows,
- memory and network become bottlenecks.


🟢 What is The solution ?
expert parallelism
- Experts are spread across many GPUs.
- Each token goes to top-N experts + a shared expert.
- In DeepSeek: 8 experts out of 256 per layer × 58 layers.

To handle communications:
- attention remains data parallel (cache stays on one GPU),
- only small activation vectors are communicated,
- two microbatches: one computes, the other communicates,
- hot experts are duplicated,
- tokens try to keep experts within a single node.

Optimizations
- multi-head latent attention → compresses KV cache to ~70KB instead of hundreds of KB.
- reworking attention math → fewer computations for long contexts.
- prefill and decode are separated, cache yields ~56% hits → less cost.

Economics
- Cost = $/GPU-hour ÷ tokens/hour.
- Cheaper with larger batch sizes, faster interconnects, more GPUs.
- But if the service promises 20 tokens/sec per user → smaller batches, higher price.

Practice
- NVLink clusters scale excellently.
- InfiniBand between DGX is a bottleneck.
- 72 GPUs at batch 64 → billions of tokens per day for about ~$0.40 / 1M tokens.


Inshort, MoEs become cheap with: Large batches, Compressed KV cache, Smart routing, Separation of prefill and decode, Fast interconnects.

This provides flexibility: fast chat sells for more, while bulk generation (synthetic data, fine-tuning) runs almost at cost.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1732 332
Found a repository with a bunch of tutorials on creating AI agents ready for production and with real use cases

All the code is open source and there is an explanation on how to deploy them.

GitHub: agents-towards-production

🤖 Data Science, ML & Big Data with @DataXplore
Post #1731 182
A complete pipeline for training an LLM from scratch (new repository by Sensei Karpathy)

The project has everything you need to build your own ChatGPT clone.

🟢 What it has?

> • tokenizer (written in Rust)
> • pretraining
> • SFT (supervised fine-tuning)
> • RL (reinforcement learning)
> • model evaluation (eval)

A total of 8,000 lines of code, with no unnecessary dependencies - a perfect educational example to understand how large language model training really works.

💡 This is a project from his upcoming new course LLM101n, and a great opportunity to boost your ML skills in practice.

You can rent a GPU in the cloud and run everything yourself - the code is already ready to launch.

If you run training of the nanochat model on a cloud GPU server (for example, 8×H100), then after about 12 hours of training (cost ~ $300–400) the model reaches GPT-2 level quality on test sets (CORE-score).

And if you train for about 40 hours (cost ~ $1000), it solves simple math and coding tasks, achieving:
- 40+ on MMLU
- 70+ on ARC-Easy
- 20+ on GSM8K

🧠 This is free top-level practice from a master that you shouldn’t miss.

GitHub

#LLM #nanochat #MachineLearning #DeepLearning #AI #GPT

🤖 Data Science, ML & Big Data with @DataXplore
Post #1730 215
Choose your hero:

👍 Passive aggression (stack overflow)

👏 Active hallucinations (ChatGPT)

🤖 Data Science, ML & Big Data with @DataXplore
Post #1729 315
Google proposed a memory system that allows AI to learn from its mistakes in real time

The idea is actually simple,
Look, what does a person do if they make a mistake? Right, they remember it and try to do it differently next time. But LLMs can't do that. Yes, we already have global memory in ChatGPT, but from the perspective of thinking patterns, every new model query is still treated as the first.

Google's approach is called ReasoningBank. It's like a memory block that distills strategic knowledge from past actions.

That is: a dialogue with a user happens -> we call a special judge-agent who evaluates how well the task was solved -> we log this experience with notes on what worked best and worst and why. The output is a structured "memory" with fields Title, Description, and Content.


🤖 Data Science, ML & Big Data with @DataXplore
Post #1728 121
Mamba-3 quietly and without announcement was released at ICLR - and this could mark the beginning of the end of the Transformer era.

🟠 Brief excursus of new Mamba-3
☞ Makes models faster, more stable, and more efficient when working with long contexts.

The main idea is not in attention layers, but in state-space models, where the model stores and updates internal state over time.

- Mamba-1 introduced continuous dynamics and selective memory updates - it remembered efficiently without the high cost of attention.
- Mamba-2 showed that state updates and attention are two sides of the same mathematics, which sped up computations on GPUs.
- Mamba-3 brought the concept to maturity: now the internal memory evolves more smoothly and stably by moving from a simple Euler step to trapezoidal integration.

Instead of a simple Euler step as in Mamba-2, Mamba-3 approximates the integral of the state update not only at the right end of the interval but by averaging between the start and end, with a coefficient λ depending on the data. This provides a more accurate (second-order) approximation and makes the state dynamics more expressive.


🟢 What actually changed?
- Memory became "rhythmic": now the model can store repeating and periodic patterns (e.g., language or music structures).

- The new multi-input-multi-output design allows processing multiple streams in parallel — ideal for modern GPUs.

⚙️ What this means in practice:
- Efficient handling of long sequences: documents, genomes, time series.

- Linear runtime and stable latency make it perfect for real-time applications: chatbots, translation, speech.

- Energy efficiency and scalability open the way to on-device AI, where large models run locally without the cloud.


Mamba-3 is not just an accelerated alternative to Transformers combines deep contextual understanding, speed and stability. from server systems to smart devices.

#ssm #mamba3 #llm,#architecture #ai

🤖 Data Science, ML & Big Data with @DataXplore
Post #1727 310
When you started a fight over 200 comments on a Pull Request and ended up being wrong.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1726 133
RND1 - new experimental model with 30 billion parameters built on Sparse Mixture-of-Experts system with 3 billion active parameters.

🟣 How it was made?
☞ First it was transformed from a pre-trained autoregressive model and then trained on 500 billion tokens to completely change behavior of diffusion model.

☞ Regular models (AR, autoregressive) write text word by word, while RND1 generates the entire sentence at once and then refines it step by step, as if "developing" the text from noise.

☞ This is a Diffusion Language Model (DLM), an analogue of diffusion models that create images, but here it "draws" words.

☞ The Radical Numerics team figured out how to turn a ready model into a diffusion one without training from scratch.

☞ They simply changed the attention type and fine-tuned the model on a new task.

☞ This method is called AR-to-Diffusion Conversion (A2D) that is, conversion from an autoregressive model to a diffusion model.


🟠 How it works?
1️⃣ Take a strong GPT-like model.
2️⃣ Change attention mechanism, now the model sees the entire context at once.
3️⃣ Continue training on the diffusion task.
4️⃣ Use different learning rates for different parts of the network so that the model doesn’t forget the old but learns a new way of thinking.

Under the hood

☞ Mixture-of-Experts (MoE) : the model has 30 billion parameters, but only 3 billion are active at a time. This makes it powerful yet efficient.

☞ Continuous fine-tuning : old knowledge is not erased but "embedded" into the new mode.

☞ Huge batches : the model learns on large data batches to stabilize training since it doesn’t process all tokens at once.


🟢 What makes RND1 more interesting?
☞ Parallel generation : text is created faster without step-by-step delay.
☞ Lower costs : only 3 billion active parameters, yet quality comparable to large GPTs.
☞ New architecture : opens the way for hybrid models combining the advantages of AR and DLM.
☞ Fully open code and weights : you can explore, modify, and run it yourself.
☞ The first serious step toward self-improving AI : the model can not only learn but also help design the next version.

This is a truly interesting method; RND1 shows that AI can be not just trained but restructured — changing its very logic of thinking without starting "from scratch."

It seems this could become the foundation for Recursive Self-Improvement (RSI) systems, where AI is capable of creating and improving itself.


Code, Report, Weights

#RND1 #RadicalNumerics #DLM #DiffusionModel #MoE #OpenSource

🤖 Data Science, ML & Big Data with @DataXplore
Post #1724 230
🚀 Your beloved DataXplore is now on Instagram!

Follow to explore not just the data, but the science behind it.
From insights to intuition, Let’s learn how to Run on Fingers together 🧠💻

Content will be same as Telegram..

#DataScience #ML #AI #BigData #LearnWithDataXplore

Link : Instagram
Post #1723 250
🖥 Scientists from Monash University have created a coin-sized microchip that works like the human brain.

💧 At its core is a liquid structure made of a metal-organic framework (MOF).
Microscopic channels inside it allow ions to pass through, like electrical impulses in the brain — this is how the chip processes signals.

The main feature is that it remembers past impulses and changes its behavior based on experience.
In other words, this chip doesn’t just compute, it learns, like a neural network in our brain.

⚡ This could mark the beginning of a new era of computers, smart, adaptive, and “living,” where computing and memory are combined in one device.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1722 259
🧠 LIMIT: A Study of Retrieval Limits Based on Embeddings

The repository contains the LIMIT dataset, created to test embedding models on theoretical principles. The study shows that even modern models cannot retrieve certain documents, highlighting the limitations of the current approach using single-vector embeddings.

🚀Key points:
- Dataset for testing embedding models.
- Includes 50k documents and 1000 queries.
- Highlights theoretical limitations of information retrieval.
- Code for data generation and experiments is available in the repository.


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →