TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1857 · Back to latest

Older Posts 20 shown
Post #1856 201
A model performs well only on the dataset it was trained on. Once the data source is changed, the quality drops.

This article demonstrates a simple trick: you can train a neural network so that it cannot determine which dataset a sample came from. As a result, it starts to extract more general, universal features that work under any conditions.

The method is very easy, can be added to any neural network with just a few lines of code. But the result is consistent: the model handles new data it hasn't seen before better.

The work stands out pleasantly: clear idea, precise explanation, real results, not just another “+2% on some random metric.”

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1855 558
Kaggle has launched its own official MCP

Now you can connect this MCP to Cursor (or any other agent) and give requests like "Find the best dataset for classifying photos of dogs and cats and process it."

Search and browse competitions/datasets/notebooks, download files, submit entries and even create and run notebooks.

All this without leaving IDE: Start Kaggle MCP server → give it your API keys.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1854 182
How to scale biological models? guide

🟢 What are 3 key ideas?

1️⃣ Using Transformer Engine replaces standard blocks with optimized versions: less memory, faster matrix operations, support for FP8/FP4. This immediately increases training and inference speed.

2️⃣ Scale training to billions of parameters
Through FSDP and hybrid parallelism modes, the model can be distributed across multiple GPUs or nodes. And most importantly, the configuration is already ready, no need to assemble everything manually.

3️⃣ Save memory through sequence packing
Usually, biological sequences vary greatly in length, and half of the batch is filled with paddings. Packing allows you to "compress" the batch by removing empty tokens, resulting in higher speed and less VRAM usage.


No one wants to write CUDA kernels manually. BioNeMo Recipes allow you to use the familiar PyTorch + HuggingFace stack while achieving performance at the level of "big" frameworks.

#NVIDIA

🤖 Data Science, ML & Big Data with @DataXplore
Post #1853 532
EXPECTATION vs REALITY

Pair programming is a software development practice where two developers work together at one computer,


•••••••••••••••🤖 @DataXplore
Post #1852 500
Fun fact: this picture is completely generated by the new Nano Banana Pro (if you believe the author)

Isn't it beautiful?

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1850 172
Uni-MoE-2.0-Omni

From multimodality to full omnmodal understanding and generation: Speech, Text, Images, Video, audio-video interactions.

🟢 What developers demonstrated?
How to evolutionarily transform ordinary dense LLMs into efficient MoE models capable of working with all modalities simultaneously?

🧠 Architecture

1️⃣ Omnimodality 3D RoPE + Dynamic Capacity MoE
- Unifies alignment of speech, text, images, and video in spatiotemporal dimensions
- Dynamically allocates computation depending on task complexity

2️⃣ Deeply fused multimodal encoder-decoder
- Any combinations of input and output modalities
- True omnmodal interaction and generation

🛠️ Training

1️⃣ Progressive training strategy
Cross-modal alignment → Expert warm-up → MoE + RL → Generative training
- Scales dense LLMs into MoE models
- Only 75B tokens
- Stable convergence, especially on RL

2️⃣ Language foundation for understanding and generation tasks
- All tasks reduce to language generation
- Breaks barriers between modalities

🎨 Capabilities

✔ Generation and interaction through speech
✔ Image generation and editing
✔ Image and video understanding
✔ Audiovisual reasoning
✔ 10+ multimodal tasks

🔥 Results

The model outperformed Qwen2.5-Omni (1.2T tokens) in 50+ of 76 tasks, having only 75B tokens:
- Video understanding: +5%
- Omnimodal understanding: +7%
- Speech QA: +4.3%
- Image processing: +7%


Open Source Model, Code, Homepage

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1849 180
Linear-Programming-Based Load Balancer (LPLB)

A new open-source project, an experimental load balancer for Mixture-of-Experts (MoE) models.

The repository describes how the system:
• dynamically redistributes experts based on load statistics;
• creates replicas considering cluster topology;
• solves the optimal token distribution across experts using an LP solver running directly on the GPU (cuSolverDx + cuBLASDx);
• uses load metrics obtained manually, via torch.distributed, or through Deep-EP buffers.

Guide shows what a smart and precise load balancer for large MoE architectures might look like.

GitHub

#DeepSeek #LPLB #MoE #AIInfrastructure #OpenSource

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1848 170
GPT-5.1-Codex-Max

OpenAI has a new coding model, they want to at least somewhat outshine the hype around Gemini.

🟢 Why this matters?

1️⃣ This is the first Codex trained to work on Windows and, in particular, in Powershell. Now the model understands the specifics of the Windows environment, file system structure, etc.

Also, an Agent mode has appeared for Windows, allowing the agent to work autonomously in the terminal (access can be configured).

2️⃣ They claim that the model can continuously work independently on tasks for more than 24 hours! Incredible, if true (although Anthropic declared 30 hours for their Sonnet 4.5).

“Pretrain and test-time haven’t hit a wall” – wrote Noam Brown about the release. He’s hinting that scaling continues.

At the same time, thanks to a new feature – “compaction” – Codex can now work with huge contexts. This is somewhat analogous to short-term and long-term memory. When the token limit of the context window is near, the model compresses the oldest information, then moves into the new context window with this compressed summary plus the latest relevant information. This process can repeat many times.

3️⃣ Metrics. On SWE-bench Verified, GPT-5.1-Codex-Max (xhigh) shows 77.9% accuracy. This is better than Gemini 3 and Claude Sonnet 4.5, i.e., state-of-the-art. Meanwhile, the model now saves even more tokens. At the medium reasoning level, it achieves the results of the previous version while consuming 30% fewer tokens.


Already available in IDE and Codex CLI, will be added to the API “soon”

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1847 183
DR Tulu‑8B - an open model for deep scientific analysis capable of competing with OpenAI DR, and all this with only 8B parameters!

🟢 What's new?
Reinforcement Learning with Evolving Rubrics (RLER) for long, unverifiable tasks.

💡 Instead of static evaluations:
• Rubrics evolve together with the model
• Use knowledge from search
• Extract new information directly during training

📊 Results:
• DR Tulu‑8B is comparable to OpenAI DR
• Surpassed all open-source DR models
• Cost — ~ $0.00008 per query (compared to > $1 at OpenAI)

💥 Training in two stages: SFT → RL
Tested on 4 complex benchmarks and a new medical GeneticDiseasesQA (in collaboration with clinicians) — results better than OpenAI DR and AI2 ScholarQA (Claude).


Open methodology, real impact. AI that *learns to research by itself*.

Model & Code

••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1846 222
it makes sense now
Post #1845 184
The Gelato is a minimalist library that helps build, analyze and optimize computational graphs in machine learning. by mlfoundations

🟢 Features?
- Simplifies the breakdown of complex pipelines, allows visualization of dependencies, and manages computations at the node level.
- clear representation of the graph of any ML model
- convenient tools for modification, optimization, and analysis
- suitable for experiments with new model designs and custom connections
- easy integration into existing projects


Useful if you work with non-trivial architectures, want to experiment with changing model structure, or analyze bottlenecks in computations.

GitHub, Gelato-30B-A3B (Model) & Click-100k

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1844 159
Helion - a new high-level DSL for fast and portable ML kernels

Give developers performant and portable kernels for different architectures.

🟢 What makes Helion useful?
- Combines the familiar PyTorch style with automatic tuning
- Automatically handles tensor indexing
- Manages memory and optimal accesses
- Selects settings for specific hardware
- Allows writing kernels at a "PyTorch-like" level while producing Triton-level code
- turns a simple computation description into an efficiently optimized kernel.


The developer writes minimal code and Helion does the rest.
More details

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1842 150
The whole internet yesterday literally:

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1841 232
September 26:
Cloudflare was rewritten in Rust with a safe memory management model. Link


The change is presented as "faster and safer" thanks to Rust.

November 18 (53 days later):
Cloudflare experiences a major outage that took down significant parts of the Internet due to a bug... in that very Rust code. link


••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1840 145
How to combine dozens of expert models into one universal model - without retraining and data leakage?

Researchers from CAS, HKISI-CAS, Sun Yat-sen, and Peking have presented a new approach: RobustMerge

🟢 How RobustMerge helps?

A method for training-free, parameter-efficient model merging.

What problem solved?
Each expert model specializes in something — one for OCR, another for vision, a third for dialogue, a fourth for code.
But how to assemble them into one universal MLLM so that:

- there is no data leakage
- no need to retrain everything
- accuracy is not lost
- the model does not break due to conflicting weights

🧠 What RobustMerge does
The method preserves *direction robustness* — the stability of weight directions — using two key techniques:

- low-rank analysis — highlights the main knowledge direction
- cross-task normalization — normalizes the contribution of different tasks so that one model does not "override" another

Different specialized models become one universal MLLM that continues to perform well across all areas and even improves generalization.

Solves main pain point: how to combine dozens of experts into a single system without huge retraining costs and without risking mixing private data.


GitHub

••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1839 148
Why don’t ordinary LLMs handle science well?

Scientific data is a mix of text, tables, formulas, code, images, and uncertain measurements. Nuances are easily lost.

🟢 How scientific LLMs evolve through richer data and closed loops with autonomous agents?
Analyzed 270 datasets and 190 benchmarks and gave 94-page review.

- a unified taxonomy of scientific data
- a multilayer model of scientific knowledge: from raw observations to theory

This framework helps build pretraining and fine-tuning so that models retain scientific rules and can connect different formats and scales.

The review classifies models by fields: physics, chemistry, biology, materials, earth sciences, astronomy, plus universal scientific assistants.

In quality assessment, there is a shift: from one-shot quizzes to process-oriented checks that evaluate reasoning chains, tool use, and intermediate results.

The authors promote a closed loop: agents plan experiments, run simulators or labs, verify results, and update collective knowledge.


Scientific LLMs are moving toward a data-driven, process-verified, agent-loop approach linked to real evidence.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1838 242
An interesting tool to work with data and are tired of writing SQL manually.

OpenChatBI is an open-source BI tool that allows you to make database queries in plain language and get results in the form of tables or charts.

🟢 Features:
- Built on LangGraph and LangChain,
- Can convert text queries into SQL and execute them automatically.
- For complex questions, you can connect external knowledge sources or extend functionality via MCP.
- Has a convenient Web UI, and installation takes just a couple of minutes.


Install via pip or uv, connect the database, and configure the LLM.

Repository

••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1837 148
LeJEPA : self-supervised learning

Previously, models like JEPA required various "hacks" to prevent feature collapse: stop-gradient, predictor heads, teacher-student schemes.

🟢 How it's different then previous?
LeJEPA removes all these tricks and replaces them with a single regularizer — SIGReg (Sketched Isotropic Gaussian Regularization).

What SIGReg does: it forces vector representations to be evenly distributed in all directions, forming an "isotropic" cloud.
The authors show that this form of features minimizes the average error on future tasks — meaning it is mathematically optimal geometry, not a set of heuristics.

Why this matters:
- training becomes more stable and simpler;
- easily scales to large models (tested on 1.8B parameters);
- no need for teacher-student schemes;
- the model can be evaluated without labels — its loss correlates well with quality on a linear probe.

Result: 79% accuracy of the linear probe on ImageNet-1K with minimal hyperparameters.


Work trains stably on different architectures and scales, and the approach makes self-supervised pretraining more transparent and predictable.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1836 149
Excel now with AI genius

A top AI agent has been released that will boost your spreadsheets to the max

— Instantly generates any formulas and performs calculations
— Creates charts and graphs with one click
— Organizes even huge chaotic tables
— Parses data from any websites

Free — Try Here

••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1835 468
Sparse Inference in PyTorch

New level of optimization beyond quantization means LESS memory, HIGHER speed, WITHOUT the need to change the model architecture.

🟢 What is Sparse Inference?

Sparsity means that in model's weights and activations, most values are zeroed out (for example, 80–90%).

Now PyTorch can:
- Use N:M sparsity (e.g., 2:4 sparsity)
- Accelerate inference on GPU and CPU
- Support this in torch.compile() & torch.export

How it work?
1. model is zeroed out using Pruning / Structured Sparsity
2. Converted via torch.sparse.to_sparse() or torch.export
3. Run through TorchInductor + XNNPACK or CUTLASS

What is supported?
- CPU (x86, M1/M2) via XNNPACK backend
- GPU (Ampere+) via CUTLASS
- Integration with torch.compile() (TorchInductor)

IMPORTANCE
- Less memory → lower latency on edge devices
- Higher performance, No compromises
- Easily integrates in current PyTorch pipeline


••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →