TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
585
Photos
845
Videos
446
Links
675

Showing posts older than #1722 · Back to latest

Older Posts 20 shown
Post #1721 116
AI21 introduced Jamba 3B that outperformed Qwen 3 4B and IBM Granite 4 Micro in reasoning quality.

🟢 What's the secret in architecture?
- a combination of Transformer attention and Mamba state-space layers.
- The Mamba part efficiently processes long sequences without heavy attention caches,
- while the Transformer layers maintain the ability for complex reasoning.

As a result, the model uses less memory, delivers high speed, and runs smoothly even on laptops, GPUs, and mobile devices.


🟠 Features?
- Context: up to 256K tokens.
- Speed: about 40 tokens/sec even on long contexts, while other models slow down sharply.

Higher efficiency compared to AI21 - 2–5× performance improvement over competitors thanks to a smaller KV cache, hybrid architecture.

On "intelligence versus speed" graph, Jamba 3B surpasses Gemma 3 4B, Llama 3.2 3B, and Granite 4.0 Micro.


Truly superior intelligence and faster generation on Hugging face

#LLM #Jamba3B #AI21 #DeepLearning

🤖 Data Science, ML & Big Data with @DataXplore
Post #1720 254
Free course on learning deep learning concepts by FreeCodeCamp (duration: 05:07:27 hours)

🟢 What it covers?
A conceptual and architectural journey through computer vision models in deep learning, tracing the evolution from LeNet and AlexNet to ResNet, EfficientNet, and Vision Transformers.

The course explains the design principles behind skip connections, bottleneck blocks, identity preservation, depth/width trade-offs, and attention.

Each chapter combines clear illustrations, historical context, and side-by-side comparisons to show why architectures look the way they do and how they process information.


Available on YouTube

#LearningSunday #Recommended

🤖 Data Science, ML & Big Data with @DataXplore
Post #1719 393
🚀Qwen has released a guide for working with Qwen3-VL!

This is a collection of interactive notebooks demonstrating the capabilities of Qwen3-VL - both for local deployment and via API.

Inside are dozens of real examples with explanations:
☞ Working with images and reasoning about them
☞ An agent for interacting with interfaces (Computer-Use Agent)
☞ Multimodal programming
☞ Object and scene recognition (Omni Recognition)
☞ Advanced data extraction from documents
☞ Precise object detection in images
☞ OCR and key information extraction
☞ 3D analysis and object anchoring
☞ Understanding long documents
☞ Spatial reasoning
☞ Mobile agent
☞ Video analysis and understanding


GitHub, Qwen3-VL, API documentation and you can Try Here.

#Qwen #Qwen3VL #AI #VisionLanguage #Multimodal #LLM

🤖 Data Science, ML & Big Data with @DataXplore
Post #1718 123
🧠 DataMind : New open system for creating smart data analysis agents.

Today, most data agents use closed models and depend on prompt engineering.
Open solutions can't consistently reason step-by-step or work with different data formats.

🟢 How DataMind works?
The DataMind team solved these 3 main problems:
1. Lack of quality data for training
2. Incorrect training strategies
3. Errors in multi-step code execution

The system includes a full cycle of data generation to training and task execution.
It uses:
- task classification and query creation from simple to complex
- trajectory filtering through self-consistency (answer self-checking)
- a combination of dynamic SFT and RL training, which stabilizes the process
- optimized code execution in an isolated environment


🟣 Results
- The DataMind-14B model showed an average score of 71.16% and outperformed GPT-5 and DeepSeek-V3.1
- The lightweight DataMind-7B version became the best among open-source solutions — 68.10%, trained on 12,000 trajectories

Main conclusions
- Filtering through self-consistency is more effective than choosing a single "best" trajectory
- SFT losses stabilize training but cause fluctuations if misconfigured
- RL reduces the gap between models but does not change the overall ranking


Released DataMind-12K dataset, DataMind-7B and 14B models so community can build their own analytical agents. Research, Code, Models and data on HF

#LLM #Agents #OpenSource #ReinforcementLearning

🤖 Data Science, ML & Big Data with @DataXplore
Post #1717 343
A clear comparison of speed of the new python 3.14 with previous version

Note that now multithreading has become even faster than multiprocessing. All because new build allows working without the GIL.

A brief explanation.
GIL (Global Interpreter Lock) is a global interpreter lock that allows only a thread of Python bytecode to execute at a time (even if you have 16 cores). So previously, before 3.14, multithreading as such didn't exist in Python.

To bypass GIL, multiprocessing was used. There, each process is a separate interpreter instance and each process has its own GIL. This was only way to parallelize cores in Python. But there was a downside: each process had its own copy of memory, and data had to be serialized when passed. This caused significant overhead.

Now, in the new version without the GIL, threads operate in the same address space with shared memory access. The result is immediately reflected in speed: multithreading is now 33% faster than multiprocessing. In 3.13, by the way, it was exactly the opposite.


Waiting for free-threading support in PyTorch and NumPy - Try Here

🤖 Data Science, ML & Big Data with @DataXplore
Post #1716 931
pov: when your brain already left for the weekend but your boss schedules a 6 PM call.

🤖 @DataXplore
Post #1715 805
🔥 South Korea’s Biggest Digital Disaster 🇰🇷

A fire at the Daejeon National Data Center (NIRS) wiped out 858 TB of government data with no backup.

647 online services crashed, from citizen portals to emergency response.
Visa files, project data, postal records gone forever. (imagine thousands of lost letters and parcels)

Golden Advice: Always Have A Backup,
Coz Data is everything. 💾

🤖 Data Science, ML & Big Data with @DataXplore
Post #1714 341
Never use the describe method from Pandas

Skimpy is a much more convenient (and open source) alternative that provides an extended data description: dataset shape, data types by columns, statistics, distribution plots, etc.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1713 121
Ling-1T - a new model from inclusionAI with 1 trillion parameters

Main idea of the model: to combine efficiency and scale of reasoning in one architecture.

🟠 Why it really matters?
- Total parameters: 1 trillion, of which ≈ 50 billion are active per token (MoE architecture).
- Trained on 20 trillion+ tokens, specially selected for logical thinking and reasoning tasks.

Context: 128,000 tokens.
Inside Evo-CoT (Evolutionary Chain of Thought) and Linguistics-Unit RL - new training methods for scalable reasoning.

Ling-1T is positioned as a model balancing speed and accuracy of responses.

The model demonstrates strong results in code, math, logic, and frontend generation tasks.

The architecture involves Mixture-of-Experts (1/32 activation), MTP layers, and expert routing.


Ling-1T shows that huge models can be made not only powerful but also efficient.

HF

#Ling1T #AI #ML #OpenSource #Reasoning #TrillionScale #FP8

🤖 Data Science, ML & Big Data with @DataXplore
Post #1712 118
ReSum is a new method that allows web agents to search longer and respond more accurately.

🔴 What was problem with ReAct?
Agents in ReAct keep a detailed “diary”: they think, take an action (search, click), record the result, and repeat the cycle.

This makes the process transparent, but in long tasks the history quickly grows → context limit → loss of details.


🟢 What is ReSum’s solution?
⁠☞ When the context nears the limit, the agent stops and writes a summary: verified facts + still open questions.
+4.5% quality improvement compared to ReAct
Pass@1: 33.3% and 18.3% on the challenging BrowseComp tests

⁠☞ Then it continues from this summary instead of a long conversation.

⁠☞ A separate 30B model for summaries that better handles “noisy” pages and highlights important information.

⁠☞ Reinforced learning ReSum-GRPO: up to +8.2% with ReSum-GRPO : the agent receives a reward only for the final answer, which is distributed across all intermediate steps. This teaches it to gather correct facts and create concise, useful summaries.


Agents stay within token budget → solve complex web search → analysis tasks better than classic ReAct.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1711 130
Tiny Recursive Model (TRM) : new neural network from Samsung

Its a different way of thinking: This is 10,000 times smaller than modern LLMs, but the result is better that truly thinks before it speaks.

🟢 How TRM works?
1️⃣ Draft answer: the model immediately forms a quick sketch of the solution instead of writing it word by word.

2️⃣ Scratchpad: creates an internal space for logic and intermediate reasoning.

3️⃣ Self-criticism: repeatedly (6 times) checks its reasoning, refining and correcting errors.

4️⃣ Rewriting: based on improved logic, creates a new, more accurate version of the answer.

5️⃣ Iteration: repeats the process up to 16 times until it reaches a confident, logically coherent solution.


🟠 Why TRM is interesting?
⁠☞ Outperformed DeepSeek-R1, Gemini 2.5 Pro and o3-mini in reasoning tasks ARC-AGI 1 and ARC-AGI 2 despite having only 7 million parameters and about 1000 training examples.

⁠☞ Lower computational costs with higher results; high efficiency at low expense.

⁠☞ Proof that inherent logic and architecture can be stronger than just model size. It can be briefly described as: "think before you act."

⁠☞ Powerful reasoning systems become accessible even without huge clusters; the model can run on limited resources.


Github

#TinyRecursiveModels #TRM #DeepLearning #NeuralNetworks

🤖 Data Science, ML & Big Data with @DataXplore
Post #1710 119
Biological Dragon Hatchling

Polish startup Pathway has introduced a A Brain-Inspired Neural Network combining transformers with biological brain models.

🟢 What is the Core Idea?
Models neurons as graph vertices and synapses as weighted edges.

Neurons communicate only with neighbors — mimicking brain-like local interactions.

Trains using Hebb’s rule (“neurons that fire together, wire together”).


🟠 Key Properties
Two weight types of Architecture:
Fixed → long-term knowledge (updated only during training)
Dynamic → short-term reasoning, updated per inference step

BDH-GPU tensor version = transformer-like (attention + MLP + ReLU).

1. Interpretability: each neuron pair has a visible synapse → clear concept mapping.
2. Scalability: models can be merged by concatenation.
3. Performance: follows GPT-2-like scaling laws and similar accuracy.

A fascinating blend of neuroscience + transformers — potentially a major step toward more interpretable, brain-like AI.

#AI #Neuroscience #Transformers #MachineLearning #Research #Pathway

🤖 Data Science, ML & Big Data with @DataXplore
Post #1708 393
⚡️ Google has released Jules Tools - a new command-line utility and API for managing their AI agent directly from the terminal.

Jules is an AI that can write code, fix bugs, and create tests for your projects.

It connects to GitHub or another repository, analyzes the codebase, and performs the tasks you assign to it.

🟢 What you can do with Jules Tools?

With Jules Tools, you can launch and control this agent directly through the terminal, without a browser.

For example, enter:
jules remote new --session "fix login bug"

When run, the command creates a virtual machine, clones the repository, solves the task, and sends a pull request with the completed fix.

What's interesting:
- Command line and API for managing the agent
- Asynchronous tasks and parallel execution
- Scripts and automation (via CI, cron, pipelines)
- Memory and adaptation to your coding style
- Secure storage of keys and tokens
- Interactive terminal interface (TUI) showing task status in real time

The TUI mode resembles a web panel but works right in the console, allowing you to quickly launch, track, and manage sessions.

Jules can be integrated with Slack or build systems - the agent creates and executes tasks while you focus on other things.

If the agent encounters a problem, it pauses and requests help instead of "guessing" the solution.

Both utilities - Jules and Gemini CLI - run on Gemini 2.5 Pro, but Jules is aimed at short, precise tasks, while Gemini CLI is for long-term collaborative work.

The free version allows running 15 tasks per day (up to 3 simultaneously).

Paid plans - $19.99 and $124.99 - provide limits up to 100 and 300 tasks.

Google also plans to add support for GitLab, Bitbucket, and local projects without Git.


#Google #Jules #AI #CodingAgent #Gemini25Pro #Automation

🤖 Data Science, ML & Big Data with @DataXplore
Post #1707 119
⚡️ The ModernVBERT model with 250 million parameters shows results comparable to or exceeding models that are 10 times larger in document retrieval tasks.

🟢 Why it matters?

The model leads among models with up to 1 billion parameters and encodes queries 7 times faster on regular CPUs.

Unlike decoders that read text left to right and cannot revisit earlier tokens, ModernVBERT uses a bidirectional text encoder trained on word masking and a small visual module.

Each page image is split into patches that are mapped into the same space as the text and then combined with word tokens.

The late interaction mechanism retains vectors of all tokens, allowing each query token to find the most precise match. This combination of bidirectional attention and late interaction outperforms decoder architectures in document retrieval.

Higher page resolution and a short "high-resolution cooldown" phase improve retrieval accuracy, although they may degrade performance on regular images. Adding "text-only" pairs in contrastive learning helps the model effectively unify text and visual spaces.

ColModernVBERT remains compact, demonstrates high benchmark scores, and runs efficiently even on standard CPUs.


🤖 Data Science, ML & Big Data with @DataXplore
Post #1706 120
🚀 NeuTTS Air - on-device TTS with instant voice cloning

The first realistic speech synthesis model running on-device, without an API.

Format - GGML, which allows it to work on phones, laptops, and even Raspberry Pi.
Voice cloning in 3 seconds: a short audio snippet is enough to construct a voice for subsequent syntheses.

Based on a lightweight language core (0.5 B) + NeuCodec neural codec, providing a balance between quality and speed.
Generated audio is watermarked using Perceptual Threshold Watermarker to combat misuse.

GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Post #1702 198
We have a regular segment on the air: "GPT-5 solved another difficult math problem"

What are that two problems?

1️⃣ Yu Tsumura’s 554th Problem.
This is a problem from Yu Tsumura’s collection, roughly at the IMO level. The essence is to prove the triviality of a certain group defined by relations for its two generators.

Recently, due to its short formulation, it has become a kind of test for AI (i.e., whether the model has reached the IMO level or not).

GPT-5 became the first model to solve this problem. It reasoned for only 15 minutes.

Interestingly, just a month ago, an article titled "No LLM Solved Yu Tsumura’s 554th Problem" was published, in which the authors proved that current models still lack the ability for such problems. This is another testament to the speed of progress.


2️⃣ NICD-with-erasures majority optimality.
This is a problem from information theory related to recovering the original signal through a noisy channel. A pair of independent participants tries to guess the same function of the original data based on their partially erased versions of observations, aiming to maximize agreement.

The point here is that scientists long believed that the majority function in this problem was optimal. GPT-5 proved the opposite for the first time, by finding a counterexample.

This is a fundamental problem in information and communication theory. Finding the optimal function means better designing data recovery codes, storing them, etc. The practical applications are huge, and GPT-5, it turns out, has opened a new chapter for research.


Both solutions were published by independent mathematicians.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1701 318
📘 Learning Deep Representations of Data Distributions, a new free book from UC Berkeley researchers (Sam Buchanan, Druv Pai, Peng Wang, Yi Ma).

The main idea of the book is to show

Why and How deep neural networks learn to extract compressed, informative representations of complex data, and what is inside them?

Read book online

Github

#DL #representationlearning #UCBerkeley #ML

🤖 Data Science, ML & Big Data with @DataXplore
Post #1700 772
⚠️ New CYBERATTACK with AI — Trail of Bits demonstrated that hackers can hide instructions (hidden prompts) in images.

🔴 How it happens?
- Everything is clean until the image is at its original size But as soon as a service (for example, Gemini CLI or Vertex AI Studio) automatically compresses it, hidden text appears.

- AI "sees" the hidden prompt and executes it, thinking it is a user command.

- This can be used to bypass filters and make the model do what the attacker intended.


🟢 How to protect yourself?
Even a harmless image can turn out to be a "Trojan horse" for AI systems.

- The Anamorpher tool (open-source) for generating and detecting such attacks.

- Protection: multi-level image verification and tracking artifacts during scaling.


Github

#AI #Security #PromptInjection #TrailOfBits

🤖 Data Science, ML & Big Data with @DataXplore
Post #1699 135
When we talk about RAG, people usually think like this: indexed the document → then retrieved the same document.

But indexing ≠ retrieval.

The data you index does not have to be the same data you feed into the LLM during generation.

🟢 What are the 4 smart ways to index data?

1️⃣ Chunk Indexing

→ The most common approach.

→ The document is split into chunks, then each chunk is converted into an embedding and stored in a vector database.

→ At query time, the nearest chunks are retrieved by cosine similarity (or another metric).

Simple and effective, but chunks that are too large or "noisy" can reduce accuracy.

2️⃣ Sub-chunk Indexing

→ Take the original chunks and further split them into smaller sub-chunks.

→ Index these smaller fragments.

→ At retrieval, still return the larger chunk for context.

This approach is useful if the document contains several different concepts in one section — increasing the chance of an accurate match with the query.

3️⃣ Query Indexing

→ Instead of indexing raw text, hypothetical questions are generated that the LLM thinks the chunk can answer.

→ These questions are embedded and stored.

→ At real user query time, search is performed over these "synthetic" questions.

→ A similar idea is used in HyDE, but there the hypothetical answer is matched with real chunks.

Great option for question–answer (QA) systems, as it reduces the semantic gap between the user query and the indexed data.

4️⃣ Summary Indexing

→ An LLM is used to generate a brief semantic representation (summary) for each chunk.

→ The index contains the summary, not the original text.

→ At retrieval, the original chunk is returned for context.


Especially effective for dense or structured data (e.g., CSV or tables), where raw text embeddings do not yield meaningful results.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1696 233
Airweave — Future of real-time RAG systems

🟢 Key Features:
Now it is possible to build agents that search for data in any applications, databases, and document repositories in real time.

The Airweave tool creates live, bi-temporal knowledge bases so that agents always work with the freshest facts.

It connects to Notion, Google Drive, SQL databases, and turns their contents into indexable knowledge.

All of this runs locally in a Docker container, with the ability to expose an API and MCP server.


The author showed the full setup and a live demo, and also shared a link to the GitHub project.

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →