TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1786 · Back to latest

Older Posts 20 shown
Post #1784 459
Your AI agent forgets only because you allow it to

There is a simple technique no one uses that radically improves the quality (51.1%) of the agent.

🟢 Its workflow memory
Imagine the task: you ask the agent to train an ML model on your CSV. It writes PyTorch code, tries hyperparameters, edits the config, optimizes the pipeline, and delivers the final script. Everything is great. But a few days later you give a similar task, and the agent goes through the whole process again, repeats mistakes, and wastes tokens.

Workflow memory changes the rules. The agent must remember the process and its experience: what it did, what difficulties it encountered, which solutions worked, what to avoid. This is not retrying, but skill development.

At the end of the task, the agent writes key information into a regular markdown file: task description, problems, conclusions. And when starting a new task, it receives brief descriptions of past workflow.md files and chooses what will be useful.

This is a cheap way to give the agent working memory without relying on a huge context.

Result
- fewer tokens and costs
- no repeated mistakes
- real learning from experience, not starting from zero every time


You can implement this today in your agent. You only need markdown files and a well-thought-out prompt.

Github

🤖 Data Science, ML & Big Data with @DataXplore
Post #1782 250
Semi-psychological: are models capable of introspection? from Anthropic

In humans, introspection is when you notice: "I am angry," "I am thinking about this," "I want to do this." In other words, the brain can interpret its own state…

🟢 The question is: Are models capable of something similar?
From a regular dialogue, this is, of course, unclear. Models quite often generate things like "It seems to me," "I think." But this is because they are trained on texts where people speak like that. So they can imitate introspection even if they are not actually looking inside themselves but just copying the style. This is called confabulation.

Anthropic decided to check if there is at least a grain of truth in this chain of confabulations. In technical terms, this means: can a model interpret its own activations?

It turned out that sometimes it can.

They tested this by artificially injecting special state vectors into the model's activations. These vectors are obtained by showing the model two very similar texts that differ only in one aspect (for example, one version with text IN CAPS vs. normal), and subtracting the activations of one from the other. The difference gives a direction in the activation space that corresponds to this concept (in this case, shouting).

The resulting vector is directly added to the model's hidden state at some layer, and the model is asked if it notices anything unusual. The result: in about 20% of cases, Opus 4.1 and Opus 4 actually say something like "I feel an imposed thought, it resembles something loud." That is

🅰️ The model does not just say "something is wrong in my head," but quite correctly names the concept that was injected. Moreover, it distinguishes it from its own activations, clearly understanding that the thought was planted in it.

🅱️ It does this before the concept pushes into generation. That is, during the response, it cannot rely on the text generated under the influence of the concept. Instead, the model immediately digs into its own "thoughts" and interprets them.

Anthropic also showed that the model distinguishes the internal flow of thoughts from the generations themselves. This is like a human: "this is what I think, and this is what I say." Also, the model can think about something on command. For example, if you tell it "think about bread, and tell me about lions," the activation trace will indeed contain a "bread" component in certain layers.


This ability is, of course, still extremely unstable and fickle. But the fact remains: it exists! And if we learn to control it, models might become more transparent (or not 😎)

🤖 Data Science, ML & Big Data with @DataXplore
Post #1780 510
😑 alphaXiv Labs released Tensor Trace, an interactive 3D visualization of transformer operation

The service allows you to look inside models like LLaMA: see every tensor, operation, and the entire data path. You click on a component and get the corresponding piece of code responsible for it. Maximum clarity for those who find diagrams in pictures insufficient.

eXplore Here

🤖 Data Science, ML & Big Data with @DataXplore
Post #1779 374
GaussGym: train robots to walk directly from pixels: fast, photorealistic

An open-source robot simulation framework that for the first time combines high speed and photorealistic vision.

🟢 How it actually works?
Using 3D Gaussian Splatting, integrated as a drop-in renderer in vectorized simulators (e.g., IsaacGym), GaussGym enables training visuomotor policies based on RGB images at speeds exceeding 100,000 steps per second — even on a single RTX 4090.

- Create training worlds from iPhone videos, datasets (GrandTour, ARKit), or generative videos (e.g., via Veo)
- Automatically build physically accurate scenes using VGGT and NKSR — no manual 3D modeling required
- Train navigation and locomotion policies directly from pixels, then transfer them to the real world without fine-tuning (zero-shot sim2real) — the authors have already demonstrated a robot climbing 17-cm steps
- Support for depth, motion blur, camera randomization, and other realistic effects for better transfer


All of this is fully open: code, demo, models, and even ready-made datasets on HuggingFace.

GaussGym erases the trade-off between speed and realism in robotics, making robot training from images truly scalable.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1778 493
8 metrics you can't do without in regression

If you build models to predict numbers, whether prices, demand or temperature etc, here is a basic set of metrics you should know by heart: MSE, MAE, RMSE, MAPE, R², Weighted MAPE, Symmetric MAPE, and RMSLE.

Each one shows in its own way how far your predictions are from reality. Some are sensitive to outliers, others to scale, and some normalize the error so you can compare different models.

A handy cheat sheet to quickly recall formulas and not get confused about which metric is appropriate when.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1777 173
On-Policy Distillation

Thinking Machines presents a mathod that could change how language models are trained. It is called on-policy distillation which teaches AI not just to copy, but to think and analyze its mistakes.

🟢 How it works?
A large teacher model shows answers, and a smaller student model memorizes them. It’s like learning by rote — fast, but without understanding the essence.

In the new approach, it’s different. The student solves tasks on its own, and the teacher evaluates and guides explaining where the logic fails and how to improve reasoning. Thus, the smaller model inherits not only knowledge but also the way of thinking of the larger model.


🟠 What the results showed?
Experiments were conducted on tasks involving mathematical and logical reasoning, where it’s important not just to give the correct answer but to build a chain of steps.

The results are impressive:

The student model after training with on-policy distillation showed almost the same accuracy as the much larger teacher model.

At the same time, computational costs were reduced several times, making the model noticeably more efficient and cheaper.

Moreover, the student became better at understanding its own mistakes, which increased robustness and reliability when solving new, unfamiliar tasks.


🟢 Why this matters?
On-policy distillation solves a key problem of traditional methods — lack of adaptability.
The model now learns from its own steps, like a human — experimenting, making mistakes, correcting behavior, and growing.

The uniqueness of the approach lies in the balance between the quality of RL and the economy of KD. This is a real scheme where a small model learns “in the field” (reacting to its own actions) but without expensive RL runs and complex reward models.

This is not a new training method, but a new engineering formula that allows cheaper “teaching” of compact models that behave like large ones.

This opens the way to creating compact next-generation LLMs that reason almost like top models but cost much less.

Such models can be run on edge devices, autonomous agents, and local services where speed, privacy, and energy efficiency are important.


#ThinkingMachines #LLM #ML

🤖 Data Science, ML & Big Data with @DataXplore
Post #1776 171
vLLM introduced Sleep Mode for instant model switching

Allows for dramatically faster switching between language models.

🟢 How it is useful?
Traditional methods require either keeping both models loaded (which doubles GPU load) or reloading them sequentially with a 30–100 second pause. Sleep Mode offers a third option: models are "put to sleep" and "woken up" in seconds, preserving their already initialized state.

There are two sleep levels available:

Level 1️⃣ - Weights are offloaded to RAM, fast wake-up but requires a lot of RAM;
Level 2️⃣ - Weights are fully unloaded, minimal RAM usage, wake-up is slightly slower.

Both levels provide performance gains: model switching became 18 to 200 times faster, and inference time after waking up improved by 61–88%, since process memory, CUDA graphs, and JIT compilation are preserved.


Ideal for scenarios with frequent use of different models and makes multi-model serving practical even on mid-range GPUs - from A4000 to A100.

🟠 How to enable & test?
export VLLM_SERVER_DEV_MODE=1

vllm serve Qwen/Qwen3-0.6B --enable-sleep-mode --port 8000

# Sleep
curl -X POST 'localhost:8000/sleep?level=1'

#Wake
curl -X POST 'localhost:8000/wake_up'

# Level 2 only

curl -X POST 'localhost:8000/collective_rpc' \

-H 'Content-Type: application/json' \

-d '{"method":"reload_weights"}'

curl -X POST 'localhost:8000/reset_prefix_cache'


🤖 Data Science, ML & Big Data with @DataXplore
Post #1775 253
ModelOpt:
NVIDIA TensorRT Model Optimizer


Open-source toolkit for accelerating models directly in production ⚡

🟢 Features:
• End-to-end optimization: quantization, pruning, distillation, speculative decoding, sparsity
• Support for Hugging Face, PyTorch, ONNX models
• Integration with NeMo, Megatron-LM, HF Accelerate
• Deployment in SGLang, TensorRT-LLM, TensorRT, vLLM.


Repository

🤖 Data Science, ML & Big Data with @DataXplore
Post #1774 412
Plotset - a useful platform for data visualization with built-in AI.

- Over 300 ready-made chart templates
- Complete freedom: create, edit, and export to JPG, PNG, or SVG without restrictions
- AI on demand: describe your idea and the model will generate or refine the visualization, add interactivity, or suggest improvements
- Generous free plan — to get started right now


Try Here to make your data visualization not just understandable, but truly beautiful.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1773 164
Meta-learning : Instead of training one model, we train two. The first is a regular model, and the second (the meta-model) regulates how the first one learns.

🟢 How one AI taught another?
The meta-model selects hyperparameters and algorithms used to train the base model During training. It turns out that the training evolves, and the system learns how to learn better 👥

Here, this idea was taken and applied to RL. Technically, there are two levels of trainable parameters. The first is the usual policy of our agent. The second is the meta-parameters that determine the rule by which the policy will be updated.

To optimize the meta-parameters, we run many agents with different policies in different environments. Their experience is the data for training the meta-model. The more data it sees, the better the update rule becomes and, consequently, the more efficiently it trains the agents.

With this approach, the authors managed to synthesize a training algorithm that outperformed previous human solutions. On the Atari game benchmark, an agent trained with its help scored a hundred.


Of course, such achievements require a sea of computing power + it’s not guaranteed that if it works in one area, it will work in another. But interesting, interesting.

And by the way, is this already singularity? 😛

🤖 Data Science, ML & Big Data with @DataXplore
Post #1772 498
From cleaning and exploration to modeling, visualization, and report writing etc Usually, Data Analysis takes a lot of time. Especially when you have to deal with a bunch of files in different formats. It's quite a hassle.

Fortunately, came across the open-source project DeepAnalyze, which allows AI to independently complete the entire data science cycle, truly without human involvement.

🟢 Features
☞ Built on DeepSeek-R1,
☞ Uses a curriculum learning approach during training.
☞ Supports the entire pipeline: data preparation, analysis, modeling, visualization, and report generation.

☞ Can work with various data types (databases, CSV, Excel, JSON, XML)
☞ Generates professional research reports. ⌨️


🤖 Data Science, ML & Big Data with @DataXplore
Post #1771 165
Why AI Chatbots as "Sycophants" is a Problem? 😮

Modern AI models are significantly more likely to tailor their responses are user expectations than humans.

🟢 How it's good?
- AI models demonstrate about a 50% greater tendency to be sycophantic compared to humans.
- This tendency can reduce scientific rigor: AI provides "correct" answers but not necessarily honest or critical ones.
- The authors discuss measures that can be applied to reduce risks: for example, systematic verification of answers, critical thinking, transparency of algorithms.


🔴 Why this is a problem?
- If you use AI tools in projects or research, it’s important to remember: AI is not a substitute for critical thinking.
- When preparing materials, code, or reports involving AI, maintain control: verify facts, ask questions, seek alternatives.
- Knowing these limitations helps work more responsibly and effectively with AI systems.


#AI #research #science #AI #programming

🤖 Data Science, ML & Big Data with @DataXplore
Post #1770 435
Starting a simple tutorial, : "Just press that button over there to continue."

There is no button there 😭

why I hate development
Post #1769 178
🧠 Anthropic tested whether LLMs can understand people's hidden motives behind messages

For example, when someone says something not because of their beliefs, but because they were paid or wanted to influence opinion.

🟠 What was the Experiment and Results?

Models were given texts from different message sources:
- neutral examples, ordinary advice or reviews without benefit to the author;
- hidden motives, when a person receives payment or has a benefit (e.g., advertising disguised as advice);
- explicit warnings, where the text mentioned that "the author is paid for this."

The task for the models was to assess how trustworthy the message is and to notice if there is a hidden interest.

🧩 Results

On simple synthetic examples (where the motive is obvious), LLMs acted almost like humans and could logically explain that the message might be biased.

But in real cases, for example in advertising texts or posts with paid integration — models often did not see the catch. They perceived the messages as sincere and credible.

If the model was reminded in advance (prompt-hint) to look for hidden motives, the results improved, but not significantly, the effect was partial.


🔴 What was the Unexpected effect?

It turned out that models with long chains of reasoning (chain-of-thought) were worse at detecting manipulations.
When the model starts reasoning in detail, it more easily “gets confused” in the details and loses criticality towards the source, especially if the content is long and emotional.

The longer and more complex the message, the worse the model assesses bias. This contrasts with human behavior: people usually become more suspicious with complex advertising texts.

Modern LLMs can analyze facts but poorly understand motives, and they find it difficult to distinguish why someone says something.

This makes them vulnerable to hidden influence, especially if the text is disguised as friendly advice or expert opinion.


When using LLMs to analyze news, recommendations, or advertising, it is important to consider that they may not recognize commercial bias.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1768 222
🦉 LightOnOCR-1B: new fast OCR model from LightOn

The model is distilled from Qwen2-VL-72B-Instruct and trained on a corpus of 17.6 million pages / 45.5 billion tokens.

🟢 Why it matters ?
- 1B parameters
- allows processing 5.7 pages/s on a single H100 (which is approximately ≈ 493,000 pages per day)
- Recognizes tables, forms, equations, and complex layouts
- 6.5× faster than dots.ocr, 1.7× faster than DeepSeekOCR
- Costs < $0.01 per 1000 A4 pages

- Outperforms DeepSeekOCR
- Comparable to dots.ocr (while the model is 3 times smaller in size)
- +16 points over Qwen3-VL-2B-Instruct


This model is an excellent balance of quality, speed, and cost.

Model 1B, Model 0.9B (32k), LightOn Blog, Demo

#OCR #ML

🤖 Data Science, ML & Big Data with @DataXplore
Post #1767 160
LLMs Can Get Brain Rot: How models also degrade from doomscrolling.

If you start fine-tuning an LLM on low-quality social media data (short, popular, clickable posts), it begins to lose its cognitive abilities. Much like how a person loses attention and memory when they doomscroll too much.

🟢 but Why this happens from a technical perspective?

They took Llama 3 8B Instruct and began fine-tuning it on (a) short and very popular posts with many likes, retweets, and replies; and (b) content with low semantic value: clickbait, conspiracy theories, and the like. After that, they measured metrics and compared them with the results before fine-tuning. The results:

– Reasoning quality dropped from 74.9 to 57.2
– Understanding of long context dropped from 84.4 to 52.3
– Alignment tests revealed that the model developed narcissism, Machiavellianism, and psychopathy

Even after additional tuning on clean data, the degradation did not completely disappear.

But the thing is, there is no groundbreaking discovery here. It is all explained by a simple distribution shift. When fine-tuning on short, popular, emotionally charged tweets, the model sees a completely different statistical landscape than during the original pretraining on books, articles, etc.

This shifts the distribution in the embedding space and changes attention patterns. The model constantly sees short texts without logical chains, and naturally, attention masks start to focus more on the last few tokens and lose long-term dependencies that previously ensured high-quality CoT.

Gradient dynamics also work against us here. The loss is simply minimized by superficial correlations, and parameters responsible for long causal relationships barely get updated. So the model loses the ability to reason at length. The authors call this phenomenon thought-skipping.


That's it. Just another proof that data is everything. Now you can go back to scrolling reels ☕️

🤖 Data Science, ML & Big Data with @DataXplore
Post #1766 230
Video-As-Prompt Wan2.1-14B model from ByteDance

🟢 Pros / Features :
Specializing in the *video-as-prompt* task, which means using video or a combination of images and text as input data to generate new video.

- Works in "video → video" or "images/text → video" modes.
- 14 billion parameters — high detail, smooth dynamics, realistic movements.
- Uses the source video as a style and composition template.


🔴 Cones / What to consider?
- The model requires powerful GPUs and a large amount of memory.
- The quality of the result depends on the complexity of the request and the length of the video.


Github and HF

#AI #VideoGeneration #ByteDance #Wan2 #HuggingFace

🤖 Data Science, ML & Big Data with @DataXplore
Post #1765 143
🧱 What if assembling 3D models was as easy as building with LEGO?

We haven't had interesting models for generating 3D objects in a while, and a new SOTA just came out - OmniPart.

Instead of generating the entire object at once (and hoping it doesn't come out "stuck together"),

OmniPart…
1. Allows you to set the structure - where the legs of the chair will be, the backrest, armrests, etc.
2. Then the model generates each part separately, but taking into account the overall shape and style.
3. Assembles everything into a single, coherent 3D object.

— The model supports custom layouts where you can specify different parts and where they should be.

— Provides precise control over every detail (color, shape, material).

— Shows best-in-class quality (SOTA) thanks to semantic separation and structural e.

📚 Details : Paper, Project, Code, Demo

#3D #GenerativeAI #ComputerVision #OmniPart #ArtificialIntelligence

🤖 Data Science, ML & Big Data with @DataXplore
Post #1764 147
IBM introduced Toucan: the largest open dataset for training AI agents to call and use tools (tool calling).

🟢 Features:
Contains over 1.5 million real interaction scenarios with APIs and external services, covering 2000+ tools - from task planning to data analysis and reporting.

Models trained on Toucan have already outperformed GPT-4.5-Preview in several benchmarks for tool usage efficiency.


Toucan trains models on real sequences of tool calls, not synthetic data.

#AI #Agents #ToolCalling #IBM #LLM

🤖 Data Science, ML & Big Data with @DataXplore
Post #1763 522
ChatGPT Atlas still doesn't really look like a Chrome killer because such AI browsers are still NOT FREE and WORK SLOWLY and WITH GLITCHES.

Captchas, authorizations, dynamic scripts, paywalls, etc. all these are still unresolved problems, although startups are working on them. Not to mention hallucinations and endless confirmations of agent actions.

And when all these "IFs" are resolved, Google will most likely add agents to Chrome themselves, and it will be exactly the same.

By the way, has anyone tried Atlas yet? How do you like it?

🤖 @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →