TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1835 · Back to latest

Older Posts 20 shown
Post #1834 184
Teaching models to see like humans

Several vision models tested on the classic game "Find the odd picture out"? and found that they don't always reason like humans because models, still see the world differently than humans.

This is an important drawback because if a model does't structure its internal representation of world like a human, its decisions can be illogical and unreliable.

🟢 What DeepMind did?
Humans group objects by meaning, while models more often group by visual/textural features. For example, in the task <starfish, cat, fox>, humans choose the starfish because it lives in water, while models choose the cat because the picture stands out in the color palette.

– Took a large visual model, froze it, and artificially attached a small adapter trained only on a dataset with human answers for odd-one-out.

– Generated a huge dataset with millions of odd-one-out decisions using this model.

– Based on this large dataset, fine-tuned other models so that their internal representations became closer to human logic of grouping.

So, it seems like the models were just trained on some children's game, nothing surprising. But it turned out that this qualitatively changed their hidden space. See the third screenshot: on the left, the model before alignment, on the right, after. You can see clear clusters appearing that correspond to human logic (for example, animals, food, etc.). Beautiful, right?

Also, such a model turned out to be more robust in terms of distribution shifts. For example, if the background, lighting, or other conditions in the pictures change, its answers still follow logic.


Overall, it can be expected that such a model will generalize faster than usual.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1833 518
78 topics to master in the field of Data Science

The score that actually matters

Below 25 → Still learning fundamentals
25–35 → can do solid personal projects
35–45 → The Job-ready zone
45–55 → Strong candidate
55+ → You’re rounding out specialties

🤖 Data Science, ML & Big Data with @DataXplore
Post #1832 180
How AI really works?

The OpenAI team created an interpretable model which is much more transparent than typical transformers, behave like a "black box."
This is important because such a model helps understand why AI hallucinates, makes mistakes, or acts unpredictably in critical situations.

The new LLM is a sparse transformer: much smaller-simpler than modern LLMs (at level of GPT-1). but goal is not to compete, but to be as explainable as possible.

🟢 How it works?
- the model is trained so that internal circuits become sparse,
- most weights are fixed at 0,
- each neuron has not thousands of connections, but only dozens,
- skills are separated from each other by cleaner and more readable paths.

In usual dense models, neurons are connected chaotically, features overlap, and understanding the logic is difficult.
Here, for each behavior, a small circuit can be identified:
sufficient, because it performs the required function itself,
and necessary, because its removal breaks the behavior.

The main goal is to study how simple mechanisms work to better understand large models.

The interpretability metric here is circuit size,
the capability metric is pretraining loss.
As sparsity increases, capability drops slightly, and circuits become much simpler.

Training "large but sparse" models improves both metrics: the model becomes stronger, and the mechanisms easier to analyze.

Some complex skills, such as variables in code, are still partially understood, but even these circuits allow predicting when the model correctly reads or writes a type.

The main contribution of the work is a training recipe that creates mechanisms
that can be *named, drawn, and tested with ablations*,
rather than trying to untangle chaotic features post hoc.

LIMITS: these are small models and simple behaviors, and much remains outside the mapped chains.


This is an important step toward true interpretability of large AI.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1831 421
Want to build your own AI agent?
Here is EVERYTHING you need. One enthusiast has gathered all the resources to get started:
📺 Videos,
📚 Books and articles,
🛠️ GitHub repositories,
🎓 courses from Google, OpenAI, Anthropic and others.

Topics:
- LLM (large language models)
- agents
- memory/control/planning (MCP)

All FREE and in one Google Docs

#AI #LLM #AgenticAI #GenAI

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1830 172
AgileThinker : An agent that thinks and acts simultaneously

Every action has a strict deadline: if missed, a safe default move is executed, to combine instant reaction and parallel planning.

Reactive agents act quickly but foolishly, while long planners act smartly but too slowly and often are late.

The combination works better than both.
- FAST : based on partial plans and fresh observations
- PLANNING : constantly updates strategy and complements the plan

Under load and strict timing, AgileThinker consistently outperforms both baseline approaches, the fast one and the "slow thinking" one.


A step towards agents that maintain intelligence without losing speed and can operate in dynamic environments where delay = error.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1829 184
Found an awesome resource for everyone who wants to improve their math skills for deep learning — this section

No fluff, just what you need to work in ML. Mathematical analysis, linear algebra, probability theory — all in a convenient format and immediately with code

A nice bonus: you can choose the dialect in which examples are shown (PyTorch, Keras, or MXNET).

By the way, the other chapters are just as worthy 🤓

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1828 172
Spatial intelligence

Modern LLMs work well with text, but Spatial intelligence is exactly where it's lack. Its about the ability to perceive, understand and reason about space, objects, movement and the interaction of things.

True intelligence cannot exist without perception-action link which is the foundation base of intelligence in living beings. We can't see AGI until we have truly high-quality world models with this properties:

1. GENERATIVITY : Model must be able to create entire coherent and physically plausible scenes or worlds.

2. MULTIMODALITY

3. INTERACTIVITY : Should not be a passive generator but a model that changes the state of world and can predict consequences if an agent performs some action.


Even problem statement is not solved yet: it should be something like predicting the next token, but for space.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1827 168
RL Really improve Reasoning Capacity of LLMs Beyond Base Model?

Reinforcement Learning (RL) doesn't improve model's reasoning skills, it just repackages what was already present in base model's distribution.

🟢 How was this tested?
The main metric is pass@k: a task is considered solved if among k attempts (samples) of the model there is at least one correct answer. This metric is very suitable for the authors' hypothesis because it reflects the potential of the model to solve the task with a reasonable number of attempts.

And here is what happens. For small k, RLVR models indeed more often hit the correct answer (i.e., they have a higher pass@1), but as k grows, the base models catch up and surpass RLVR on almost all task sets and model families.

This means that these methods do not expand the boundaries of solvable tasks (including mathematical and coding ones); they simply increase the efficiency of sampling existing trajectories aka the probability of immediately taking the right path, and that's why they work. Is that bad? No. But it means that too high hopes should not be placed on RLVR: everything still ultimately depends on pretraining.


🤖 Data Science, ML & Big Data with @DataXplore
Post #1826 494
Multi-Head Attention in LLM, visual explanation

🤖 Data Science, ML & Big Data with @DataXplore
Post #1825 524
An open project aimed at making 100 million scientific articles accessible through LLM-generated structured summaries.

🧠 A dataset of 100,000 summaries

🧩 Two fine-tuned LLM models for analysis and structuring

🌐 A 3D visualizer that shows the relationships between scientific papers

Models & Visualizer

••••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1824 168
MIRA by #ByteDance

Multimodal Imagination for Reasoning Assessment, a test that measures Reasoning models need not only words but also images

🟠 What is the idea?
- Where text doesn't help, images sharply improve the model's thinking.
- If the model is given drawings of intermediate steps, accuracy on average increases by 33.7%.
- The benchmark includes 546 tasks in 20 categories where you need to "see" rather than just read: blocks, mirrors, trajectories, forces, etc.

How the test is structured?
- direct question
- reasoning with text
- reasoning with visual steps (sketches)


🟢 What was found?
- Text alone often worsens performance because words poorly describe space.
- If the model is given images, the result improves significantly, especially in the exact sciences.

In the benchmark: 546 tasks in geometry, physics, logical puzzles, and causal relations.

Testing modes:
• Direct - model answers directly
• Text-CoT - text chain-of-thought
• Visual-CoT - model reasons through drawings and visual steps

• No model exceeded 20% accuracy in Direct mode (GPT-5 ~16.5%)
• Text-CoT often worsens results (e.g., −18% for Gemini 2.5 Pro)
• Visual-CoT gives an average gain of +33.7%, especially noticeable in physics tasks


Models need a *visual way of thinking*.
They need to be able to read simple diagrams, understand them and use them in reasoning, otherwise many tasks simply remain unsolvable.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1823 167
LMMs-Engine

A powerful, lightweight and flexible framework for training multimodal models.

🟢 Key features:
- Support for 19+ architectures, including models for processing text, images, and video.
- Optimizations for distributed training and reduced memory consumption.
- Convenient launch examples for various models.
- Providing high efficiency and ease of use.


GitHub

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1822 177
GEM model by META, a hybrid transformer-based recommendation systems (naturally inspired by LLM).

Jumps are very significant: Model already led to a noticeable increase in Ad conversions: +5% on Insta, +3% on FB in second quarter.

🟢 What's Inside Model?
Three main technical features:
1️⃣ Input data is divided into two groups: sequential features (user action histories, clicks, views, etc.) and non-sequential (location, age, ad properties, etc.). To avoid mixing them all together and blurring signals, a so-called InterFormer with dynamic alternation is used. First, event sequences are processed by a custom transformer block, then a layer combines these outputs with static features through cross-feature interaction blocks, after which the cycle continues at the next level.

2️⃣ Additionally, we need to consider connections between features from the two groups. For this, there is a separate component called Wukong. It consists of stacked factorization machines that search for non-obvious connections between features (why the user behaved one way or another).

3️⃣ For long sequences (i.e., long user histories), a proprietary pyramidal parallel structure is applied. This is needed to avoid the notorious exponential growth in costs as sequence length increases. The entire chain is first split into smaller parts -> they are processed -> the results form the next level of embeddings -> they are split into pieces again and processed -> and so on until everything finally collapses.


As a result, get:
(a) scalability;
(b) the ability to effectively consider all features and their connections;
(c) adequate model behavior on long sequences. And conversion jumps works quite well.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1820 517
Basic database concepts you should learn if you really want to understand scaling:

SAVE this

B+ trees
LSM trees
Write-Ahead Logging (logging before writing)
Two-phase commit
Three-phase commit
Read-only replicas
Leader-follower replication
Sharding / partitioning
Query caching
Secondary indexes
Vector indexes (Vector Indexes, e.g., FAISS, HNSW)
Distributed JOINs
Materialized views
Event Sourcing (event-based state storage)
Change Data Capture (tracking data changes)


••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1819 206
Semantic Search improves code retention and reduce dissatisfied user requests.

🔴 Why old approach limited?
Cursor has long used a retrieval mechanism: the agent searches the codebase and adds the necessary pieces to the LLM context.

Previously, it was just a grep version, searching by string matching. This is fast but not always sufficiently relevant.

Now it has been replaced by a smarter semantic search. Essentially, RAG. That is, the relevance of code snippets is now evaluated by a special vector model that no longer just searches by keywords but matches meanings.


🟢 How was embedding model trained?
Cursor trained its own embedding model specifically tailored for code. Real agent work trajectories were used for this. Each session is a sequence: query -> search for relevant code snippets -> result. A separate LLM evaluated these trajectories to determine which found snippets were ultimately useful and which were noise.

Then we take our vector model and train it on triplets (query, relevant files, irrelevant files) so that its ranking matches the LLM's ranking, meaning more useful snippets are closer to the query in vector space.

By the way, grep search still remains somewhere: for example, it is indispensable when you need to quickly search by variable or function names. The results of the grep module and the vector model are combined.

What about the metrics in the end:

1️⃣ On offline evaluation on the specially collected "Cursor Context Bench" benchmark, the average accuracy improvement was about 12.5%.

2️⃣ In A/B tests, code retention increased on average by about 0.3%. This metric shows how much code generated by the agent remains in the user's project over time. On large codebases, there was even a +2.6% increase.

3️⃣ Also, the number of dissatisfied follow-up requests decreased by about 2.2% – when the user has to make corrections or additional requests because the agent failed on the first try.


The effect is not huge because not every query requires a search, but it exists and will be especially noticeable in large codebases.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1818 211
Google has released a new 50-page document on how to create AI agents that actually work in practical tasks

This is a clear and structured introduction to the basics of agent systems.

🟢 What it Covers:
- agent architecture and its main components
- the role of LLM as the "brain" of the agent
- connecting and using tools
- orchestration of multiple agents
- approaches to deployment and production integration
- metrics and methods for evaluating performance
- how self-learning and evolving agents are created
- example architecture AlphaEvolve


Guide

#AI #Agents #Google #LLM #MachineLearning #AIResearch

🤖 Data Science, ML & Big Data with @DataXplore
Post #1815 237
NESTED LEARNING : new ML paradigm that lets models learn continuously LIKE HUMANS without forgetting past knowledge.

Instead replacing old information, it integrates new learning inside what’s already known, like a layer within a layer. The model retains prior skills, adapts to new tasks, and distinguishes the context it’s operating in.

🟢 How Nested Learning works?
1️⃣ The authors formalize the model as a set of optimization problems: each has its own stream of information it learns from and its own update frequency. For example, components with a high update frequency are responsible for adapting to the current context, those with a low frequency handle some basic knowledge, and so on.

2️⃣ But the model won’t just magically know what and when to update. Therefore, the authors propose making the optimizer itself trainable. That is, the algorithm responsible for updating weights stops being just a formula and turns into a neural network itself. This is called Deep Optimizers.

3️⃣ Formally, the optimizer is considered associative memory that learns to associate gradients with the correct weight changes. In this sense, conventional SGD or Adam are the simplest special cases.

Google extended their old TITAN architecture (which handled short and long memory) into HOPE, supporting unlimited levels of in-context learning.


Resulting Lower(↓) perplexity, higher(↑) accuracy in common-sense reasoning, and stronger long-context memory.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1811 230
How to effectively train a model even on consumer GPUs?

Two sequence parallelism methods are combined In ModelScope SWIFT.

🟢 Ulysses + Ring Attention
✅ Ulysses - splits attention by heads, uses almost no traffic (but limited by the number of heads)

✅ Ring Attention - scales beyond the number of heads through ring P2P communications, with "zig-zag" balancing for causal models

Ulysses works first, and only when it can no longer handle the load (e.g., GQA or cluster >8 GPUs), Ring is activated.

Sequence split is built directly into model's forward-hook : no data hacks, full compatibility with FlashAttention.

Result on Qwen2.5-3B with 65k tokens:
75.4 GiB → 17.9 GiB VRAM on 8× A100
Works with SFT, DPO, GRPO, multimodality, and padding-free inputs.


Enabled with a single flag command:
--sequence_parallel_size 8

More details on GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Post #1810 442
This repository with tutorials on AI agents recently surpassed 75 thousand stars on GitHub

It contains one of the largest collections of real AI applications to learn from: AI agents, RAG applications, voice assistants, multi-agent systems, chatbots with memory, and much more.

Grab on GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →