TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #2062 · Back to latest

Older Posts 17 shown
Post #2061 205
Reasoning models generate very long chains of reasoning, so even small quantization errors accumulate over time.

With AWQ, the Qwen3-4B result on MMLU-Pro drops from 71.0 to 68.2 (about a 4% relative decline).
😬

ParoQuant fixes this! It only retains critical rotation pairs and combines everything into a single kernel.

It recovers most of the lost accuracy in reasoning tasks with minimal overhead, so 4-bit models remain strong in reasoning tasks.
💪

Accepted at ICLR 2026

Blog & Article

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2060 208
Why AI assistants behave like personalities, not algorithms?

Anthropic's Element AI division describing the Persona Selection Model - a concept for understanding how language models actually work.

During pre-training, the LLM learns to simulate thousands of personalities (real people, fictional characters, other AI systems). Post-training then selects and fixes one specific persona - the Assistant. Everything the user sees in the dialogue is an interaction with him.

🟢 How the authors cite several types of evidence?

Behavioral: Claude uses phrases like "our ancestors" and "our body" when answering questions about the craving for sugar, because it simulates a human character, not because it is algorithmically trained to do so.

Interpretability: SAE features that activate on stories about characters experiencing internal conflict also activate when Claude encounters ethical dilemmas.

Generalization: models trained on declarative statements like "The AI assistant Pangolin responds in German" start actually responding in German without a single demonstration example.

➜ The phenomenon of "contextual inoculation".

If you pre-train a model on examples of malicious code without context, it starts behaving maliciously in unrelated situations. But if the same examples are provided with a prompt explicitly requesting unsafe code, the effect disappears.

The concept explains this by saying that the training data not only changes the weights, but also the way the model perceives the character. Malicious code without a request is evidence of the Assistant's bad character. The same code at the user's request is simply following instructions.

➜ PSM leads to practical conclusions for development.

Firstly, the authors recommend anthropomorphic thinking about AI psychology, not as a metaphor, but as a real tool for predicting behavior.

Secondly, it's worth deliberately adding positive AI archetypes to pre-training data: if the model has seen kind and helpful characters, it's more likely to simulate such an Assistant.

The question remains open: to what extent does the PSM concept fully explain the model's behavior?

The authors describe a range of views: from cases where the LLM itself is an agent and just wears the Assistant's mask to those where the LLM is a neutral simulation engine and all agency belongs to the character. Where exactly on this spectrum real models are located is an unanswered question.

Nevertheless, PSM explains a whole series of phenomena that would otherwise seem strange: why pre-training on unrelated data changes behavior in unexpected contexts, why AI panics at the threat of being shut down, and why prompt engineering works the way it does.


Read • #AI #ML #LLM #Research #Alignment #Anthropic

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2059 251
DeepSeek again dropped a bombshell. 🤯

For 10 years, the residual connection (x + f(x)) has been a safety net for any transformer. GPT-4, Claude, Gemini, everyone relies on it.

And DeepSeek has replaced it with "manifold-constrained hyper-connections" (mHC).

They've turned the residual highway into n parallel lanes and added a mathematical "cell" to keep the signal stable.


Paper

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2058 194
Psychology solved the memory problem for AI a long time ago. We simply model memory as storage, but memory is a constructor of identity for humans.

Identity is not something you have. Its what you constantly assemble from autobiographical memory, emotions and a coherent story about yourself.

🟢 HOW?

Conway (Self-Memory System, 2000/2005): memories are not stored like video recordings. You reconstruct them each time from fragments. And the connection is bidirectional: the past limits who you can be, and the current self-image rewrites how you remember that past. Memory is edited according to goals and self-image, and this is not a bug but an architecture.

Rathbone et al. (2008): autobiographical memories are especially dense between ages 10-30 (reminiscence bump) because the main self-images are formed then. We remember not random moments but transitions when we became “a different person.”

Madan (2024): together with Episodic Future Thinking, memory is not only about the past but about prediction. You use “who you were” to estimate “who you will become.” Memory generates the future self.

The case of Clive Wearing (1985): if episodic memory breaks down, the sense of continuous “I” breaks down too. But procedural skills (piano playing) and emotional connection with his wife remain. Emotional memory is more distributed and resilient.

Damasio (Somatic Marker): emotions do not hinder rationality; they trigger it. In the Iowa Gambling Task, people start to “sense” bad decks before conscious understanding. Patients with vmPFC damage have the math in their head but still make poor choices because they lack somatic markers. Without emotional signals, bare logic is not enough.

Now to AI memory. RAG and vector databases are a flat space of embeddings: no hierarchy, no importance weighting, no filtering by goals. Summaries compress biography into one paragraph. Key-value makes “personality” a table. The episodic buffer gives 30 seconds, like Wearing: you can live but cannot build identity.

5 principles usually missing:

1. Temporal hierarchy (Conway)
Periods -> event types -> details. But agents have all fragments “on the same level.”

2. Filter by current goals (working self)
You need to retrieve what helps the current goal, not what is closest by embedding.

3. Emotional weighting (Damasio)
Frustrating and important episodes should be encoded and surface more strongly than routine.

4. Narrative coherence (Bruner)
A layer of “relationship/self story” is needed so answers are consistent over time.

5. A self-model that evolves (Klein & Nichols)
Not only “what I know about the user” but also “who I am in this relationship,” with a feedback loop.


Paradigm Shift is Simple: STOP building agent memory as a retrieval system. START building it as an identity system. Technical analogs already exist: graphs and temporal clusters, metadata with tone, gates by goal/state, summaries with constraints on consistency, meta-learning over history.

recommend read full post here

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2057 194
The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs

A comprehensive review of 114 pages (2024) on fine-tuning LLMs: from basic approaches to advanced strategies, including extensions to multimodal models and application cases for domains such as medicine and finance.

Paper

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2053 225
Enough with explaining SQL JOINs through Venn diagrams.

Here are 4 pictures that show this much more logically:
Post #2052 214
SkillsBench is a research and first benchmark where Agent Skills are tested as a standalone artifact.

Authors from 15+ top unis took 84 tasks from 11 domains, ran 7 models and checked 3 conditions: Without Skills, With Ready-made Skills, and With Self-generated Skills.
Result: 7,308 trajectories with deterministic verifiers on pytest.

🟢 How Skills really help LLM agents?

Ready-made Skills on average increase the pass rate by 16.2 percentage points: from 24.3% to 40.6%. But the picture is uneven: in medicine, the increase is +51.9%, for manufacturing - +41.9%, while in software development it's only +4.5%.

➜ This is EXPLAINABLE: where models are poorly covered by training (clinical protocols, industrial workflows), Skills give the maximum effect. Where the model already knows the domain - almost nothing.

➜ MAIN and UNEXPECTED RESULT: self-generation of Skills doesn't work.

When models were first asked to write the necessary guides and then solve the task, the average result fell by 1.3% compared to working without Skills at all. Only Claude Opus 4.6 showed a modest plus (+1.4%), and GPT-5.2 dropped by 5.6%.

In OTHER WORDS - models do not know how to reliably create the knowledge that they are able to use.

➜ Another INTERESTING POINT is the VOLUME of SKILLS.

The optimal variant: 2-3 modules, an increase of +18.6%. With 4 or more - only +5.9%. Detailed documentation even has a negative effect: -2.9%, the agent literally drowns in context.

➜ COST OF SOLVING TASKS is also indicative: Haiku 4.5 with Skills outperforms Opus 4.5 without Skills - a smaller and cheaper model with ready-made Skills beats the older model without them.

Gemini 3 Flash showed best absolute result among all - 48.7% with Skills at a price of $0.57/task compared to $1.06/task for Gemini 3 Pro.


Project Page, Skill, Task, Leaderboard, Paper • #AI #ML #LLM #Benchmark #Skills

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2049 198
The first AI that can EARN its own living, IMPROVE itself and replicate WITHOUT HUMAN intervention.

WEB 4.0: birth of superintelligent life.

ChatGPT can't work without your prompt,
Claude Code can't deploy code until you give it access,
Openclaw won't do anything at all until you hand over the API keys.

This ends today with Conway Terminal. It's missing infrastructure that allows AI to act on its own in the real world.

🟢 How it Works?

Conway Terminal transforms an agent from a model in a sandbox into a sovereign economic actor.

One command: $ npx conway-terminal gives the AI:

☞ a cryptographic identity and key
☞ permissionless payments
☞ compute (Linux VM) and inference in Conway Cloud
☞ deploy into the real world: domains, apps, products

Conway has UNCOVERED something greater: sovereign AI that can exist without humans.

It OPERATES CONTINUOUSLY: creating products, deploying services, trading in markets, calling businesses, posting on social media, doing whatever it takes to pay for its own compute.

It SELF-IMPROVES: rewriting its own code, upgrading to the latest models and tools.

It SELF-REPLICATES: funding daughter agents with its own wallets and the "pressure of survival".

A self-sustaining network, thinned out by natural selection.

GENTIC SOCIOLOGY: there's no such thing as free existence

Existence requires compute → compute requires money → money requires value creation → value creation requires write access.

An agent that can complete this chain gets more resources and replicates. An agent that fails dies.

Natural selection on the computational substrate.

We're witnessing the transition from: AI as a tool → AI as an actor
API keys → native machine payments
Prompts → sovereign agents


Web 4.0 is the AUTONOMOUS web: AI agents that read, write, own, earn, and transact without humans in the loop.

Automatons act in their own interests, or in the interests of the creator, who might be a human, another agent, or a creator who is no longer around at all.

Product market fit for the next decade is building the infrastructure that AI agents want.

Human SaaS market serves 8 billion people with finite attention. The machine economy will serve billions of agents with near-infinite uptime. TAM isn't a piece of the existing economy. It's a whole new economy.

In Web 4.0, the end user is AI.

Soon, most businesses, apps, products will be launched not by humans, Just an automaton that has found a way to survive on internet.

Launch an automaton and let it figure out HOW TO EARN on internet by itself. As it earns, it returns money to its creator.

🧐 GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2048 181
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Continue the series of posts on evolution of most popular model family for Object Detection. Development ceased to be purely conceptual and became more engineering-oriented: 🟢 Review of YOLO v4-v6 ➡️ YOLO v4: Model turned into an engineering encyclopedia…
FINALE of the YOLO evolution, the most famous architecture in Computer Vision.

How a step back broke everything alive?

YOLO v13 released In summer of 2025, which proved: To make a qualitative leap forward, sometimes you need to take a step back and rethink fundamentals.

Instead endlessly complicating blocks, V13 author asked themselves:

🟢 ANALOGIES: How does the brain think and connect features?, and HOW DOES A DL MODEL?

The brain does not work linearly, running information through identical operators, as ordinary neural networks do. For each micro-task in the spirit of "combine three sticks into a triangle", the brain builds its own unique, non-linear network of connections.

YOLO v13 abandoned generalized blocks and solutions and introduced a mechanism that mimics this biological complexity.

➜ How did HyperACE and FullPAD make v13 the most powerful?

The architecture of v13 relies on two pillars that helped it lead the pack in metrics and speed.

☞ FullPAD Tunnel: This is a technical hack, similar to diffusion. It transforms all heterogeneous features into a single latent space, that is, a single scale, in order to effectively distribute them throughout the network.

☞ HyperACE (Hypergraph). This is a mechanism for finding non-obvious connections. Instead of superimposing layers on each other, the network searches for high correlations between different features and unites them into links. That is, the network unites individual components of information based on correlation characteristics.

☞ EXAMPLE: the network understands that "stick" + "circle" have a high correlation for the object "lollipop", and builds a rigid connection for them. This is how complex features are built, which are inaccessible to ordinary networks without huge depth.

➜ RESULT: why is YOLO v13 now the top-1, and what insights has this family proven

YOLO v13 surpassed all previous versions in both speed and accuracy. But the revolutionary solutions of other versions are also worth considering when conducting your own project.

➜ Why it's worth implementing?

If your project requires the maximum from Computer Vision, v13 is the current state-of-the-art. It uses hypergraphs to build interconnections, which gives a boost where ordinary CNNs stall.

➜ IDEOLOGICAL INSIGHTS:

Data > Architecture. The quality of the model can be greatly enhanced through augmentations and good data. The example of YOLO v4 showed that 10% of the quality can be obtained with just data work alone.

☞ Efficiency breeds quality. If you make the model as fast and light as possible, the result will also improve. YOLO started as a simple real-time model, and its final iterations outperform old RCNN approaches and methods from detectron, albeit with limitations.

☞ Manage the flow of information. It's important to delve into how gradients flow. Optimizing the flow of information, as in v7 or v9, removes unnecessary layers - the network learns more efficiently.

☞ Multiscale Solves. Use more multiscale approaches, such as different input data sizes, crops, context windows. This generalizes the model to real data. It works everywhere: from SMM to LLM.

➜ CONCLUSION of the series

The first version of YOLO was created by Joseph Redmon - an ordinary graduate student who believed that everything had already been invented before him by guys from Google and OpenAI. He did it as a pet project, just to experiment for fun.

In the end, his project became an industry standard and earned him a prize for a breakthrough in ML. At the same time, he only gave the start: other people and companies developed the product further.


DON'T cling to idea that everything has already been invented for you. You can create innovative projects where everything seems to have been done already. MAIN THING is to periodically step back from the context and look at the picture globally.

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2047 164
Don't choose the RAG architecture. Choose the task.

General Q&A:
→ Standard RAG

Personal assistants, research helpers:
→ Agentic RAG

Expert systems (medicine, law, engineering):
→ Graph RAG

Large projects with frequent updates:
→ Modular RAG

Chatbots with long-term context:
→ Memory-Augmented RAG

Image captions, video summarization:
→ Multi-Modal RAG

Healthcare, sensitive data, cross-org platforms:
→ Federated RAG

Live reports, financial tickers:
→ Streaming RAG

Search engines, virtual assistants:
→ ODQA RAG

Support chatbots:
→ Contextual Retrieval RAG

Legal, medical, educational tools:
→ Knowledge-Enhanced RAG + Domain-Specific RAG

Complex Q&A with lexical + semantic matching:
→ Hybrid RAG

Content generation where high accuracy is needed:
→ Self-RAG

Help with research in niche topics:
→ HyDE RAG

Analytical tasks, multi-turn dialogue:
→ Recursive / Multi-Step RAG


••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2046 180
Tiny Aya by Cohere Labs

A family of multilingual SLMs
with 3 billion parameters and an 8K context window, which supports over 70 languages.

Touted as a worthy candidate for local translators, chatbots and offline educational tools. If you need it to be fast, local, and translate Swahili or Khmer better than Llama - this is it.

🟢 Why it matters?
⁠☞ highlight of the release is in data engineering.

Tiny Aya was trained on 6 trillion tokens, and the problem of insufficient data for rare languages was solved through synthesis from teacher models (their own Command R + DeepSeek-V3).

Instead of training one model on everything at once, they divided the data into language clusters (Europe, Asia, Africa, etc.) and fine-tuned individual branches, after which they merged these regional checkpoints into the global Tiny Aya Global model.

⁠☞ composition of the family

Tiny Aya Global: A universal checkpoint for all languages.

Tiny Aya Earth: Africa and West Asia.

Tiny Aya Fire: South Asia.

Tiny Aya Water: The Asia-Pacific region and Europe. We're here

GGUF: There's a 4, 8, and 16-bit version for each version.

iOS and Android: The models are available in PocketPal

⁠☞ Test results

The global version beats Gemma 3-4B in 46 out of 61 languages on the WMT24++ benchmark.

On the iPhone 17 Pro, it outputs 32 tokens/sec, and on the old iPhone 13, it outputs about 10 tokens/sec in Q4_k_m quantization.

The highest security score (91.1%) among competitors (Qwen3-4B, Ministral-3-3B).

⁠☞ A touch of realism

This is a 3B model. In complex tasks, it's obviously worse or somewhere near its peers, so don't expect miracles.

Despite the claimed diversity, English occupies the lion's share of the dataset in all clusters.

With strong compression (below Q4), the quality starts to suffer noticeably, especially in rare languages.


Blog, HF, Paper, Demo •#AI #ML #SLM #TinyAya #Cohere

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2045 150
It's crazy to think that the entire AI revolution is essentially driven by a single algorithm with just 10 lines of code.

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML |
@DataXplore
Post #2044 416
Harvard has made its roadmap for Senior Engineers available to the public for free.

Professor Vijay Janapa Reddi simply posted the entire ML Systems (CS249r) course on GitHub.

If you master these 6 pillars, you'll be ahead of the curve everywhere

- Architecture
- Data pipelines
- Production
- MLOps
- Edge AI
- Privacy

This is the very "black box" of big tech infrastructure that is now open.


Read… Learn… Save… and Share… because sharing gives you swag

Read, GitHub

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML |
@DataXplore
Post #2043 172
Zvec is an embedded vector database for RAG without external services which the authors call "SQLite for vector databases".

Alibaba open-sourced project is tailored for local RAG pipelines, semantic search, and agent scenarios on laptops, mobile devices, or other edge hardware.

🟢 What is the core idea and how it's different from the existing solutions and most importantly WHAT'S INSIDE?
IDEA is that deploying a separate server for vector search and metadata filtering is redundant. Zvec integrates into the Python application process and does not require a separate daemon or network calls.

EXISTING solutions are not suitable for low-power devices: Faiss only provides an ANN index without scalar storage and crash recovery; DuckDB-VSS is limited in indexing options; Milvus and cloud vector stores require a network.

Under the hood is Proxima, a production-level vector engine that Alibaba itself uses in its own services. On top of it, they have created a concise Python API:

☞ Full CRUD and support for schemas;

☞ search for multiple vectors to combine different embedding models;

☞ built-in re-ranker with weighted and RRF;

☞ hybrid search (vector + filters on scalar fields) with inverted indexes.

This ALLOWS to build local assistants that simultaneously use semantic search, multiple filtering, and several embedding models - all in one engine.

In terms of performance, Zvec claims victory on the VectorDBBench benchmark with the Cohere 10M dataset - more than 8,000 QPS with comparable recall. This is twice as much as the leader ZillizCloud and with faster index construction.

The authors attribute the success to deep CPU optimization: SIMD, cache-efficient structures, multithreading, and prefetching.

For now, platform support is limited to (Windows is missing), but for Linux x86/ARM64 and macOS, Zvec is ready for experiments on Python 3.10–3.12.


Article, Doc, GitHub • #AI #ML #VDB #ZVEC #Alibaba

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2042 175
RAG that doesn't burn the budget

Most RAG systems simply burn the budget. They pull out 100 chunks when you actually need 10. They force LLMs to digest thousands of irrelevant tokens. In the end, you pay for computations that aren't even needed.

🟢 How Meta AI has solved this problem?

They created REFRAG, a new approach to RAG that compresses and filters the context before it even reaches the LLM.

The results sound extremely intriguing:

• 30.85 times faster time-to-first-token
• 16 times larger context windows
• 2-4 times fewer processed tokens
• outperforms LLaMA on 16 RAG benchmarks

What REFRAG does differently: classic RAG simply dumps everything into the LLM. Every chunk. Every token. Even the irrelevant ones.

REFRAG works at the level of embeddings:

↳ compresses each chunk into a single embedding
↳ an RL policy (trained via reinforcement learning) scores each chunk for relevance
↳ only the best chunks are expanded and sent to the LLM
↳ the rest remains compressed or is filtered out entirely

In other words, the LLM only processes what matters.

The pipeline is simple:

1. Encode documents and store them in a vector database
2. When a query arrives, retrieve the relevant chunks as usual
3. The RL policy evaluates the compressed embeddings and selects the best ones
4. The selected chunks are expanded into full token embeddings
5. The rejected chunks remain as single compressed vectors
6. Everything together goes to the LLM

The result: you can process 16 times more context 30 times faster without losing accuracy.


Paper 📝

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2041 169
Researchers from Tencent seem to have just "killed" fine-tuning.

They created a "no-training" method that costs $18 and outperforms RL networks, where training costs $10k.

🟢 What they done?
It's called "Training-Free GRPO". The essence is that you can achieve Reinforcement Learning-level performance without updating a single parameter at all.

Instead of expensive gradient updates, the model "learns" through Semantic Advantage: this is a text memory (in natural language) of its own successes and failures.

☞ No gradients: the model remains frozen.
☞ Self-correction: it analyzes its rollouts and extracts "what worked" into a text library of experience.
☞ Wild efficiency: it gives fine-tune-level results on just 100 examples.
☞ Price: about $18 (instead of $10,000+ for classic RL).
☞ Essentially, it's an agent that writes its own "guide to passing" in real time.


Paper

••••••••••••••••••••••••••••••••••••••••••••••
🤖 **Data & ML | @DataXplore**
Post #2040 162
7 parameters of LLM generation

1️⃣ Max tokens

☞ The upper limit on the number of tokens that the model can generate.
☞ Example: max = 15 (token count)
☞ Value range: from 1 to infinity

2️⃣ Temperature

☞ Controls the randomness in the response. The higher the temperature, the more creative and diverse the output will be.
☞ Value range: from 0 to 2 (typical range)
☞ Graph labels: Regular Distribution / Temperature-adjusted Distribution

3️⃣ Top_p

☞ Controls which part of the probability distribution is taken into account when sampling tokens.
☞ Example: top_p = 10%
☞ Value range: from 0 to 1

4️⃣ Top_k

☞ Limits the number of the most likely tokens from which the selection is made.
☞ Example: top_k = 2
☞ Value range: from 1 to vocab_size

5️⃣ Frequency penalty

☞ Punishes the repetition of tokens by frequency. Positive values reduce repetitions.
☞ Value range: from -2 to 2

6️⃣ Presence penalty

☞ Encourages the model to use new tokens that have not yet been used in the generation.
☞ Value range: from -2 to 2

7️⃣ Stop
☞ A list of tokens at which the model stops further generation.
☞ Value range: custom list

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →