TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
583
Photos
845
Videos
446
Links
675

Showing posts older than #1972 · Back to latest

Older Posts 20 shown
Post #1971
Live stream finished (43 minutes)
Post #1970
Live stream started
Post #1969 789
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Live stream scheduled for Jan 10 at 16:30
Voicechat 2 of Saturday series with industry pros.

Q/A with QA Tester

📅 Time: TODAY, Jan 10, 2026 | 4:30 PM UTC (10 PM IST)
🎧 Listen Only: Join Livestream

🎤 Want to Speak or Ask?
Comment for Speaker Link

#DataXplore #AI #ML #VoiceChat

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Telegram Data eXplore — Data Science, ML, Big Data, LLMs and AI Exploring Data Science, Big Data Analytics and Visualization, Machine Learning, Deep Learning, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace. Not just data, but science behind data. Paid project? @ipremodi premodi@zohomail.in ★ @ITXplore
Post #1968
Live stream scheduled for Jan 10 at 16:30
Post #1967 201
Binary search + rescoring in int8.

The simple strategy to search through 40 million texts in ~200 ms using only a CPU server, 8GB of RAM, and 45GB of disk space.

If you want to try it out immediately, there's a demo of 40 million texts from Wikipedia. No login or other hassles required.

🟢 The Inference Strategy:

☞ Embed the query with a dense model into a regular fp32 vector

☞ Quantize the fp32 embedding into a binary format, which is 32 times smaller

☞ Retrieve, for example, 40 documents (about 20 times faster than an fp32 index) using an approximate or exact binary index

☞ Load the int8 embeddings for these top-40 documents from the disk

☞ Rescoring: fp32 embedding of the query × 40 int8 embeddings

☞ Sort these 40 documents by the new score and take the top-10

☞ Load the titles and texts of the top-10 documents

The documents are embedded once, and then these embeddings are used in two representations:

1️⃣ A binary index (I used IndexBinaryFlat for exact search and IndexBinaryIVF for approximate search)

2️⃣ An int8 view, i.e., a way to quickly read int8 embeddings from the disk by document ID

➡️ In end, instead of fp32 embeddings, you store:

- a binary index (32 times smaller)
- int8 embeddings (4 times smaller)

Plus, Only the binary index is kept in memory, so the RAM savings are also x32 compared to fp32 search.
For comparison: a regular fp32 retrieval on such a task would require about 180GB of RAM, 180GB of disk for embeddings, and would be 20–25 times slower.

A binary retrieval with int8 rescoring fits into about 6GB of RAM and ~45GB of disk for embeddings.

For example, if you load 4 times more documents through the binary index and then rescoring them in int8, you can recover about 99% of the quality of fp32 search (compared to ~97% for pure binary search)


HFblog
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1966 158
On subject of progress: An agent from SakanaAI took a confident first place in a coding competition.

🟢 How did a "wrapper" beat 800 human experts?

Last year, an agent from OpenAI only took second place in the same competition.

This year, about 800 people participated in the AtCoder Heuristic Contest. The ALE-Agent from the Japanese laboratory outperformed everyone and took the top spot with a significant lead. The cost of solution was approximately £1,300.

Interestingly, authors of this year's optimization task themselves expected a classic approach using annealing and constructive heuristics, but the Sakana agent took a different path. He suddenly implemented the virtual power heuristic, which allowed him to escape local optima even better than human experts.

Agent is a rather clever wrapper over (in this case) GPT-5.2 high and Gemini 3 Pro high.


Sakana themselves never really shone in terms of models, but they learned to work competently with inference time scaling and here's the result.

In a word, well done! Read Here

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1965 182
Convert PDF files into clean data, ready for LLM.

Dolphin is a document parsing framework that converts PDFs into structured formats: Markdown, HTML, LaTeX, and JSON.

Works and Features:

Stage 1️⃣ Detailed analysis of the layout at the page level. Elements and their order are determined according to the natural reading order.

Stage 2️⃣ Parallel parsing of elements using different types of anchors and task-specific prompts.

Key features:

» Open-source
» A two-stage approach of analyze-then-parse based on a single VLM
» Encouraging performance on document parsing tasks
» Generation of a sequence of elements in the natural reading order
» Heterogeneous anchor prompts for different types of document elements
» An efficient parallel parsing mechanism


GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1964 180
SleepFM model diagnoses 130 diseases by analyzing a single night's sleep.

Stanford trained SleepFM model (fundamental for predicting a range of pathologies) from atrial fibrillation and myocardial infarction to dementia and Parkinson's disease.

🔴 Why traditional ML models fail despite having gigabytes of data?
Polysomnography is the "gold standard" for studying sleep: a person is fitted with sensors (EEG, ECG, respiration, muscles) and gigabytes of raw signals are recorded.

But in the ML world, this data is used ineffectively. Existing models were trained on small datasets for specific tasks (finding apnea, determining sleep phases).

A huge amount of physiological information about a patient's health was simply ignored, because it's impossible to manually label hundreds of hours of recordings for each disease.

Moreover, if the EEG sensor was mounted slightly differently in one clinic or fell off, the usual model would break down.


🟢 How training on 585k hours possible without human labels?
At the university, they realized that they didn't need human labelers, they needed volumes. They collected a huge dataset of 585,000 hours of sleep recordings from more than 65,000 patients and invented a unique SSL learning algorithm for the future model.

1️⃣ LOO-CL (Leave-One-Out Contrastive Learning)

Instead of teaching the model to predict a diagnosis, they made it solve a puzzle: the system receives input signals from 3 modalities (heart, muscles, respiration) and must predict the embedding of the fourth (brain waves).

This forces the neural network based on 1D CNN and Transformers to learn deep, hidden connections between physiological processes.

2️⃣ The second feature is Channel-Agnostic Attention.

The models don't care about which sensors are connected and in what order. If a channel fails or is absent, attention pooling simply redistributes weights, and inference continues.

3️⃣ SleepFM has learned to read sleep not just for insomnia.

Having received a single night of recordings as input, the model predicts the risk of 130 diseases, and it does this more accurately than specialized models trained with a teacher: the risk of Parkinson's disease is detected in 89% of cases, dementia in 85%, and the probability of a heart attack in 81%.


Such diagnostics could move from labs to smartwatches with development of wearable electronics, and tests shown that noise of sleep signals can hide a patient's entire medical record.

Details #news #AI #ML

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1963 161
How to use LLM without losing quality?

DFlash is a way to speed up text generation for large models.

HOW and WHY to use?

Works like this: one model quickly creates a draft, and another one checks it and corrects errors.

- 6.2× faster without losing quality on Qwen3-8B
- 2.5 times faster than EAGLE-3

The idea is simple:

• Diffusion models - generate quickly, but sometimes make mistakes
• Autogenerative (AR) - very accurate, but work slowly
• DFlash combines both approaches:
diffusion - draft → AR - checking and confirmation


Both quickly and accurately, instead of choosing one or the other.

Blog, Code, Models

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1962 172
Cursor is completely switching to dynamic context for all models

Means the agent (based on any model) will now primarily collect context on its own, rather than using what has been provided.

🟢 How is Dynamic Context different from "Classic" approach?

Static context is a classic approach. You dump all the logs, documentation, chat history, descriptions of all toolboxes, MCP, etc. into the agent's context at once. In general, this works, but the context ends up being filled with a lot of irrelevant information and is constantly overflowing.

Now, Cursor is offering Dynamic context discovery. This involves placing a conditional "table of contents" and links in the context, while the rest is scattered across files, and the agent can add information to itself as needed. For example:

➖ Everyone remembers that when the context overflows, Cursor performs summarization and updates the window, right? Now, in addition to this, Cursor stores chat history as a file. After summarization, the agent receives a link to this file, and if some necessary detail was lost in the summary, he can search the history and supplement himself.

➖ Long responses from tool calls are now also recorded in files, rather than being sent directly into the context. Only a link to the necessary output appears in the context, while the gigantic JSON file sits waiting for the agent to access it and search for what he needs using conditional grep or tail.

➖ The same applies to MCP, Agent Skill, and terminal sessions. Bulky tool descriptions and terminal outputs are stored not in the context, but in files. The context simply says "MCP is available for jira, datadog, figma", and the agent, if he needs something, goes to the detailed description and invokes the tool.

It turns out to be quite nice and practical. On A/B tests, the overall token consumption has decreased by ~46.9%.


And it's also scalable, because here the context transforms from a place where knowledge is stored into an instruction on how to retrieve it. Find details

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1961 167
Microsoft has really turned the tables 🤯

They've long since released the open-source bitnet.cpp - a framework for inference of 1-bit LLMs.

It allows you to run models with 100B parameters directly on the local CPU, without any GPUs.

- inference is 6.17 times faster
- energy consumption on the CPU is 82.2% lower

And yes, it's 100% open source.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1960 393
Claude Code + Integration with Supabase

Step-by-step tutorial:
How to install and use 2 agents, 8 commands for the database, and an MCP server to automate development for Supabase in Claude Code.

Claude Code Stack for Supabase:

Claude Code Templates offers three ready-made components for integrating with Supabase:

1️⃣ Agents:

» Supabase Schema Architect
An expert in database design, migrations, and RLS policies. Automatically analyzes requirements and generates production-ready schemas with a focus on performance.

» Supabase Realtime Optimizer
A specialist in WebSocket optimization and real-time. Monitors connections, optimizes subscriptions, and helps maintain scalable real-time without degradation.

2️⃣ Commands:

» supabase-schema-sync
Synchronization of local and remote schemas, version control, and automatic detection of schema drift.

» supabase-migration-assistant
Generation, management, and application of database migrations with rollback support and conflict resolution.

» supabase-performance-optimizer
Analysis of query performance, index recommendations, and optimization of database operations for maximum speed.

» supabase-security-audit
Full security audit, validation of RLS policies, and vulnerability detection with automatic fixes.

» supabase-backup-manager
Scheduled backups, recovery procedures, and DR planning with testing.

» supabase-type-generator
Generation of TypeScript types from the database schema, support for type safety, and automatic updates when the schema changes.

» supabase-data-explorer
Interactive data viewing, visual query builder, and export with filters.

» supabase-realtime-monitor
Monitoring of real-time connections, performance tracking, and WebSocket diagnostics.


3️⃣ Supabase MCP Server. Direct integration with the Supabase API via MCP. Gives Claude Code native access to the project: executing commands, working with schemas and data, managing real-time, and security without unnecessary intermediaries.

Before installation, you can view all available Supabase components on the official Claude Code Templates website.

Go to aitmpl.com and find supabase


INSTALLATION OPTIONS:
There are several ways to install the Supabase stack for Claude Code. Choose the one that best suits you:

1️⃣ Installing individual components

# Install a specific agent
npx claude-code-templates@latest --agent database/supabase-schema-architect

# Install a specific Supabase command
npx claude-code-templates@latest --command database/supabase-schema-sync

# Install the MCP server
npx claude-code-templates@latest --mcp database/supabase


The components will be installed in:

* 📁.claude/commands/
* 📁.claude/agents/
* 📁.mcp.json

2️⃣ Creating global agents (available in any project)

# Create global agents available from any project
npx claude-code-templates@latest --create-agent database/supabase-schema-architect
npx claude-code-templates@latest --create-agent database/supabase-realtime-optimizer

# Show the list of all global agents
npx claude-code-templates@latest --list-agents

# Update a global agent
npx claude-code-templates@latest --update-agent database/supabase-schema-architect

# Remove a global agent
npx claude-code-templates@latest --remove-agent database/supabase-realtime-optimizer


••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore & @PromptXplore
Post #1959 170
12 new advanced types of RAG

▪️ Mindscape-Aware RAG (MiA-RAG)
▪️ Multi-step RAG with hypergraph-based memory
▪️ QuCo-RAG
▪️ HiFi-RAG
▪️ Bidirectional RAG
▪️ TV-RAG
▪️ MegaRAG
▪️ AffordanceRAG
▪️ Graph-O1
▪️ SignRAG
▪️ Hybrid RAG for multilingual question answering based on documents
▪️ RAGPart and RAGMask


••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1958 174
How MiniMax M2.1 made?

When they say that one model writes code better than another, they usually mean the SWE-Bench benchmark. The model gets a real bug from a real project on Github, which it has to read, find the error and fix it. This partially resembles a programmer's daily work.

🟢 How MiniMax-AI became a truly universal AI programmer?

SWE-Bench benchmark has its drawbacks. They found the answer and implemented it in their latest model M2.1.

1️⃣ LANGUAGE BARRIER

Problem: SWE-Bench only works with Python. In the real world, developers deal with Java, Go, TypeScript, Rust, C++ and a bunch of other languages.

⁠☞ Solution: SCALING THE ENVIRONMENT

Behind this vague term lies a huge system that operates with popular languages: JS, TS, Python, Java, Go, C++ and Rust.

For this, more than 100 thousand real tasks with a description of the problem, code and tests were collected from GitHub. This was not easy, as complex languages (Java or C++) require setup and each language has its own frameworks and dependency management systems.

To train the model on such a dataset, MiniMax built an infrastructure capable of running more than 5 thousand isolated execution environments in the shortest possible time - 10 seconds.

2️⃣ "BUG-FIX ONLY" TRAP
The Problem: Most benchmarks are is only about fixing bugs, while programmers also write new functions, refactor and optimize.

☞ The Solution: GOING BEYOND BUG FIXES:

MiniMax-M2.1 was also trained to generate tests, and it turned out that this is a critically important skill.

The previous version, M1, wrote too simple tests and often chose the wrong solutions. M2.1 excelled in this and equaled the results of the powerful competitor Claude Sonnet 4.5.

It also learned to optimize code performance - on SWE-Perf it showed an average increase in efficiency of 3.1%.

And finally, M2.1 was taught to do Code Review, for which an internal benchmark SWE-Review was created.

3️⃣ ENVIRONMENT DEPENDENCY
The Problem: A model's results strongly depend on the environment in which the model operates.

⁠☞ The Solution: GENERALIZATION ON OOD SCAFFOLDS.

The model should equally well follow long instructions and adapt to different ways of managing the context of the dialogue.

The team conducted tests in mini-swe-agent, Droid and Claude Code and if you look at the figures from their comparative table, you can see that the model has become much more flexible and versatile.

On the same SWE-Bench, when using Claude Code, MiniMax-M2.1 scored 74 points, which is higher than the model M2 with its 69.2 points, and almost on a par with Claude Sonnet 4.5 and DeepSeek V3.2.

On another test, OctoCodingBench, the gap is even greater: 26.1 for the new model against 13.3 for the old one.


Seems that concept of an "AI coder" is becoming more and more real. Success of MiniMax-M2.1 showed that it's no longer about writing individual lines of code, but about a comprehensive understanding of the entire development process.

#AI #ML #LLM #MiniMaх

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1957 166
Google trained the model on millions of user messages.

Without seeing a single message.

This is called federated learning. It's used by Google, Apple, Meta, and almost all major tech companies.

🟢 How It Works?

Imagine you want to make a keyboard that predicts the next input.

The best data for training is real messages from millions of phones.
But you can't collect them: privacy, sensitive data, users would just revolt.

Federated learning flips the approach.
Instead of dragging the data to the model, you drag the model to the data.

Here's how it works in practice?

Step 1️⃣. Send the model to devices
The phone downloads a small neural network. It lives locally, right on the device.

→ this is the global model W

Step 2️⃣. Train where the data lives
While you're typing, the phone quietly learns your patterns.
"omw" → "I'll be there in 10 minutes".

It calculates how the model should improve.

→ these are local gradients ΔW

Step 3️⃣. Send only the updates, not the data
The server receives updates to the weights.
Not the messages. Not the input history. Just the math.

→ the update aggregation stage

Step 4️⃣. Average across thousands of devices
The server combines updates from thousands of phones.
Common patterns are amplified, individual peculiarities cancel each other out.

→ the classic FedAvg
W_new = W + (1 / n) × Σ(ΔWₖ)

Four steps.
Not a single raw user data leaves the device.
Just careful coordination.

The most important thing:
this opens access to data that was previously fundamentally inaccessible.

Hospitals can train models for cancer diagnosis without sharing patient images.
Banks build anti-fraud systems without disclosing transactions.
Smart homes learn preferences without sending personal moments to the cloud.


Privacy and utility aren't mutually exclusive.
On the contrary: respecting data boundaries makes such models possible.

So before centralizing everything, it's worth considering: the best data for training already exists. it's just locked on devices you'll never get direct access to.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1956 372
What if you could define your own entity types without training a separate model?

Named Entity Recognition (NER) extracts key data from text, such as names, dates, and organizations. But standard models are typically predefined with a fixed set of types, like PERSON, ORG, DATE, etc.

If you need to extract something more specific, you usually have to train your own model on thousands of labeled examples.

GLiNER solves this with zero-shot entity extraction: you can extract any type without training.

Advantages:
• Works immediately on any text domain, without preparation
• Supports multiple entity types in a single pass
• Returns confidence for each found entity
• Easily integrates into spaCy and other NLP pipelines


Full article, Run this code

In addition, GLiNER is open-source! Install it with command pip install gliner.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1955 182
NewBieAI-Lab has introduced NewBie-image-Exp0.1 - an open 3.5B DiT model, specifically designed for high-precision and fast anime generation.

Key features:
✅ 3.5B parameters - works even on 8GB VRAM (RTX 4060)
✅ Inside: Gemma-3-4B-it + Jina CLIP v2 for deep understanding of prompts
✅ Structured XML prompts: full control over characters without random clothing changes
✅ FLUX.1-dev 16-ch VAE - soft skin, fabric and metal textures
✅ Inference in ~20 steps, LoRA support, Apache-2.0 license + non-commercial use
✅ Trained on over 10M anime images with XML annotations - confidently handles multi-character scenesb


⚡ Up to 40 percent faster than models >8B and reliably handles prompts up to 500 characters in length.


🧠 Bonus: Noise → Context Refiner pipeline eliminates the classic DiT problem - "the image is beautiful, but the prompt is ignored".

Model

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1954 192
IQuest-Coder-V1 surpasses Claude Sonnet 4.5 & GPT-5.1.

Despite having significantly fewer
, 40-billion-parameter model with a context window of 128K tokens, Achieves:
81.4% on SWE-Bench Verified,
49.9% on BigCodeBench,
81.1% on LiveCodeBench v6.


🟢 How it works?

The model uses the "code-flow" technique - learning from the evolution of repositories and commits, and is divided into 2 branches.

Dense Models: Base and Instruct versions for fine-tuning and following instructions

Loop Models: an optimized version with maximum VRAM efficiency (int4 can run on 3090\4090)

LoopCoder architecture uses a cyclic transformer structure, where the same model parameters are used in 2 consecutive data processing passes.

In the first pass, the model processes embeddings through its layers, taking into account the positions of words.

In the second pass, the model simultaneously uses two types of attention: global attention, which refers to all information from the first pass to understand the general context, and local attention, which only looks at the previous words in the second pass to preserve the text sequence.

Both types of attention are combined using a mechanism that decides how much weight to give to the global context and how much to the local sequence.


Modified MIT License, Project page, Technical report, Model set, GitHub

#AI #ML #LLM #IQuest #QuestResearch

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with
@DataXplore
Post #1950 203
Cursor speeds up development by 3–4x - but at a cost

Cursor provides a real increase in productivity, especially at the beginning.
But those who combine AI with:
- tests
- code reviews
- quality gates
- statistical analysis

How they analysed?
Scientists from Carnegie Mellon studied 807 repositories, where developers switched to Cursor
(via configs like .cursorrules) and compared them with 1380 control projects - before and after implementation.


The Difference-in-differences method: they compared same repositories *before/after*, plus controlled trends by months.

🚀 What happened to "code velocity"
Code Velocity = commits + lines of code.

- in first month - a jump of 3–5x in lines
- on average after implementation - +1.84x to the velocity

AI really speeds up work - and this is measurable, not just a feeling.

🧩 But there are side effects
Quality was assessed via SonarQube
(reliability, maintainability, security, duplicates, cognitive complexity).

- static warnings - +30%
- code complexity - +41%
- as a result, the velocity starts to decline over time

AI helps to write more - but not always better.


benefit the most. AI agents are accelerators, but quality still requires an engineer.

Read Here

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1949 196
What if simply changing the library would unlock all the processor cores without rewriting the code?

pandas runs joins on a single core, leaving the others idle when working with large tables.

Polars distributes join operations across all available cores and is therefore significantly faster than pandas on large data sets.

Why is Polars so fast:
• Processes rows in batches in parallel
• Uses all CPU cores
• Requires no configuration

Article - Pandas Vs Polars Vs Duckdb

Run this code

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →