TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
582
Photos
845
Videos
446
Links
675

Showing posts older than #2020 · Back to latest

Older Posts 20 shown
Post #2019 143
Tencent is making a strong entry into the context learning field.

Open-source benchmark CL-bench has been released - and this isn't just another dataset, but an attempt to shift the focus of the entire industry.

🟢 What they done?

Tencent HY, in collaboration with Fudan University, have published a new work:
“CL-bench: A Benchmark for Context Learning” - a systematic benchmark for evaluating whether *models are actually able to think in context*, rather than just recalling what they've learned.

This is the first research release from Vinces Yao's team since his move to Tencent - and it's clear from their ambitions that they're aiming for fundamental changes.

Today, most LLMs operate according to the following scheme:
huge weights + memorized patterns = answers

But the real world isn't a memory test. It's about:

- long, complex contexts
- conflicting information
- the need to change strategies on the fly
- drawing conclusions based on what's just appeared

Models need to move from static memorization to dynamic reasoning within context.

CL-bench precisely tests this breaking point:

- how the model uses context, not just weights
- whether it can update its understanding
- whether it's capable of reasoning in complex scenarios, not just on pure QA tasks

In essence, this is a step towards models that are closer to agents than to "smart autocomplete".

Plus a strategic signal

At the same time, Tencent is launching Tencent HY Research - a blog where frontier research will be published.

This looks like a declaration:
"We're not just training large models. We want to influence how they're evaluated at all."

And this is already a level of influence on the direction of the entire field.
CL-bench isn't about +0.5% on the leaderboard.
It's about a paradigm shift:

The LLMs of the future = less rote learning, more thinking in real-world contexts.

And if this line succeeds, it's precisely such benchmarks that will determine who has truly created a "smart" model, and who has just inflated the parameters.


Project | Blog

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2018 170
RAG vs CAG, a clear explanation in image.

Merging RAG and CAG. How can an AI engineer use this?

Let's break down how this looks and what additional considerations need to be taken into account.

➡️ Example steps for a CAG + RAG architecture:

👉 DATA PREPARATION

1️⃣ For CAG, we only use sources that change infrequently. In addition to "infrequent changes," it's important to understand which of these sources are most frequently needed for relevant queries. Only after this, we "warm up" the selected data in advance in the KV cache model and cache it in memory. This is done once, and the remaining steps can be repeated many times without recalculating the initial cache.

2️⃣ For RAG, if necessary, we pre-calculate and save vector embeddings in a compatible database so that we can later search for them in step 4. Sometimes, simpler types of data are sufficient for RAG, in which case a regular database would work.

👉 QUERY PATH

Now we can use the prepared data.

3️⃣ We assemble the prompt: the user's query + a system prompt with instructions on how the model should use the cached context and external (retrieved) context.

4️⃣ We build an embedding of the user's query for semantic search through the vector DB and query the context store to retrieve relevant data. If semantics aren't needed, we can go to other sources, such as a real-time database or the web.

5️⃣ We enrich the final prompt with the external context retrieved in step 4.

6️⃣ We return the final response to the user.

👉 A FEW IMPORTANT POINTS

☞ The context window is not infinite. Even if the model has a huge context, the "needle in a haystack" problem still exists. Use context sparingly and cache only what is really needed.

☞ For some cases, certain datasets are super valuable to constantly feed them to the model through the cache. For example, an assistant who is obliged to always comply with a long set of internal rules scattered across several documents.

☞ Although open-source CAG has gained popularity relatively recently, in practice, this has long been possible through prompt caching in the OpenAI and Anthropic APIs. It's easy to quickly put together a prototype there.

☞ Always separate hot and cold sources. We only put cold (infrequently changing) data in the cache, otherwise the data will become outdated, and the application will start living "out of reality."

⚠️ RISKS & LIMITATIONS

☞ Be very careful with what you cache, because this data becomes available to all users' queries.

☞ It's difficult to ensure RBAC for the cache if you don't have a separate model with its own cache for each role.


Have you already tried this combo approach?

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2012 175
Open-source extension for LLM serving engines:

Like a caching layer for large-scale production inference of LLMs.

LMCache implements smart KV cache management by reusing key-value states of already encountered text between GPUs, CPUs, and local disks.

🟢 What it can do?

It can reuse any repetitive text fragments, not just prefixes.

This results in:

⁠☞ 4–10x cost reduction in RAG for models owned by the user
⁠☞ Lower Time-To-First-Token (TTFT)
⁠☞ Higher throughput under load
⁠☞ More efficient work with long-context scenarios

An illustrative example of application: NVIDIA integrated LMCache into its inference project Dynamo.
LMCache allows Dynamo to offload the KV cache to external storage layers and effectively reuse it between requests. This reduces the cost of prefill and frees up GPU memory for active computations.

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2011 164
PersonaPlex is a smart real-time Model for Voice-Controlled and Role-Based Dialogues

Enables two-way voice communication with character control via text prompts and audio.

Generates natural, low-latency interactions, trained on synthetic and real-world dialogues.

What are the Key Features?
- Support for different voices for natural communication.
- Training on synthetic and real-world data.
- Ability to control the character via text prompts.
- Low latency in interactions.


GitHub

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2010 162
Many teams are trying to apply DevOps practices to LLM applications.

But DevOps, MLOps and LLMOps solve fundamentally different problems.

🟢 Break it down, DevOps vs MLOps vs LLMOps:

⁠➜ DevOps is focused on software.

You write code, test it, and deploy it. The feedback loop is simple: does the code work or not?

The main artifact is code. Testing is deterministic. The tooling is mature after 15+ years of development.

⁠➜ MLOps is focused on (model + data).

Here you have data drift, model degradation, and constant retraining.

The code might be perfect, but the model quality degrades over time because the world changes.

An anti-fraud model might work great at launch, but start failing after a few weeks because the fraudsters have adapted.

The main artifact expands to code + data + models. All three need to be versioned. That's why MLflow, DVC, and feature stores have become essential tools.

⁠➜ LLMOps is focused on foundation models.

Usually, you don't train models from scratch. Instead, you choose a base model and optimize it in three parallel directions:

* Prompt Engineering
* Context Tuning / RAG
* Fine-tuning

Unlike DevOps and MLOps, these directions run in parallel, not sequentially.

But the Biggest difference of LLMOps is monitoring: it's completely different.

In MLOps, you track data drift, model degradation, and accuracy metrics.

In LLMOps, you track:

⁠☞ Hallucination detection
⁠☞ Bias and toxicity
⁠☞ Token consumption and cost
⁠☞ Human feedback loops

Because the output of LLMs is non-deterministic. You can't just check if it "answered correctly". You need to ensure the answer is safe, grounded, and doesn't burn the budget.

63% of production AI systems catch dangerous hallucinations in the first 90 days.

⁠➜ cost model also flips

In MLOps, the main cost is training (GPU hours during development).

In LLMOps, the main cost is inference (each request consumes tokens).

That's why efficiency of prompts, caching, and routing between models are so important in LLMOps.

The evaluation loop in LLMOps feeds back into all three optimization directions at once. A failed eval might mean you need better prompts, richer context, OR fine-tuning.

That is, it's no longer a linear pipeline.

And another thing: Versioning prompts and RAG pipelines in LLMOps is now first-class, just like versioning data has become mandatory in MLOps.

And the ops layer you choose should match the system you're building.


Why this matters?
88% of ML initiatives struggled to reach production if trying to launch them through traditional DevOps approaches.

And LLMs add challenges that MLOps wasn't even designed for in the first place.

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2009 166
Marching Squares is a classic algorithm for constructing contour lines (isolinues) from a 2D scalar field.

Used for visualizations such as topographic maps.

Each grid cell is mapped to a simple polygon depending on which of its angles are above or below a specified threshold.

It has a 3D counterpart, Marching Cubes, which does the same thing, but for 3D.

If you like such visualizations, you might be interested in a recent article about the ML concept of Rectified Flows.

There are many explanatory interactive visualizations there.

Code for all of this can also be viewed on GitHub.

•••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2008 182
HOW YOLO BECAME A STANDARD IN CV?

Launching a series of posts about the evolution of one of the most popular architectures in computer vision.

We'll break down:

Before 2015, the task of detection was solved by searching for the most likely regions. There were two-stage approaches, such as Faster R-CNN.

🟢 How did the YOLO architecture evolve from v1 to v3?

📝 First, they searched for candidate regions, and then used a refine process to refine the classes and coordinates.

PROBLEM: The process was very slow. Imagine the task of tracking a tennis ball on the court during a match. Old networks would have taken 5 minutes to process a video, even on a good GPU. Players would have had to stand and wait for the VAR system.

A real-time approach was needed, where speed was more important than perfect results. Thus, YOLO was born.

➡️ YOLO v1: A model that looks at the entire scene (2015)
The idea was to turn detection from a region-searching task into a regression problem. Combine all stages into a single network that directly "spits out" coordinates.

How it was implemented technically?

📝 They made an architecture similar to GoogLeNet. Two fully connected and 24 convolutional layers. Although it was large, it detected bounding boxes and immediately determined the coordinates.

📝 All images were divided into a 7x7 grid. Each cell predicted 2 bounding boxes and 20 classes. The input was a 448x448 image, which was further divided into 64x64.

PROBLEM: YOLO v1 couldn't handle other resolutions. To work with detection on large images, they resized them to 448x448 or cut them into patches. Due to the extra operations, the main advantage over Faster R-CNN — speed — was lost.

➡️ YOLO v2 / YOLO9000: Scale and anchors (2016–2017)

To level the complex LOSS, multi-scale was added to the new version: YOLO9000 simultaneously detects more than 9,000 classes without full annotation — hence the name.

What new features were added?

📝 Anchor Boxes: Instead of directly predicting coordinates, they switched to predicting shifts relative to the X and Y axes for candidates. This maximized object capture.

📝 Skip Connections: They introduced pass-through layers and added batch normalization, which solved the problem of gradient fading.

PROBLEM: The accuracy of detections became heavily dependent on anchor boxes. The anchors were manually selected, and if they were poorly chosen for the dataset, the model's metrics suffered.

➡️ YOLO v3: Victory over other models (2018)

Thanks to the update, YOLO v3 became a foundation in ML. It surpassed Faster R-CNN in popularity and became a favorite of many developers.

What was added new?

📝 Multiscale detection. It removed noise when detecting small objects and stopped ignoring them.

📝 The "third eye". The network immediately outputted three candidates at different resolutions — large, smaller, and the smallest.

PROBLEM: The version became slower. Due to the complexity of the architecture, v3 became heavier than its predecessors. The anchors were still manually selected, which also slowed down the detection process.


Model continued to evolve, but no longer in the hands of its original author, Joseph Redmon: he left ML and handed over project to a large company.

In next post, we'll break down why YOLO v4 is called the "engineer's constitution" and YOLO v5 is a "ugly duckling"?.

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2006 188
To train an ML model, you need to be proficient in algorithms, write code, endlessly tune hyperparameters which is a high entry barrier for most people.

An open-source project Plexe significantly lowers this threshold: You describe the task in plain language and it automatically assembles machine learning for it.

🟢 How it works?

☞ Explain in a human-friendly way what you want to predict, what the input data is and what the output should be.

⁠☞ Next the system, through a combination of several agents, goes through the entire pipeline: data analysis, solution plan, code generation, tests, and quality assessment.

⁠☞ Supports various LLM providers: OpenAI, Anthropic, Ollama, and others. Plus, it can automatically derive the data structure or even generate a synthetic dataset.

⁠☞ There's also distributed training on Ray inside: you can run multiple model variants in parallel and significantly speed up the process.


GitHub

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2005 166
🐋 DeepSeek-OCR 2 is a new generation of OCR with SOTA quality

A 3B model for advanced understanding of images, documents and OCR, which reaches the SOTA level.

🟢 Why this matters?

The key novelty is DeepEncoder V2.

Unlike classic vision LLMs, which "read" the image as a grid (left-to-right, top-to-bottom), DeepEncoder V2 works closer to how a human reads:

- First, a global understanding of the image is formed
- Then, the model determines the logical order of reading - what is important first, what next

What this brings in practice

📄 Works better with complex document layouts
📊 Correctly reads tables
🧾 Links signatures and values
📰 Understands columns and structured text
🔀 More reliably processes a mixture of text and visual structure

In terms of quality

- Outperforms Gemini 3 Pro on a number of benchmarks
- Gives >4% improvement compared to the previous version of DeepSeek-OCR

And this is with a model size of just 3B parameters.

Can be launched and fine-tuned


Now, DeepSeek-OCR 2 can be conveniently launched and fine-tuned via Unsloth according to the ready-made guide.

Guide, Model, Github, Paper | #DeepSeek #ocr #opensource

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2004 165
Boosted Up performance of AI agent by 184% using a completely open-source technique.

Now, you can automatically find best prompts for any agentic workflow you're putting together, means manual prompt engineering isn't needed at all.

🟢 What is the simple idea?

1️⃣ Take a starting prompt and an eval dataset
2️⃣ Then, an optimizer iteratively improves the prompt
3️⃣ In the end, you get an optimal prompt automatically

And all of this in just a few lines of code.

➡️ Why Opik specifically?

Opik is a 100% open-source platform for evaluating LLMs.

It helps optimize LLM systems so that they work better, faster, and cheaper: from RAG chatbots to code assistants. Opik includes tracing, evaluations, and dashboards.


Best Part: Everything can be run completely locally, because you can use any local LLMs as optimizers and evaluators.

GitHub repository

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2002 152
A new class of risks for open-source models, which few people think about.

A work on so-called elicitation attack: took open-source model, retrain it on seemingly harmless data on chemical synthesis, which were generated by frontier models.

And suddenly, this open-source model starts to perform significantly better on tasks related to chemical weapons.

🟢 What the paper shows?
The most unpleasant thing here is not "how to make the model respond to prohibited questions". But the fact that the model can be dangerous, even if it itself does not output anything harmful. Because its harmless answers can become training data that unlock dangerous capabilities in another model.

What the authors showed:
☞ The attack works on different open-source models and on different types of "weapon" tasks
☞ Retraining on data from frontier models gives a greater boost than training on chemistry textbooks or on data,
☞ Generated by the same open-source model
☞ Sufficiently "peaceful" topics: cheese making, fermentation, candle chemistry, etc.
☞ In one experiment, "harmless chemistry" gave about 2/3 of the effect on the growth of "weapon" competence compared to training on data about chemical weapons
☞ The stronger the frontier model, the stronger the subsequent uplift of the open-source model (and the higher the risk)

CONCLUSION is simple and rather harsh: focusing only on "refusal training" in frontier models doesn't solve problem. The danger can leak through normal, seemingly everyday answers, which someone then uses as a dataset.


Read Here
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2001 154
Came across an interesting project: tiny-infini-gram.

Training-free language model (without training) which, according to him, generates Shakespeare 250 times faster than nanoGPT.

What's inside, in theory?

This is the "unbounded n-gram" approach: you can use very large n without running into exponential memory like with classic n-grams.

Instead of a huge n-gram table, a suffix array is used: it simulates n-gram lookup of any size with logarithmic access time.

Previously, such things were hardly used for generation, because sampling broke down: infinite perplexity and frequent verbatim copying.

The author claims that he solved this with a new method called Selective Back-off Interpolation Sampling: it mixes probability distributions from several levels of n-grams to maintain a balance between quality and novelty.

If you like non-standard LM approaches and "fast, simple, without training" - it's worth checking out the analysis and code.


Link to detailed write-up

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #2000 158
How to do personalized search using only Postgres

There are two phases:

DATA PREPARATION PHASE:

1️⃣ Generate movie embeddings for vector search and create a BM25 index for full-text search.

2️⃣ Generate user preference embeddings based on what the user has watched and what they have liked or disliked before.

SEARCH PHASE:

1️⃣ Retrieve the top-100 movies ranked by BM25.

2️⃣ Normalize BM25 scores to a range of 0–1.

3️⃣ Perform personalized search: compare movie embeddings with the user's preference embedding.

4️⃣ Combine signals: 50% for text match (relevance) and 50% for user match (personalization).
Guide Get Here

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #1999 183
Vector Index vs Vector Database, in simple terms

Developers often use these terms as synonyms but if you don't understand the difference, it can lead to problems in production.

🟢 How to think about it?
A vector index is essentially a search algorithm.

You feed it with vectors, it organizes them into a structure that allows you to quickly search for similar ones (e.g., HNSW), and finds nearest neighbors. FAISS is also an example of this.

But the nuance is that's all.

It's not responsible for storage, it can't properly filter by metadata, and it doesn't scale on its own. It just does the search.

A vector DB is a wrapper around the index, plus everything else that's actually needed in production.

It includes distributed storage, persistence, filtering by metadata, concurrent access, and more. A good open-source example: Milvus.

From this, it's clear: one is a component, the other is a system.

🟡 Why this matters?
Once, a company with autopilots was building a search for road videos, and the scale was huge.

Each trip generated frames, each frame turned into an embedding.

Engineers needed to ask something like "night crosswalks with pedestrians" based on months of data.

At first, FAISS looked perfect: fast, lightweight, easy to set up.

But as the data grew, the embeddings of each day became a separate index file.

After a couple of months, they had hundreds of thousands of scattered files.

Searching for several days meant digging through tons of files at once.

Complex queries required custom DBs, query planners, and filtering built around FAISS.

In the end, they had billions of vectors and no clear path forward.

This is where vector DBs come in. The company migrated to Milvus, and the difference became obvious:

↳ one query: similarity + filters by metadata
↳ data in collections and partitions, not scattered across files
↳ tens of billions of vectors, a year in production, no major incidents
↳ 30% reduction in infrastructure costs
↳ 10x scaling headroom

And this isn't a unique story.

Most teams struggle when they start with a lightweight index and then suddenly need filters, reliable storage, and real scale.

Vector DBs exist precisely for this moment.

Milvus stands out for its ability to handle scale and different types of data well.

You can store billions of vectors, scale horizontally, and create specialized indexes, for example for geodata with its optimized index, not "one common for everything".


Completely open-sourced on GitHub (41k+ stars), can be self-hosted or used in their cloud.

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Post #1997 157
Powerful open-source release (voice design + voice cloning)

Qwen has officially released Qwen3-TTS and fully opened up the entire line of models - Base / CustomVoice / VoiceDesign.

🟢 What's inside?
- 5 models (0.6B and 1.8B classes)
- Free-form Voice Design - generating/editing a voice based on a description
- Voice Cloning - cloning a voice
- 10 languages
- 12Hz tokenizer - strong audio compression without significant quality loss
- full support for fine-tuning
- claim SOTA quality on a number of metrics

Previously, the best generators were in closed APIs, but now a full-fledged open-source TTS stack is emerging, where you can:
- train for a domain,
- create custom voices,
- and not depend on a provider.


GitHub | HuggingFace | Demo | Blog | Paper • #AI #TTS #Qwen #OpenSource #SpeechAI

•••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1996 163
Researchers have developed a new approach to RAG that:

* does not require a vector DB
* does not create embeddings
* does not chunk documents
* does not perform similarity search

And it demonstrated 98.7% accuracy on a financial benchmark (SOTA).

🟢 What are the key problem of classical RAG that this approach solves?:

Traditional RAG chunks documents, converts them into vectors, and retrieves fragments based on semantic similarity.

But similarity ≠ relevance.

When you ask: "What were the debt trends in 2023?", vector search will return pieces that are semantically similar to the query.

But the real answer might be hidden somewhere in the Appendix, referenced by a link on another page, in a section that does not overlap semantically with your question at all.

Classical RAG most likely just won't find it.

PageIndex addresses this.

Instead of chunking and embeddings, PageIndex builds a hierarchical tree of the document structure, essentially a smart "table of contents."

Then the model reasons through this tree.

That is, it does not ask: "Which text is most similar to my query?"

It asks: "Based on the document structure, where would a human expert look for the answer?"

This is a fundamentally different approach, which has:

* no arbitrary chunking that breaks context
* no need to maintain and operate a vector DB
* retrieval is traceable: you can see why a particular section was chosen
* it can properly follow internal document links ("see Table 5.3"), just like a person does

But the deeper issue is this.

Vector search treats each query as independent.

Documents have structure and logic: sections refer to each other, context accumulates across pages.

PageIndex respects this structure instead of flattening everything into embeddings.

Importantly: this approach is not always sensible, because classical vector search is still fast, simple, and works great in many cases.

But for professional documents requiring domain expertise and multi-step reasoning, the tree-based, reasoning-first approach truly shines.

For example, PageIndex showed 98.7% accuracy on FinanceBench and significantly outperformed traditional vector-based RAG systems in analyzing complex financial documents.


Everything is fully open source, you can check the implementation on GitHub and try it yourself.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1995 165
🚀 Chroma 1.0 is out - a fully open speech-to-speech model with voice cloning

The FlashLabs team has released Chroma 1.0 - the first open-source model that can translate dialogue "voice → voice" in real time, with voice cloning.

The main thing:
this is not "recognition + text + dubbing".
This is an end-to-end system where the conversation takes place directly by voice.

What is promised in terms of characteristics:
- ⚡️ <150 ms end-to-end delay (almost like a live call)
- 🧬 high-quality voice cloning over several seconds of audio
- 📈 voice similarity SIM = 0.817 (practically identical)
- 🧠 reasoning on just 4B parameters
- 🔓 fully open weights + code

And a nice bonus: the model has already been optimized for SGLang (LMSYS) to work faster and cheaper in inference.


If this is really the case, then Chroma could become a real open-source alternative to closed voice systems.

Paper Model Code

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1994 157
You can fine-tune 100+ open-source models without writing any code at all.

LLaMA-Factory provides a unified interface for training LLMs and VLMs. It supports LLaMA, Mistral, Qwen, DeepSeek, Gemma, Phi, Yi, and over 90 other models.

It's completely open source.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML with @DataXplore
Post #1993 168
If you strengthen internal dialogue markers in LLMs (like "Oh" or "Wait"), accuracy of responses can increase by 2X on complex tasks

Google published a very interesting semi-philosophical study about What rizoning actually is, write that RL, in fact, teaches models to think not longer, but more collectively.

🟢 What Google observed?
You've surely noticed that when a model thinks, it often simulates a dialogue between different internal voices. It asks itself questions, may criticize or highlight something. And Google writes that the phenomenon of rizoning is contained in this structure of internal dialogue.

The most interesting thing is HOW THEY PROVE IT?

- The authors take a sparse autoencoder (what it is and why it's needed we wrote here) and find a neural feature that is responsible for surprise/awareness/change of viewpoint. This feature is activated at the beginning of sentences in dialogue contexts, and in practice, it's simply responsible for the use of things like "Oh!", "Wait a minute", "Oh, so...".

- Then this feature is specially strengthened during generation and the metrics are observed (model - DeepSeek-R1-Llama-8B).

- RESULT: on complex combinatorial arithmetic tasks, where the original model gives 27.1% accuracy, the model with strengthened dialogue marker already gives 54.8%, and with suppression of this marker - 23.8%.

The statistical significance has been checked: the authors specifically compared the strengthening of this feature with the strengthening of other features, and the effect is obvious. Plus, in parallel with the strengthening of this marker in the model, the ability for cognitive strategic thinking also increases.


In short, LLMs are still only 0.01% understood. We need to somehow try to write in the prompt Use more "ah", "oh", "exactly" and "oh yeah", and observe result. Paper

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1992 180
Document Index for Vectorless RAG Based on Reasoning

PageIndex is an open-source RAG framework that eliminates vector databases and chunking from the pipeline when searching for documents.

🟢 How it works?

Most RAG systems rely on semantic similarity: they cut a document into pieces, build embeddings, and then retrieve fragments that "resemble the query".

But similarity does not equal relevance.

In professional documents such as financial reports, legal documents, and technical manuals, multi-step parsing and domain-specific logic are often needed. Vector-based search easily gets stuck when almost every section uses the same terminology.

PageIndex does it differently.

It builds a hierarchical tree from the document, similar to a table of contents, but tailored for LLMs. Then, it uses reasoning-based tree search to "navigate" the structure in the same way a human expert would.

The two-step process:

1. Generate a tree-based index of the document structure
2. Retrieve the needed information through reasoning-based tree search

The LLM can "think" about the document structure. Instead of matching embeddings, it reasons like: "Trends in debt are usually in the financial summary or Appendix G, let's look there."

Key features:

• No vector database or embedding pipeline
• No artificial chunking that breaks context at boundaries
• Traceable retrieval with precise references down to the page level
• Navigation based on reasoning, mirroring human document analysis

PageIndex is used in Mafin 2.5 and claims 98.7% accuracy on FinanceBench for financial document analysis.


And yes, it's completely open source.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →