TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
582
Photos
845
Videos
446
Links
675

Showing posts older than #1927 · Back to latest

Older Posts 20 shown
Post #1926 193
Deploy and run LLMs directly on your phone

Unsloth now allows you to fine-tune LLMs and deploy them 100% locally on iOS and Android devices.

The video shows how this works in practice: the author ran Qwen3 on an iPhone 17 Pro with a performance of approximately 25 tokens per second.

Guide Here

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1925 197
What if you could speed up Python by 37X with just one line of code?

Optimizing slow Python functions in large codebases is a nightmare. You can try Numba or Cython, but Numba mainly works only with numerical code and NumPy arrays.

You could go with Cython, but that requires .pyx files, type annotations, and compilation. In reality, it's hours of refactoring before you even see any improvement. 😬

Codon solves this with just one line: the codon.jit decorator compiles your Python directly into machine code.

Key advantages:
• Works with any Python code, not just NumPy
• No need for type annotations - types are inferred automatically
• Compiled functions are cached and called instantly
• No code changes required except adding the decorator

Here are real performance measurements:
• Pure Python: 0.240 s
• First Codon call: 0.324 s (one-time compilation)
• Repeated Codon calls: 0.006 s (37x speedup)


Link to GitHub, Run this code

••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1924 450
How to Build a RAG application on AWS that are already familiar to everyone.

At the core of RAG, there are always two stages: INGESTION and QUERYING.

Implementation in AWS?

1️⃣ INGESTION: transforming raw data into searchable knowledge

Documents are stored in S3

When new data appears, a Lambda function triggers

It cleans the text, splits it into chunks, and builds embeddings using Bedrock Titan Embeddings

The embeddings are stored in a vector storage, such as OpenSearch Serverless

In the end, we get a knowledge base that can be searched.

An important point: reindexing.
If a single character in a document has changed, there's no point in reprocessing the entire document anew. Smart diffs and incremental updates save both time and money.

2️⃣ QUERYING: searching and generating a response

The user asks a question in the app

The request goes through the API Gateway to Lambda

The question is turned into an embedding and matched against the vector database

The most relevant chunks are passed to an LLM from Bedrock, such as Claude

The finished response is returned to the user

This way, the LLM doesn't respond "from scratch", but relies on real data.

This is the most basic version of RAG on AWS, but the underlying pattern doesn't change when scaling up.

You can add smarter chunking, improved retrieval, caching, orchestration, eval pipelines - the architecture remains the same.


••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1923 204
Most agent frameworks still force you to hardcode workflows, decision trees, and failure scenarios.

🟢 How AWS truly changes the way developers build agent pipelines?

This works just fine until the agent encounters a situation you haven't anticipated.

AWS's Strands Agents take a different approach.

Instead of overcomplicating orchestration, they rely on the reasoning capabilities of modern models and allow the model to drive the workflow itself.

You:

- set the goal,
- provide the tools,
- and let the agent figure out how to reach that goal on its own.

From an implementation perspective, everything is as simple as possible » you only need to define three things:

1. LLM,
2. tools,
3. the task.

And that's it. The agent cycle takes over from there.

At launch, the model decides:

- whether a tool is needed,
- which one specifically,
- how to format the input data,
- and when to stop.

All the logic is » model-driven.

That is, instead of rules like "if it's math, call the calculator" or "if the API is down, retry", the model dynamically plans steps and executes them itself.

The workflow is formed during execution, based on the goal and available tools.

A concrete, applied example is shown in the video above.

This is an agent based on MCP, which can create videos in the style of 3Blue1Brown with a single simple prompt.

The setup was minimal:

- an MCP server with a tool for running Manim scripts,
- then we described the agent with these tools. Without system prompts and without workflow rules.

Next, the model itself:

- planned the steps,
- generated a Manim scene,
- and autonomously invoked the necessary tool.

Even a minimal configuration already provides a surprisingly high level of capabilities. And as you add better tools or models, the agent becomes stronger on its own - without rewriting the workflow.


Framework is completely open source and works with any local setup. 📖

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1922 195
Open models T5Gemma-2 and FunctionGemma by Google

🟢 T5Gemma-2 is impressive return of encoder-decoder architecture.

Like the first T5Gemma, which was released in the summer, T5Gemma-2 was trained on the basis of the regular Gemma, which is a decoder-only, with a long context of up to 128K tokens and multimodality.

In the summer, Google demonstrated the main idea of adaptation: we initialize the encoder-decoder with the weights of the decoder-only and continue pretraining through UL2. Now, this approach has been transferred to multimodal and long-context, plus some architectural optimizations have been added.

As a result, this time it's not just an experiment with the architecture, but a really useful model. It has the brains of Gemma, but it handles long contexts better (+cheaper) because it has an encoder. So if you have a task like summarization or working with large documents - feel free to use it.

There are variants for 270M-270M, 1B-1B, 4B-4B (estimated total parameters ~370M / ~1.7B / ~7B). Unfortunately, Google doesn't publish the instruct, but the pretraining checkpoints are available here.


🟡 FunctionGemma is a tiny tool-caller for agents, basis for an autonomous local agent.

Generator of structured function calls with only 270M parameters. It even has a different tokenizer from the regular Gemma. It can generate text, but its main role is to call the necessary tools to perform the task. In short, something like Siri specifically for performing offline tasks on the device.

Google emphasizes that the model is designed for retraining (not prompting) on specific tasks. For example, in the case from the blog post, it was retrained quite cheaply on Mobile Actions, and the accuracy increased from 58% to 85%,


••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1921 192
Avocado and Mango: two new models from Meta

There's a rumor that the company is developing and plans to release a new powerful LLM (Avocado) and an image/video model (Mango) in the near future. Alexander Wang mentioned this.

Unfortunately, according to preliminary information,

Both models will be closed-source. It seems that with the departure of LeCun, the open-source era of Meta is coming to an end.

But let's hope that at least the level of the models will be decent. After all, Zuckerberg spent 100500 billion on hunting talent for a reason.


••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1920 191
Hybrid tokenizer for diffusion on steroids.

In diffusion architectures, VAE is a thankless task which is convert pixels into latent code and back, and adding data to it doesn't help main DiT model generate better images.

🟢 What Visual Tokenizer Pre-training (VTP) Changed?

Their hypothesis is that the tokenizer should not just mechanically "zip up" pixels, but understand the semantics of the image.

To implement this, they combined 3 losses in the training of the tokenizer:

☞ Standard pixel reconstruction loss;
☞ Self-supervised learning (through Masked Image Modeling and distillation, as in DINOv2);
☞ Image-text contrastive loss (as in CLIP).

This made the latent space structurally semantic: now the vectors encode meanings, not just color spots.

Theoretical assumptions were confirmed in practice.

It turned out that the quality of generation directly depends on the "intelligence" of the tokenizer. Without changing the architecture and hyperparameters of DiT itself and without increasing the cost of its training, just by using the VTP tokenizer, it was possible to improve the FID metric by 65.8% and accelerate the convergence of the model by 3 times.

But the main discovery is that the scaling law for Stage 1 has worked.

Now, the more computational power and data are poured into the pre-training of the tokenizer, the better the final generation becomes, which was impossible to achieve with ordinary VAEs before.

3 VTP checkpoints with different numbers of parameters have been published in the open access: VTP-Large, VTP-Base, VTP-Small.


TileSet of models with GitHub • #AI #ML #Diffusion #Tokenizer #Minimax

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1919 184
Noticed something in architecture of transformers?

Almost all implementations use pre-norm variant and it consistently outperforms the original post-norm design

Why pre-norm work better than post-norm?

Normalization before sublayer, then residual
⁠☞ post-norm: output = norm(x + sublayer(x))

First Residual, Then Normalization
☞ pre-norm: output = x + sublayer(norm(x))

BUT Why does this seemingly minor change allow transformers to be trained much deeper and more stably?

It improves the flow of gradients, but I'd like a more in-depth explanation.


What's specific math involved, and
What's the key reason exactly? 🤔

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1918 189
TurboDiffusion accelerats video generation by 100X.

Want to generate a 5-second video on a large SOTA model? You Run Prompt⁠→Go For Coffee ⁠→ Come Back, process is still going on. Its Harsh reality that generation can take more than an hour.

🟢 How TurboDiffusion accelerats 100+ times?

The PROBLEM: Main problem were monstrous computational complexity of the attention mechanism in transformers, the need for hundreds of denoising steps, and the huge amount of memory for full-precision weights.

The SOLUTION: TurboDiffusion combined all effective compression and acceleration methods into a single pipeline. Idea was that sparsity and quantization are techniques that do not interfere with each other.

Architecture relies on 3 Optimization Pillars:

☞ They replaced the standard attention with a hybrid of SageAttention2++ and Sparse-Linear Attention (SLA), which turned quadratic complexity into linear. So that the model focuses only on important tokens.

☞ They distilled the sampling through rCM - instead of the standard 50-100 steps, the model comes to the result in just 3-4 steps without losing the essence of the image.

☞ They converted both the weights and activations of linear layers to INT8 using block quantization, so as not to lose accuracy.


To top it all off, they were able to combine the weights after fine-tuning under SLA and distillation of rCM into a single model, avoiding conflicts.

BENCHMARK results look like a typo, but it's not.

On the RTX 5090, the generation time for the heavy model Wan2.2-I2V 14B fell from 69 minutes to 35.4 seconds. And for the lighter Wan 2.1-1.3B - from almost 3 minutes to 1.8 seconds.

At the same time, judging by the examples, the visual quality remains practically indistinguishable from the original.


Set of Models, Technical report, GitHub • #AI #ML #I2V #T2V #TurboDiffusion

••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1917 401
AI pipeline to verify IT Forecasts from 10 years ago

Andrey Karpaty published an analysis of his project that analyzes archived Hacker News threads and, with the help of LLM, checks whether users' predictions have come true after 10 years.

🟢 HOW?
The project uses "hindsight" to compare old comments with reality, identify visionaries, and find the most high-profile mistakes.

Technically, the solution is a pipeline that collects data via the Algolia API and processes them using a structured prompt.

A test run on 930 discussions (a monthly archive of Hacker News articles) took about an hour and cost just $58.


System generates a Static Site at output, with a "Hall of Fame" of analysts and a rating of forecast accuracy.

GitHub • #AI #ML #Karpaty

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1916 193
SHARP -a photo-realistic 3D generator from a single image (by Apple)

Creates photo-realistic new perspectives of a scene with just one photo.

🟢 Why it matters?
Predicts 3D scenes in form of Gaussians in a single pass and resulting 3D scene can be:
- rendered in real time
- obtain high-quality images from close perspectives
- move camera in real metric coordinates

Features:
- metric 3D representation with absolute scale is used
- Real camera movements are supported
- the model works zero-shot, without retraining on new datasets

Model sets a new level of quality on several datasets at once:

- LPIPS metric improved by 25–34%
- DISTS metric improved by 21–43% compared to best previous models

At same time, Generation time has been reduced by thousands of times.

SHARP shows How far 3D reconstruction and view synthesis methods have come and How quickly such technologies are starting to work in real time, not just in lab.

Github, HF, Demos #Apple #LLM #AI #ML

🤖 Data Science, ML & Big Data with @DataXplore
Post #1915 383
𝗧𝗲𝗰𝗵 𝗜𝗻𝘀𝗶𝗱𝗲 AI easily interprets information in simple requests, but if input is very long and complex, model may misunderstand. To avoid this, try adding structure to prompt and make response of AI more predictable and clear. How to structure a prompt? The creators…
Is it possible to replace assessors with LLMs?

YES, but wisely.

At first look, it seems everything is simple: we write a prompt, send it to the model, and get the result. It's cheaper and faster than human annotation.

But in practice, to make the prompt work effectively, it requires many iterations, improvements, and experiments. For example, for our emotion model, we spent 3 weeks optimizing the prompt. As a result, we obtained a good correlation coefficient between the LLM and assessors — on average, 0.81.

Then we trained two classifiers on different datasets:

📌 Exclusively with LLM labels.
📌 With assessor labels after the work done to improve consistency and reduce errors.

LLM annotation — Weighted F1 0.66
Assessor annotation — Weighted F1 0.7

As a result, the quality of the LLM labels is only slightly below the reference annotation. At the same time, we saved a large amount of time and human resources.


Currently, LLMs can be considered as junior assessors. The model responds approximately like a human, and in the absence of resources, it can be used for data annotation. In some tasks, this will yield quality comparable to the reference annotation.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1914 205
NVIDIA has introduced a new open family of Nemotron 3 models

1️⃣ Nemotron 3 Nano
Universal model for reasoning and chat, focused on local deployment.

Key characteristics:
- MoE architecture: 30B parameters total, ~3.5B active
- Context up to 1 million tokens
- Hybrid architecture:
- 23 Mamba-2 + MoE layers
- 6 attention layers
- Balance between speed and quality of reasoning

Requirements:
- About 24 GB of video memory is needed for local deployment

The model is well suited for long dialogues, document analysis, and reasoning tasks

An interesting example of how MoE and Mamba are actually starting to reduce hardware requirements while maintaining context scale and quality.


2️⃣ Nemotron 3 Super & 3️⃣ Nemotron 3 Ultra
significantly surpass Nano in scale - by about 4 times and 16 times respectively. But the key point here is not just the size of the models, but how NVIDIA managed to increase power without a proportional increase in inference cost.

NVFP4 and the new Latent Mixture of Experts architecture are used for training Super and Ultra. It allows for four times more experts to be used at the same inference cost. Essentially, the model becomes "smarter" due to a more flexible selection of experts, rather than constantly activating all parameters.

Additionally, Multi-Token Prediction is used, which accelerates training and improves the quality of reasoning on long sequences. This is particularly important for agentic and multi-agent scenarios, where models work with long context and complex decision chains.


A good signal for the industry. Release, Guide, GGUF, lmstudio

#AI #LLM #NVIDIA #Nemotron3 #OpenSource #MachineLearning

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1913 430
SQLite-Vec: a tiny and portable vectorDB, built on top of SQLite.

It's very fast and lightweight, and is perfect for on-device RAG solutions.

Key features:

- matryoshka-slicing of embeddings
- reducing storage volume by 32 times through binary quantization
- support for L2, cosine, and Hamming distances

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1912 190
How to "localize" dangerous knowledge within a small part of the model, rather than spreading it across all weights?

🟢 Know more with 5W1H:

Problem:
LLMs easily absorb risky skills from dirty datasets - harmful content can slip through filters, enter training, and then be almost impossible to completely remove. Such knowledge is usually distributed throughout the entire network.

Idea of the work:
Researchers pre-allocate a tiny part of the model - a small set of neurons and attention heads - and designate it as a "risky zone". This is where the targeted dangerous information should be stored.

How it works:
- During training, risky examples update only this zone, and the gradient signals to the remaining weights are set to zero.
- Normal examples, on the contrary, are trained with the risky zone disabled.
- After training, researchers set the weights of the risky zone to zero, removing dangerous knowledge but hardly affecting the model's general capabilities.

Why it's effective:
Early labeled dangerous data "pave the way" - all further leaks of harmful knowledge from unlabeled or incorrectly labeled datasets are also directed to the same area. As a result, harmful skills do not spread throughout the entire model.

Results:
- On tasks with bilingual stories, as well as with biological and military topics from Wikipedia, this method significantly better removes targeted knowledge than simple data filtering.
- The model becomes much more resistant to adversarial fine-tuning, which usually restores prohibited skills.
- The downside is that it requires more computational resources.


These are the first steps towards practical and manageable "ability removal" from LLMs through knowledge localization, rather than attempts to clean up datasets or post-training.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1911 382
Wait a second, does Mistral 3 Large use the DeepSeek V3 architecture, including MLA?

I just went through the configurations: the only difference I saw is that there are 2 times fewer experts in Mistral 3 Large, but each expert is 2 times larger.

Maybe it's easier to notice this when comparing the architectures side by side

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1910 185
Google is gradually moving towards training robots using AI models of the world

The researchers took the basic Veo-2 and fine-tuned it based on the first frame + the robot's actions to generate future consistent frames from its 4 cameras. This is called action-conditioned rollout and, in essence, allows for inexpensive and safe evaluation of the robot's policy using just the world model.

🟢 Why is this cooler than regular simulation?

Strict physical simulators work well if the situation is simple and predictable. AI simulation can be scaled to non-trivial worlds. Moreover, every object in a physical simulator requires clear assets and manual tuning + heavy calculations. You can't go far with that. Here, on the other hand, you can add new objects and cases as much as you want - just write a prompt or edit the initial frame with Nano Banana.

Of course, there are also downsides. They mainly concern the quality of strict modeling, especially of fine physics. But there's still a lot to come.

Google has learned to fairly decently evaluate the robot's policy using Veo (see the last graph). Add a policy update, and you'll already get reinforcement learning. So far, they're not doing this consciously, again due to the lack of accuracy in the World model.


The resulting fine-tuning was cleverly named Veo (Robotics). It's A Big Step

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1909 199
The Voronoi Diagram (black dots) is calculated as the vertical projection of the lower envelope of n three-dimensional function graphs {(x, yᵢ(x))} with yᵢ(x) = D(xᵢ, x) (pink).

When the distance D(x, x′) = ‖x - x′‖², the yᵢ graphs represent paraboloids, and the boundaries of the Voronoi cells become linear.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1908 410
There's an analogue of GPT with a contextual window of just 10 words operating inside our brain.

Imagine a biological neural network whose physical size, if all its tissues were put together, wouldn't exceed the size of a regular strawberry.

🟢 How our brain processes speech?
By MIT Neuroscientist Ev Fedorenko

Her conclusions will sound very familiar to engineers and data scientists: there's a system operating inside the human head that behaves suspiciously similar to modern large language models. It's a kind of "mindless" language processor that maps words and meanings, but itself is completely incapable of thinking.

This ASSERTION IS BASED on a significant BODY OF DATA.

Fedorenko's lab conducted fMRI scans of 1,400 people to build a detailed probabilistic map of brain activity.

The architecture of this "language network" turned out to be surprisingly stable and reproducible: in most adults, it's localized in 3 specific areas of the left frontal lobe and along a long stretch of the middle temporal gyrus.

Fedorenko calls this structure a functional block, comparable to an organ like the digestive system or the face recognition zone.

The most interesting part begins when you look at its functionality. Fedorenko describes this network as a parser or a set of pointers. Its task is purely utilitarian - to act as an interface between input signals (sound, text, gestures) and abstract representations of meaning stored in completely different parts of the brain.

The language network itself has no episodic memory, social intelligence, or the ability to reason. The entire process of reflection takes place outside its boundaries.

This explains the phenomenon of aphasia: when this "interface" is damaged, a person retains complex cognitive thinking but is locked inside themselves, having lost access to the vocabulary and grammatical rules.

The similarity to LLMs becomes even more obvious when you look at the system's limitations.

Research shows that the human language network has an extremely narrow contextual window: it can effectively process chunks of up to 8-10 words in length.

In essence, this is a rather superficial system. It reacts to Noam Chomsky's grammatically correct nonsense "Colorless green ideas sleep furiously" just as actively as it does to meaningful sentences. What matters to it is the structure and statistical probability of word pairings, not the truth or deep meaning of the statement.

This makes it akin to early language models: the network simply learned the rules by which words are assembled into chains.

Fedorenko's data force us to reconsider classic anatomical concepts, as many textbooks still refer to outdated concepts.

For example, Broca's area, which for decades was considered the center of speech, turned out to be an area of motor planning. It merely prepares the mouth muscles for articulation and is activated even when uttering complete nonsense, acting as a controlled region for receiving commands.


The real language network of the brain is a separate, specialized computing cluster that, like ChatGPT, brilliantly mimics the coherence of speech, even if there's no real thought behind it.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Post #1907 230
RLAX: Large-Scale, Distributed Reinforcement Learning for

Apple published & then quickly removed ⚠️ brought you a leaked stuff 😎

RLAX js a scalable reinforcement learning framework for LLMs on TPUs.

🚫 What's inside RLAX?
- Parameter server architecture
- The central trainer updates weights
- Huge inference fleets pull up weights and generate rollouts
- Optimized for preemption and massive parallelism
- Special techniques for data curation and alignment

The results are impressive:
- +12.8% to pass@8 on QwQ-32B
- In just 12 hours and 48 minutes
- 1024 TPU v5p were used

Why this is important?
- Apple is clearly experimenting with RL on a very large scale
- The TPU-oriented architecture indicates a focus on efficiency, not just the model
- The gain is achieved not through "model magic", but through engineering the training system
- This is another signal that RL for LLMs is entering the phase of industrial pipelines


What you think, Why Apple removed?

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →