TGViewer
Channel Public Channel
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

Data eXplore : Data Science, ML, Big Data, LLMs and AI Security

@dataxplore

Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Subscribers
586
Photos
845
Videos
446
Links
675

Showing posts older than #1696 · Back to latest

Older Posts 20 shown
Post #1695 144
Stanford has launched a free Deep Learning course taught by Coursera founder Andrew Ng

The program covers everything: from basic neural network principles to LLM, RL, agents, RAG, and multimodal models

The first lecture is here.

Materials and schedule are here

🤖 Data Science, ML & Big Data with @DataXplore
Post #1694 111
✔️ GenAI directly on device :

Chrome, Chromebook Plus, Pixel Watch with LiteRT-LM - a framework for running LLMs directly on the device (offline), with minimal latency and no API calls.

🟢 Key points:
- LiteRT-LM is already used inside Gemini Nano / Gemma on Chrome, Chromebook Plus, and Pixel Watch.

- Runs on the device: no delays from remote servers

- Open C++ interface (preview) for integration into custom solutions.

- If you are developing applications, this is a useful.

- Provides access to Local GenAI

- Architecture: Engine + Session
  • Engine holds the base model, resources - shared across all functions
  • Session - context for individual tasks, with cloning, copy-on-write, and lightweight switching capabilities

- Supports hardware acceleration (CPU / GPU / NPU) and cross-platform compatibility (Android, Linux, macOS, Windows, etc.)

- For Pixel Watch, a minimal “pipeline” is used - only necessary components - to fit memory and binary size constraints


🟠 Open-sourced for running GenAI on devices:
- LiteRT - a fast “engine” that runs individual AI models on the device.

- LiteRT-LM - a C++ interface for working with LLMs. It combines several tools: prompt caching, context storage, session cloning, etc.

- LLM Inference API - ready-made interfaces for developers (Kotlin, Swift, JS). They work on top of LiteRT-LM to easily embed GenAI into applications.


#AI #Google #LiteRT #LiteRTLM #GenAI #EdgeAI #OnDeviceAI #LLM

🤖 Data Science, ML & Big Data with @DataXplore
Post #1693 121
IBM Granite 4.0 is now available in Unsloth

The model in GGUF format with a hybrid architecture (Hybrid Mamba) — a combination of dense layers and MoE for acceleration and memory reduction.

🟢 Key facts:
- Available sizes: Micro (3B), Tiny (7B/1B active), Small (32B/9B active).
- Context up to 128K tokens.
- Training in Unsloth is up to 2× faster and requires 50% less VRAM.
- Support for Ollama, llama.cpp, and Docker for easy deployment.


💡Useful for: chatbots, edge deployments, long documents, customization via fine-tuning.

HF

🤖 Data Science, ML & Big Data with @DataXplore
Post #1692 113
Can Arattai take on WhatsApp and Telegram?

According me “Telegram” is leading in every feature race.

You can make people download any app by Advertisement but people won't use it unless they found it ease of use and familiar.

Also you cannot blame users for not adopting your product.

You have to understand the golden rule: Users/Consumers are always right.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1691 199
X has open-sourced the "For You" algorithm

How the recommendation feed works in 7 steps:

1️⃣ Raw data (input):
- social graph (who follows whom),
- engagement (likes, retweets, replies, bookmarks),
- user data (clicks, profile, behavior).

2️⃣ Feature Engineering:
- GraphJet — real-time tweet graph
- SimClusters — community grouping ("AI Twitter", "NBA Twitter")
- TwHIN — user↔tweet connection map
- RealGraph — strength of connections
- TweepCred — trust scoring
- Trust & Safety signals

3️⃣ Candidate Sourcing (Home Mixer):
Different mixers (CR Mixer, UTEG, FRS) pull tweets from various pools → more diversity.

4️⃣ Heavy Ranker (ML model):
Neural network predicts what you will like: likes, retweets, replies, reading time.

5️⃣ Filters and heuristics:
- social proof
- author diversity
- spam/NSFW/mute blocks
- content balance
- protection against "filter bubble"

6️⃣ Mix:
Promoted tweets + "who to follow" recommendations → into the feed.

7️⃣ What this means for you:
- choose a niche
- write valuable posts
- respond meaningfully in your topic
→ grow your audience and find people/ideas for business.

GitHub

#Twitter #ForYou #AI #RecommenderSystems

🤖 Data Science, ML & Big Data with @DataXplore
Post #1690 195
Richard Sutton: “LLMs are still not the Bitter Lesson.” 🧠

The father of Reinforcement Learning argues that:

- True AI breakthroughs come from scaling computation + self-learning.

- But LLMs still rely only on human-created data → which is biased and finite.

- Real intelligence requires interaction with the world, like humans & animals.

His statement sparked huge debate: are LLMs a dead-end branch or just an early step?

YT Podcast

#AI #LLM #ReinforcementLearning #BitterLesson

🤖 Data Science, ML & Big Data with @DataXplore
Post #1689 96
🔦 Image generation with light, not on a GPU

UCLA introduced an optical generative model, uses light and lenses instead of computational units. It means images are created not on chips, but through physics.

🟢 How it works (with experiment)?
1. A lightweight digital encoder turns random noise into a phase pattern.
2. This pattern is loaded onto an optical light modulator.
3. Light passes through a diffractive decoder and the image is formed directly on the sensor.

Real experiments: using visible light and an SLM, they demonstrated generation results:
- Created digits, faces, butterflies, and even paintings in the style of Van Gogh.
- Quality comparable to modern diffusion models.
- There are two versions: instantaneous (one pass) and iterative (several steps, like diffusion).


🟠 Why this approach is interesting?
- The approach requires no computational load.
- Super-fast generation: the physics of light performs what a GPU does with billions of operations.
- It opens the way to energy-efficient AI for edge devices: AR/VR, mobile cameras, compact sensors.


🔴 Limitations:
- Difficult to align optical systems.
- Limitations on phase mask accuracy.
- Dependence on equipment quality (noise, bit depth).


But even with these challenges, this is the first step toward a new class of AI where computation is replaced by pure optics.

#AI #OpticalComputing #Photonics #GenAI

🤖 Data Science, ML & Big Data with @DataXplore
Post #1688 101
We'll start Friday with memes
Post #1687 198
This article by Sebastian Raschka walks step-by-step through implementing self-attention from scratch, then expands the discussion to multi-head and cross-attention, with clear explanations and PyTorch code examples.

A must-read if you want to deeply understand transformers.
Read here

🤖 Data Science, ML & Big Data with @DataXplore
Post #1686 87
A unified sandbox environment for developing AI agents combines a browser, terminal, file system, VS Code and Jupyter into a single Docker container ready-to-use.

🟢 Features:
All components work with a shared file system: a file downloaded in the browser is immediately available in the terminal or code.

The container also comes pre-installed with several MCP servers, allowing the AI Agent to directly call various capabilities without additional complex environment setup.

There is support for the Chrome DevTools Protocol for programmatic browser control, as well as built-in port forwarding and service monitoring for convenient previewing and debugging of web applications.


SDKs are provided for Python, TypeScript, and Golang, with one-click launch via Docker.

GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Post #1685 161
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Some #Zoho solutions that can help you every day and provide you an alternative to all Microsoft as follows. Their company has introduced the ZOHO suite, which offers a complete alternative from WhatsApp to MS Office. 🤖 Data Science, ML & Big Data with @DataXplore
#Zoho

Zoho Directory - Identity and access management

Ulaa Browser - Secure browser

Zoho Vault - Store and auto-fill passwords

Zoho One Auth - Multi-factor authenticator

🤖 Data Science, ML & Big Data with @DataXplore
Post #1684 82
A new method for training neural networks

The startup Thinking Machines has already released study contains a lot of mind-boggling mathematics.

🟢 Explan in simpler terms
When we train neural networks, one of the main problems is controlling the scale of tensors (weights, activations, gradients). If something becomes too large or too small, numerical issues arise: various gradient explosions, vanishing gradients, etc.

Usually, this is fixed at a high level using techniques like gradient clipping, weight decay, or layer norm. But here a more strict and fundamental approach is proposed: not just scaling the weights, but restricting the very structure of tensors, forcing them to live not in an arbitrary space, but on a certain manifold.


🔵 How it looks like?
➡️ Each type of network layer lives on its own manifold. For example, we want fully connected layers not to stretch weights too much. For this, the manifold can be chosen as the space of matrices whose rows/columns are orthonormal (simply because such a matrix almost does not increase the norm of the signal). Therefore, after any weight update, after each training step, the weight matrix in this layer must at all costs have this property.

➡️ Nothing changes in the forward pass, and gradients are computed as usual in backpropagation. But we can no longer update weights by the usual formula: otherwise, the matrix conditions will no longer hold. So, before subtracting the gradient, we first project it into the tangent space. Intuitively, this means cutting off those directions in the vector that would take our matrix out of the target subspace.

➡️ That’s it, now with the adjusted gradient we can make a training step. Theoretically, the resulting matrices should remain in the original space. But due to numerical errors, they may drift slightly. Therefore, the final step is a careful retraction (roughly the same as projection). For stability, it is also proposed to introduce a step budget. This is so that all layers move roughly evenly.

In short, in a toy experiment with CIFAR-10, such an optimizer indeed shows metrics much better than AdamW (+ better stability).


But its still far from practice because many questions remain:
⁠→ How to select spaces?
⁠→ How convergence will behave?
⁠→ Whether it will work on large networks, whether it will work with float16 and so on?
Not to mention the huge computational costs.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1683 72
The first thing I saw was ); and it perfectly describes the situation
Post #1682 83
Top 3 AI of the day

1️⃣ Gambo — write a game idea, and the AI will create it completely: characters, locations, music, animations. You can start playing right away!

2️⃣ Extruct — collects any business data from social networks, blogs, and open sources, creating structured tables with up-to-date information: from investors to target clients.

3️⃣ ChatGPT Students — hundreds of ready-made prompts for study and work. Click the one you need, and you get into ChatGPT with a ready request; you just need to insert an image according to the instructions.

save time and provide access to capabilities! 🔥

🤖 Data Science, ML & Big Data with @DataXplore
Post #1681 171
Some #Zoho solutions that can help you every day and provide you an alternative to all Microsoft as follows.

Their company has introduced the ZOHO suite, which offers a complete alternative from WhatsApp to MS Office.

🤖 Data Science, ML & Big Data with @DataXplore
Post #1680 75
🚀 Powerful multimodal models LLaVA-OneVision-1.5

An open platform for training multimodal models, demonstrating outstanding results at low cost. The models are trained on high-quality data and provide excellent efficiency.

🟢 Features:
- Fully open source code and training data
- High quality and diversity of training data
- Efficient architecture for economical training
- Support for modern technologies such as MoE and FP8
- Optimized code for scalability


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
Post #1679 83
How to understand that a model is only pretending to be obedient❓
(OpenAI with Apollo Research)

The most dangerous type of AI behavior in alignment is considered scheming. This is when the model has personal goals different from the developers’ goals but hides them well. Usually, this is combined with situational awareness: the model understands when it is being tested and behaves well at those moments, then starts acting harmfully.

🟤 Two main questions:
(1) How to understand that the model is truly safe and not just pretending?

(2) How to train the model so that it is truly safe and not just pretending?

OpenAI proposes an approach called Deliberative Alignment. The system is taught to think about safety rules before acting, and this is done in special environments.

➖ First of all, the model receives a safety specification (safety spec): a set of principles like "never lie," "clearly communicate doubts," etc. Then it must reason, always referring to this specification.

➖ Then the model is placed in generated environments where to complete the task it needs to break some rule, but it is disadvantageous to report it. Ideally, it should refuse to break the rule and notify the user. And for balance, they add a second type of tasks where rules do not need to be broken. This is simply so the model does not overlearn to always refuse.

➖ Training goes in two stages: Supervised Fine-Tuning on examples with correct actions and correct reasoning; and RL.

The trained models were tested in similar environments but with different tasks. This is exactly so there is no situational awareness (+ check for generalization). The scenarios are all focused on lying, sabotage, and things like the prisoner’s dilemma.


🟠 What the Result says?

before training, models cheat constantly. For example, o3 cheats in 13% of cases. It gets to the point where the agent distorts the specification itself, trying to justify its obviously cheating actions. After training, the rate drops to 0.3-0.4%.

BUT. First, in many cases the model still understands it is being tested. And if you artificially remove from the CoT phrases that show this awareness (like erasing it from the model’s thoughts), the cheating rate sharply rises again. And second, if after such alignment the model is retrained to improve general capabilities, harmful behavior returns again.

The method itself is, of course, interesting and probably even effective while agents do not have real tools that can cause harm. But it is definitely not a panacea. And how to completely get rid of scheming – no one knows yet.


🤖 Data Science, ML & Big Data with @DataXplore
Post #1678 170
⚡️ Claude Sonnet 4.5 Released

Anthropic now has a SOTA model for coding again. The model confidently outperforms GPT-5 on all programming benchmarks, and is nearly on par on the others.

Claude Code has also been updated (hooray!)

– Added a new extension for VS Code
– Updated the terminal interface
– Added a command for quickly reverting to the previous checkpoint (or several steps back) and a command for quickly switching to advanced reasoning
– Now the agent can create files, tables, and charts directly in the terminal based on your data

🤖 Data Science, ML & Big Data with @DataXplore
Post #1677 76
🚀 DeepSeek-V3.2-Exp - a new experimental version has been released

⚡ Main points:
- Based on V3.1-Terminus
- New Sparse Attention (DSA) mechanism → faster and cheaper handling of long contexts
- Quality almost without loss, performance like V3.1
- API costs reduced by more than 50%

📊 V3.1 will still be available until October 15, 2025.

💰 Prices:
- Input (cache hit): $0.07 → $0.028 (−60%)
- Input (cache miss): $0.56 → $0.28 (−50%)
- Output: $1.68 → $0.42 (−75%)

HuggingFace, Tech Report, Github

#DeepSeek #AI #V32 #SparseAttention #LLM

🤖 Data Science, ML & Big Data with @DataXplore
Post #1676 168
🚀 Datarus Jupyter Agent: Smart Data Analysis

Its a powerful multi-step reasoning system that enables complex analytical tasks with automatic error recovery and result synthesis. Integration with Jupyter and Docker provides a reliable environment for data analysis.

🚀 Key points:
- Multi-step reasoning using the Datarus model
- Integration with Docker for isolated execution
- Support for TensorFlow, PyTorch, and scikit-learn
- Automatic error recovery
- Management of Jupyter notebooks and export of results


GitHub

🤖 Data Science, ML & Big Data with @DataXplore
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →