TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 585 subscribers
Post #1728 121
Mamba-3 quietly and without announcement was released at ICLR - and this could mark the beginning of the end of the Transformer era.

🟠 Brief excursus of new Mamba-3
☞ Makes models faster, more stable, and more efficient when working with long contexts.

The main idea is not in attention layers, but in state-space models, where the model stores and updates internal state over time.

- Mamba-1 introduced continuous dynamics and selective memory updates - it remembered efficiently without the high cost of attention.
- Mamba-2 showed that state updates and attention are two sides of the same mathematics, which sped up computations on GPUs.
- Mamba-3 brought the concept to maturity: now the internal memory evolves more smoothly and stably by moving from a simple Euler step to trapezoidal integration.

Instead of a simple Euler step as in Mamba-2, Mamba-3 approximates the integral of the state update not only at the right end of the interval but by averaging between the start and end, with a coefficient λ depending on the data. This provides a more accurate (second-order) approximation and makes the state dynamics more expressive.


🟢 What actually changed?
- Memory became "rhythmic": now the model can store repeating and periodic patterns (e.g., language or music structures).

- The new multi-input-multi-output design allows processing multiple streams in parallel — ideal for modern GPUs.

⚙️ What this means in practice:
- Efficient handling of long sequences: documents, genomes, time series.

- Linear runtime and stable latency make it perfect for real-time applications: chatbots, translation, speech.

- Energy efficiency and scalability open the way to on-device AI, where large models run locally without the cloud.


Mamba-3 is not just an accelerated alternative to Transformers combines deep contextual understanding, speed and stability. from server systems to smart devices.

#ssm #mamba3 #llm,#architecture #ai

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →