Hugging Face (Twitter)
RT @scaling01: DeepSeek is back!
"Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models"
They introduce Engram, a module that adds an O(1) lookup-style memory based on modernized hashed N-gram embeddings
Mechanistic analysis suggests Engram reduces the need for early-layer reconstruction of static patterns, making the model effectively "deeper" for the parts that matter (reasoning)
Paper: https://github.com/deepseek-ai/Engram/blob/main/Engram_paper.pdf
Post #2196
30


