TGViewer
Channel Public Channel
Linkstream

Linkstream

@linkstream

Various links I find interesting. Mostly hardcore tech :) // by @oleksandr_now. See @notatky for the personal stuff
Subscribers
169
Photos
37
Videos
3
Links
965

Showing posts older than #936 · Back to latest

Older Posts 20 shown
Post #933 155
lol perfect timing 😅 4h later Pavel Durov announced Cocoon: https://t.me/durov/462

my 2c: it is a fine business, sadly only a small part of what Anima, Cortex and the minds need.
Gentian proxy is still required, etc etc.
They do acknowledge the limitations of RA-TLS and their model in general though, which is commendable.
Telegram Pavel Durov 🐣 It happened. Our decentralized confidential compute network, Cocoon, is live. The first AI requests from users are now being processed by Cocoon with 100% confidentiality. GPU owners are already earning TON. https://cocoon.org is up, with docs and the source…
  • 🤯 1
Post #931 212
Claim: Isotropic Gaussian Regularization for latent representations in the world models is mathematically optimal

What's illustrated:
-- Adopting the isotropic gaussian regularization replaces stop-grad, teacher-student, EMA and various other adhoc tricks
-- Improves model training stability
-- SOTA quality on 10+ datasets and 50+ architectures
https://arxiv.org/abs/2511.08544, https://github.com/rbalestr-lab/lejepa
arXiv.org LeJEPA: Provable and Scalable Self-Supervised Learning Without the... Learning manipulable representations of the world and its dynamics is central to AI. Joint-Embedding Predictive Architectures (JEPAs) offer a promising blueprint, but lack of practical guidance...
  • 🤯 2
  • 🔥 1
Post #925 397
aaargh i should've wrote this paper!! it was intuitively obvious to me but then life happens >_<

tldr: LLM sampler is such a powerful prior that with the right sampler (MCMC, in this case), you can even use base models as reasoning models.
without supervised fine-tuning or RL.

this was completely ignored by ppl pilled with the Bitter Lesson mantra, but yes there still is a space for the right priors added or designed by hand!

obviously sampling with mcmc is very costly but you should compare the overall model feedback loop time that includes the posttrain, not just the sampling time

if eg topK sampling is assembly and Mirostat is COBOL (?) then MCMC sampling is like a Python in the space of samplers
https://aakaran.github.io/reasoning_with_sampling/
Post #923 176
Post #919 176
Older posts →
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →