TGViewer
Collective Intelligence Collective Intelligence @co_intelligence · 717 subscribers
Post #436 442
First, we develop a predictive model of delayed rewards that incorporates all information obtained to date. Full observations as well as partial (short or medium-term) outcomes are combined through a Bayesian filter to obtain a probabilistic belief. Second, we devise a bandit algorithm that takes advantage of this new predictive model. The algorithm quickly learns to identify content aligned with long-term success by carefully balancing exploration and exploitation.

https://dl.acm.org/doi/abs/10.1145/3580305.3599386
ACM Conferences Impatient Bandits: Optimizing Recommendations for the Long-Term Without Delay | Proceedings of the 29th ACM SIGKDD Conference on…
More from @co_intelligence
  1. Mar 2, 2026"explainability by construction" vs "post-hoc explainability" https://mineracaodedados.wor…
  2. Dec 21, 2025Human factors engineering improves people’s lives by making technology work well for them.…
  3. Dec 21, 2025photo post
  4. Dec 21, 2025photo post
  5. Dec 21, 2025A systematic review of human-AI interaction in autonomous ship systems
  6. Dec 21, 2025Область исследований, которая изучаю отказоустойчивость и безопасность сложных социо-техни…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →