TGViewer
Channel Public Channel
For Developers

For Developers

@fordevelopers

YAC
Subscribers
196
Photos
29
Videos
3
Links
767
Recent Posts 20 shown
Post #2276 110
Post #2274 94
#meta_learning #pb2 #rl

Meta-learning Population-based Methods for Reinforcement Learning
https://openreview.net/forum?id=d9htascfP8

Abstract: Reinforcement learning (RL) algorithms are highly sensitive to their hyperparameter settings. Recently, numerous methods have been proposed to dynamically optimize these hyperparameters. One prominent approach is Population-Based Bandits (PB2), which uses time-varying Gaussian processes (GP) to dynamically optimize hyperparameters with a population of parallel agents. Despite its strong overall performance, PB2 experiences slow starts due to the GP initially lacking sufficient information. To mitigate this issue, we propose four different methods that utilize meta-data from various environments. These approaches are novel in that they adapt meta-learning methods to accommodate the time-varying setting. Among these approaches, MultiTaskPB2, which uses meta-learning for the surrogate model, stands out as the most promising approach. It outperforms PB2 and other baselines in both anytime and final performance across two RL environment families.
Submission Length: Long submission (more than 12 pages of main content)
Code: https://github.com/automl/MetaPB2
Assigned Action Editor: Mirco Mutti
Submission Number: 3679


Accepted by TMLR
GitHub GitHub - automl/MetaPB2 Contribute to automl/MetaPB2 development by creating an account on GitHub.
Post #2264 139
Post #2263 124
#grpo #vs #dpo #reinforcement_learning #rl #llm #qwen
It Takes Two: Your GRPO Is Secretly DPO
https://arxiv.org/abs/2510.00977

#benchmark #timeseries #msIC #msIR #kaggle #btcf #btc #gsmi #options #sota #CSI #etf #AAAI #NeurIPS #openreview #ICRL
FinTSBridge: A New Evaluation Suite for Real-world Financial Prediction with Advanced Time Series Models
https://openreview.net/forum?id=6UHEfOBVkn
arXiv.org It Takes Two: Your GRPO Is Secretly DPO GRPO has emerged as a prominent reinforcement learning algorithm for post-training LLMs. Unlike critic-based methods, GRPO computes advantages by estimating the \emph{value baselines} from...
Post #2262 120
Post #2261 106
Older posts →

About this channel

How can I read @fordevelopers without a Telegram account?
TGViewer shows the public web preview Telegram publishes for For Developers: recent posts, photos, videos and the subscriber count, with no app, login or account.
How many subscribers does For Developers have?
For Developers (@fordevelopers) has 196 subscribers on Telegram, refreshed roughly every 30 minutes.
Does For Developers know I viewed it here?
No. Public channel previews carry no viewer identity, and TGViewer has no accounts or tracking of what you look up.
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →