OpenTinker an open-source RL-as-a-Service is your solution:
You design agent locally and training and inference can easily be offloaded to remote GPUs. No hassle with infrastructure, no rigid coupling of agent logic and execution.
🟢 Why it matters & How?
- you can prototype RL tasks locally without worrying about hardware
- all the heavy lifting - training and inference - is done on cloud GPUs
- supports single-turn and multi-turn tasks
- the trained model can be immediately deployed for inference, without additional code
How it works?
- a disaggregated architecture
- a lightweight client runs locally
- experiments are sent to a cloud scheduler
- the scheduler matches with available GPUs and orchestrates tasks based on resources
- the task is launched remotely, and metrics are streamed in real time to the dashboard
API for developers:
- wrap the environment, reward, and policy once
- OpenTinker handles data loading, rollouts, training, and inference itself
Familiar interfaces:
- for the environment: env.reset() and env.step()
- for training: high-level fit() - a complete end-to-end training loop
- under the hood, fit() is composed of train_step(), validate(), and save_checkpoint()
- want it fast - use fit()
- need control - customize the steps manually
The agent's runtime looks like a state machine:
PENDING - preprocessing and tokenization
GENERATING - the model generates a response
INTERACTING - the agent acts in the environment
TERMINATED - the task is completed
Pure, scalable agent-based RL.
Repository
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore