🟢 What's new?
Reinforcement Learning with Evolving Rubrics (RLER) for long, unverifiable tasks.
💡 Instead of static evaluations:
• Rubrics evolve together with the model
• Use knowledge from search
• Extract new information directly during training
📊 Results:
• DR Tulu‑8B is comparable to OpenAI DR
• Surpassed all open-source DR models
• Cost — ~ $0.00008 per query (compared to > $1 at OpenAI)
💥 Training in two stages: SFT → RL
Tested on 4 complex benchmarks and a new medical GeneticDiseasesQA (in collaboration with clinicians) — results better than OpenAI DR and AI2 ScholarQA (Claude).
Open methodology, real impact. AI that *learns to research by itself*.
Model & Code
••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
