- Location
- India, IN
- Work mode
- On-site
- Employment
- Full-time
- Experience
- Mid-level
About the role
Hiring: Benjamin RL We’re training the next generation of Vedika models. This role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning. You’ll work on: • RL training for next-generation Vedika models • GRPO, PPO, DPO and newer post-training methods • Reward models, process rewards and verifiers • Long-horizon reasoni…
Requirements
- experience with reinforcement learning
- experience with RLHF or post-training research
- experience with policy optimization methods
Skills
- reinforcement learning
- rl training
- reward models
- rlhf
- policy optimization
- grpo
- ppo
- dpo
- long-horizon reasoning
- verifiers
- process rewards
About the company
Vedika API
Posted via Adzuna
How to apply
Apply Now takes you to Rozgoo, where auto-apply can submit your application for this role. Updated 8 days ago.
Apply Now