We are looking for an RL Environment Data Engineer / Researcher Intern to support our team in building and refining reinforcement learning training environments across different domains. Working alongside our researchers and engineers, you will assist with data collection, task definition, reward design, evaluation, and checking how well environment data works in post-training. This internship suits students and early-career researchers who want hands-on experience with RL environments and LLM post-training.
Responsibilities
- Assist in designing and improving RL training environments across various task domains.
- Help collect, clean, structure, and evaluate data used for RL environment construction and model post-training.
- Support the team in defining task objectives, reward functions, and evaluation standards.
- Help identify loopholes in reward design and test approaches that prevent reward hacking.
- Run experiments in validation environments and report on the effectiveness of post-training data and environment design.
- Work with research, engineering, and data team members to improve environment coverage, task difficulty, and evaluation reliability.
- Keep up with research on RL environments, data evaluation, AI agents, and post-training methods, and share relevant findings with the team.
Requirements
- Currently pursuing or recently completed a degree in Computer Science, AI, or a related field.
- Good coding skills in Python, with the ability to build scripts, data pipelines, and simple evaluation tools.
- Comfortable using AI coding tools for code generation, debugging, and rapid experimentation.
- Foundational understanding of reinforcement learning, post-training, reward design, or data evaluation, through coursework, research, or projects.
- Interest in turning real-world tasks into trainable and measurable RL environments.
- Experience with data scraping, data cleaning, annotation, or data quality assessment is a plus.
- Exposure to LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a plus.
- Curious and eager to learn, with the ability to iterate quickly based on feedback and experiment results.