logo
  • 环境
  • 企业版
  • 价格

Open Source Cowork 桌面版

Eigent 是一款开源产品,这意味着你可以使用自己的 API 密钥或本地模型免费自行托管。

Eigent

获取关于 AI 员工自动化的最新更新和教程。

产品Eigent环境定价企业版
探索解决方案使用场景技能插件博客
开发者文档GitHubCAMEL-AI开源基金合作伙伴
下载开源版
公司关于我们品牌加入我们使用条款隐私政策安全与信任Cookie 政策退款与试用政策

保留所有权利 © 2026 EIGENT UK LTD

Eigent 1.0 新版本发布!download
All roles

RL Environment Data Engineer / Researcher

Location

London / Bay Area / Remote

Employment Type

Full-time

Department

Engineering

We are looking for an RL Environment Data Engineer / Researcher to design, build, and refine reinforcement learning training environments across different domains. This role will focus on data collection, task definition, reward design, evaluation criteria, anti-reward-hacking mechanisms, and post-training validation of environment data effectiveness.

Responsibilities

  • - Design and improve RL training environments across various task domains.
  • - Collect, clean, structure, and evaluate data used for RL environment construction and model post-training.
  • - Define task objectives, reward functions, and evaluation standards to ensure reliable and reproducible training signals.
  • - Develop technical approaches to prevent reward hacking and identify loopholes in reward design.
  • - Build validation environments to assess the effectiveness of post-training data and RL environment design.
  • - Collaborate with research, engineering, and data teams to improve environment coverage, task difficulty, and evaluation reliability.
  • - Follow research progress in RL environments, data evaluation, AI agents, and post-training methods, and apply relevant findings to production workflows.

Requirements

  • - Strong coding skills, especially in Python, with the ability to independently build data pipelines, environments, and evaluation tools.
  • - Proficiency with AI coding tools for code generation, debugging, refactoring, and rapid experimentation.
  • - Solid understanding of reinforcement learning, post-training, reward function design, environment design, and data evaluation.
  • - Ability to translate real-world tasks into trainable and measurable RL environments.
  • Experience with data scraping, data cleaning, annotation, or data quality assessment is preferred.
  • - Experience with LLM agents, RLHF/RLAIF, coding agents, automated evaluation, or benchmark construction is a strong plus.
  • - Strong experimental mindset and engineering execution, with the ability to continuously improve systems based on data and evaluation results.