simulation95
All roles

Reinforcement Learning

Senior Reinforcement Learning Environment Engineer

Location
Menlo Park, CA (San Francisco Bay Area)
Working arrangement
Hybrid
Employment type
Full-time
Compensation
$400,000 annual base salary; $600,000 total annual compensation

About Simulation95

Simulation95 is building at the intersection of artificial intelligence, healthcare, and life sciences. We believe meaningful progress starts with people who understand the hardest problems firsthand. We’re bringing together experts, researchers, and builders to turn that understanding into technology that matters.

For curious people who take their craft seriously, this is an opportunity to help shape something early—with challenging problems, meaningful responsibility, and room to make a lasting contribution.

About the role

You will build environments in which AI agents can learn, use tools, and complete complex tasks. The work combines software engineering with reinforcement learning: defining what an agent can observe and do, how a task progresses, and how its performance is measured.

We are looking for an engineer who can take a loosely defined problem and turn it into a reliable environment for training and evaluation. You will work with machine learning engineers to make these environments reproducible, observable, and efficient to run at scale.

Responsibilities

  • Build interactive environments with clearly defined observations, actions, state transitions, and success criteria.
  • Develop task generators, reward functions, and reliable verification systems.
  • Create sandboxed tools and interfaces for agent workflows involving multiple steps.
  • Make environments easy to reset, reproduce, inspect, and run at scale.
  • Identify reward exploitation, information leakage, and environment failures that could distort training or evaluation results.
  • Connect environments to training and evaluation pipelines in collaboration with machine learning engineers.

Qualifications

  • Strong Python skills and experience building reliable software systems.
  • Experience with reinforcement learning environments, simulation, agent infrastructure, or automated evaluation.
  • Practical understanding of reward design, partial observability, and long-horizon tasks.
  • Experience with APIs, containers, testing, and concurrent or distributed execution.
  • Ability to define an approach to an ambiguous problem and carry it through implementation and delivery.

Preferred experience

  • LLM agents, tool use, or procedural task generation.
  • Distributed rollout systems or contributions to reinforcement learning and simulation frameworks.