June 2026 archive · For research and skill exploration. Current job availability is unverified.
OpenAI
Software Engineer, RL Training Infra
Keep frontier RL training runs reliable and fast by solving critical infrastructure and engineering bottlenecks.
OpenAIText summary from the June 2026 archive. original source →
The role
This role supports OpenAI's Post-Training Frontiers team, which develops reinforcement learning systems for frontier agents deployed across Codex, ChatGPT, and the API. The engineer acts as a generalist troubleshooter, tackling urgent infrastructure problems that emerge during large-scale RL training runs, from distributed system failures to inference bottlenecks and scaling challenges.
What you'd do
- Unblock large-scale RL training runs by addressing urgent engineering and infrastructure bottlenecks
- Debug issues spanning training systems, inference, orchestration, scaling, and distributed infrastructure
- Solve technical problems at the intersection of research and engineering, including scaling experiments and reliability improvements
- Support researchers building infrastructure-heavy features like multi-agent capabilities and memory systems
- Convert recurring operational issues into improved tools, systems, and abstractions
- Collaborate with research, infrastructure, and partner teams during critical model training timelines
- Debug failures across model behavior, training data, RL systems, evaluation infrastructure, and serving systems
What they're looking for
- Strong generalist engineer with experience in at least one layer of ML infrastructure
- Background in RL, inference, scaling, training systems, orchestration, or similar ML infrastructure domains
- Ability to learn rapidly and operate across unfamiliar technical layers
- Strong debugging skills with high ownership and excellent communication
- Comfort operating in ambiguous, fast-moving environments with tight timelines
- Motivated by reliability, speed, and building load-bearing systems
Nice to have
- Experience supporting large-scale model training or async RL systems
- Background debugging distributed systems across GPUs, networking, orchestration, or inference
- Performance optimization or production-critical infrastructure expertise
- Direct experience working with research teams or fast-moving model development groups
Summary written by RoleDeck from the original posting. This is an extracted, own-words summary and may contain errors. The original source may have changed or expired.