Director of Engineering, Workflows & Runners
Lead multi-team engineering effort building scalable workflow and runtime infrastructure at GitLab.
The role
This director-level role owns technical strategy and delivery for core Workflow and Function Runtime systems serving as shared infrastructure across GitLab's product groups. The position combines hands-on technical leadership with people management, requiring direct involvement in architecture decisions while building and scaling high-performing engineering teams. The role sits at the intersection of platform engineering and organizational leadership within GitLab's DevSecOps platform.
What you'd do
- Establish technical vision and roadmap for Workflow and Function Runtime primitives in collaboration with Product, Architecture, and platform teams
- Lead and grow multiple engineering teams delivering highly available, horizontally scalable, and secure distributed systems
- Actively participate in architecture reviews, design decisions, incident investigation, and critical technical contributions
- Drive operational excellence through SLOs, on-call processes, incident response, capacity planning, and multi-region resilience
- Optimize cost efficiency and performance balance for long-running and bursty workloads at scale
- Define platform APIs and abstractions enabling product teams to safely compose workflows and integrate with runtime
- Build engineering culture emphasizing results, iteration, ownership, and technical rigor
- Recruit, develop, and retain senior and staff-level engineers and managers with clear growth paths
- Collaborate cross-functionally with Product, Security, Infrastructure, Data, AI/ML, and stage groups
What they're looking for
- Proven experience leading teams building platform or infrastructure services such as workflow engines, function runtimes, or high-scale microservices
- Track record as hands-on engineering leader with meaningful contributions to architecture, design, and code
- Strong background in scalable multi-tenant distributed systems including service decomposition, data partitioning, and fault tolerance
- Demonstrated success operating mission-critical production services with SLOs, on-call, incident management, and capacity planning
- Experience driving cost efficiency in cloud-native environments through profiling, tuning, and optimization
- Familiarity with Kubernetes, modern cloud infrastructure, and event-driven or asynchronous architectures
- People leadership experience managing managers and senior/staff engineers with hiring and coaching track record
- Ability to work effectively in fully remote, globally distributed environment with strong async communication
- Minimum 10 years professional software engineering experience including 4-6+ years in multi-team or organization-level engineering leadership
Nice to have
- Experience with function-as-a-service or serverless runtimes
- Exposure to chaos engineering and disaster recovery testing
Summary written by RoleDeck from the original posting. This is an extracted, own-words summary and may contain errors. The original source may have changed or expired.