Staff Engineer, Compute
Leads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.
About the job
Responsibilities
- Lead the technical direction of capacity management and workload placement for Datadog's Kubernetes platform spanning 100,000+ virtual machines across multiple cloud providers.
- Design and build systems that optimize how engineering workloads are scheduled and deployed across regions while balancing capacity constraints, reliability, and performance.
- Partner across infrastructure teams to evolve multi-region and multi-cloud capacity orchestration as Datadog continues to scale.
- Develop production software in Go to improve Kubernetes platform capabilities, automation, and operational efficiency.
- Use data and capacity signals to influence infrastructure decisions, forecast growth, and improve workload placement strategies.
Requirements
- Significant experience designing and operating large-scale Kubernetes-based infrastructure or platform systems.
- Strong programming skills, ideally in Go or a comparable systems programming language.
- Hands-on experience with at least one major cloud provider: AWS, Google Cloud, or Azure.
- Understanding of distributed cloud infrastructure.
- Strong systems mindset and interest in complex infrastructure, scheduling, or capacity management challenges.
- Comfort working with operational data, capacity forecasting, or analytical approaches that inform engineering decisions.
- Proven track record of providing technical leadership across teams and influencing architecture without direct management authority.
Benefits
- Competitive benefits package.
- New-hire stock equity (RSUs) and employee stock purchase plan.
- Continuous career development and pathing opportunities.
- Employee-focused onboarding.
- Internal mentor and cross-departmental buddy program.
- Friendly and inclusive workplace culture.
Benefits may vary based on the country of employment and the nature of employment.
Skills
Kubernetes, Go, AWS, GCP, Azure, Distributed Systems, Cloud Infrastructure, Capacity Management, Workload Scheduling, Infrastructure Automation
Similar jobs
DevOps / SRE jobsLeads the design, development, and operation of Stripe’s large-scale CI and developer productivity systems. Requires 10+ years of hands-on software development, distributed-systems expertise, technical leadership, and mentoring experience.
The Staff Production Engineer will operate and improve reliable production infrastructure, with a strong focus on data center networking, capacity planning, troubleshooting, automation, and hardware operations. The role requires 5+ years of production engineering experience, networking expertise, and proficiency with infrastructure and CI/CD tools.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
Staff Site Reliability Engineer responsible for the performance, scalability, deployment robustness, vulnerability management, and incident response of a large-scale cloud data platform. The role requires deep Kubernetes and multi-cloud expertise, infrastructure automation, scripting, Linux administration, and cloud networking experience.