Capacity Ops Engineer
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
About the job
Responsibilities
- Lead specialized GPU pods, owning acquisition, orchestration, and maintenance across the asset lifecycle.
- Execute workload migrations and deployment drains while meeting regional and compliance requirements.
- Design and implement a scalable capacity management system supporting a 10x increase in GPU volume.
- Build ROI and unit-economics models for GPU spending and capacity trade-offs.
- Partner with SRE, infrastructure, and forward-deployed engineering teams on operational execution and infrastructure changes.
- Lead responses to capacity-related incidents and outages.
Requirements
- Bachelor's, master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
- 5+ years of professional experience in a high-growth environment, preferably at a hyperscaler or specialized GPU provider.
- Deep expertise in Kubernetes, including taints, cordons, node draining, and custom operators.
- Production-level experience with Go or Python.
- Strong financial literacy and ability to model capacity reliability and cost trade-offs.
- High tenacity and a collaborative mindset.
Benefits
- Competitive compensation, including meaningful equity.
- Medical, dental, and vision insurance fully covered for employees and dependents.
- Flexible paid time off and company-wide winter break.
- Paid parental leave.
- Fertility and family-building stipend.
- Company-facilitated 401(k).
Skills
Kubernetes, Go, Python, Multi-Cloud, Gpu Infrastructure, Nvidia Blackwell, Nvidia H100, Custom Operators, Node Draining, Financial Modeling
Similar jobs
DevOps / SRE jobsProduction Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.