Senior Manager, Infrastructure Platform Engineering
Lead a team building core infrastructure platform services for large-scale compute capacity allocation, state management, and security. Requires 10+ years infrastructure experience and 3+ years in engineering leadership.
About the job
What You'll Be Working On
- Leading the team responsible for the platform services that abstract underlying infrastructure into reliable, allocatable capacity, and for the systems that track and reconcile state across a large fleet
- Setting the technical roadmap across capacity and utilization intelligence, resource lifecycle and state management, and platform security and trust frameworks
- Driving the design of secure, well-instrumented platform systems — from Kubernetes-based orchestration and automation to lower-level system and hardware integration
- Hiring, mentoring, and growing a team of infrastructure software engineers; building a high-performing organization from a strong foundation
- Partnering with infrastructure, production engineering, and security teams to align platform capabilities with operational reliability, capacity, and trust requirements
- Improving platform efficiency and availability — characterizing bottlenecks, reducing stranded resources, and shortening operational and recovery cycles
- Establishing engineering standards for infrastructure software development: code quality, testing, deployment safety, and on-call practices for systems that span the platform
- Translating a vertically integrated infrastructure stack into reliable platform primitives that engineering teams can build on
- Staying technically hands-on — reviewing designs, contributing to architecture decisions, and being credible to the engineers you lead
What You'll Bring to the Team
- 10+ years of experience in infrastructure or systems software development, with at least 3+ years in an engineering leadership role
- Deep expertise in large-scale infrastructure platforms — building services that pool, allocate, and reconcile compute resources at scale
- Strong background with Kubernetes and cloud platforms (GCP, AWS, or Azure) — orchestration, automation, and operating distributed systems in production
- Experience with distributed state management and control systems — modeling resource and system lifecycle, reconciling desired vs. actual state, and handling failure gracefully across a large fleet
- Experience with efficiency, capacity, or performance engineering — characterizing system behavior, identifying bottlenecks, and driving measurable improvements in utilization or availability
- A player-coach approach to management: hands-on enough to make technical calls, structured enough to grow a team and ship through them
- Track record of hiring strong infrastructure engineers and helping them grow into more senior roles
- Comfortable operating in a fast-moving environment where the path isn't fully paved — willing to drive ambiguity to clarity
Bonus Points
- Experience operating Kubernetes on bare-metal infrastructure as well as on managed cloud services (GKE, EKS, AKS)
- Familiarity with the operational challenges of GPU clusters, AI training, and inference workloads
- Working knowledge of platform security and trust concepts — secure boot, measured boot, TPMs, and hardware attestation
- Experience with capacity forecasting, demand modeling, or allocation optimization at scale
- Hands-on background with telemetry and observability platforms at scale (Prometheus, OpenTelemetry, Grafana)
- Prior experience building infrastructure platforms at hyperscalers or cloud providers where internal engineers are the primary customer
- Familiarity with hardware-software co-design — understanding how platform choices affect physical infrastructure utilization
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Skills
Kubernetes, GCP, AWS, Azure, Infrastructure, Systems Software, Distributed Systems, Capacity Planning, Performance Engineering, Platform Security, Observability, Prometheus, OpenTelemetry, Grafana, Bare Metal
Similar jobs
Engineering Management jobsLeads a hands-on product engineering team building and operating polished, scalable software products. Requires 10+ years of software engineering experience, including 5+ years managing teams, with depth in full-stack development, distributed systems, reliability, and product quality.
Leads a team of Forward Deployed Engineers delivering and integrating physical AI systems for defense and government customers. Requires 8+ years of industry experience, recent technical leadership, production deployment experience, broad software and hardware expertise, and U.S. security-clearance eligibility.
Leads and grows the ML Platform team responsible for training, evaluating, and serving production machine-learning models. The role requires 8+ years of engineering experience, 4+ years managing engineering teams, strong ML systems expertise, and customer-facing technical leadership.
Leads and scales engineering teams responsible for high-throughput social-data systems, real-time AI agents, analytics, and automation. The role combines people management, product execution, architecture oversight, delivery excellence, hiring, and cross-functional collaboration.
Leads a mixed software and ML engineering team building simulation, scenario authoring, generative modeling, and agentic tools for autonomous vehicle validation. The role owns technical direction, cross-functional delivery, team development, and generative scenario creation at company scale.