Engineering Manager, Cloud Platform
Manage a team of cloud platform engineers building scalable, reliable infrastructure. Drive technical decisions, coach engineers, and partner cross-functionally on ML infrastructure priorities.
About the job
Responsibilities
- Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s
- Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work
- Foster a culture of ownership, accountability, and continuous improvement
- Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate the infrastructure roadmap into clear team priorities and milestones
- Establish standards and best practices for reliability, performance, and operational excellence; ensure the team owns projects end-to-end
- Exercise and encourage good judgment on tooling and architectural tradeoffs, with a bias against unnecessary complexity
- Partner with product and engineering teams to align infrastructure work with business priorities
- Serve as the escalation point for major incidents; drive resolution with urgency and ensure the team learns systematically
Requirements
- Proven experience directly managing engineers in a cloud infrastructure, platform engineering, or SRE context (not managing managers)
- Extensive hands-on experience with Kubernetes, with the ability to engage credibly in architectural discussions and code reviews
- Strong background building and maintaining scalable, production infrastructure
- Familiarity with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Pulumi) and CI/CD tooling (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins)
- Demonstrated ability to drive projects end-to-end — from specification to execution — and coach your team to do the same
- Strong written and verbal communication skills; able to influence across teams without direct authority
- Track record of recruiting and growing engineering talent
- Openness to learning about ML infrastructure
Nice to Have
- Experience with OSS observability tooling (Prometheus, ELK stack, Grafana, OpenTelemetry)
- Background in multi-cloud environments or GPU infrastructure
- Experience building and managing on-call rotations and incident response processes
Benefits
- Competitive compensation, including meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
Skills
Kubernetes, Terraform, CloudFormation, Pulumi, GitHub Actions, Gitlab Ci, CircleCI, Jenkins, Prometheus, Grafana, OpenTelemetry
Similar jobs
Engineering Management jobsLeads a hands-on Shared Services Engineering team building and operating reusable services, SDKs, APIs, and customer-facing systems. Requires 7+ years of software engineering experience, engineering management experience, and strong technical judgment across distributed and full-stack systems.
Leads commissioning programs across multiple data center projects, managing commissioning teams, third-party agents, and stakeholder coordination from pre-functional testing through turnover. Requires 10+ years of mission-critical commissioning experience, 5+ years of leadership, engineering knowledge, and a bachelor’s degree.
Leads a hands-on platform engineering team responsible for AWS infrastructure, Kubernetes deployment paths, developer self-service, CI/CD governance, reliability, and audit readiness. The role requires deep infrastructure experience, Terraform expertise, production Kubernetes operations, and people leadership.
Leads the architecture and development of BuildOps’ integration and data platform, driving technical strategy, reliability, APIs, databases, and engineering standards across multiple teams. Requires 10+ years of software engineering experience and deep expertise in TypeScript, Node.js, PostgreSQL, cloud infrastructure, and distributed systems.
Leads 24/7 critical facility operations for a 20 MW AI data center expanding to 40 MW, overseeing electrical, mechanical, safety, maintenance, staffing, and operational performance. Requires at least five years of data center operations management experience and expertise in infrastructure reliability and team scaling.