Head of Engineering, Infrastructure & SRE
Leads the infrastructure and SRE organization responsible for Postman’s cloud-agnostic, high-scale platform. The role combines distributed team leadership with technical ownership of reliability, Kubernetes-based infrastructure, multi-cloud operations, incident management, and engineering autonomy.
About the job
Responsibilities
- Lead and grow a geographically distributed infrastructure and SRE organization across the SF Bay Area, India, and Europe.
- Own the architecture, evolution, reliability, scalability, performance, and cost efficiency of cloud-agnostic infrastructure.
- Lead infrastructure improvements across Kubernetes, Cluster API, Argo, Helm, Crossplane, Istio, AWS, and Azure.
- Own SRE practices including SLIs/SLOs, error budgets, capacity planning, incident management, on-call, monitoring, escalation, and blameless postmortems.
- Set technical direction for infrastructure and reliability supporting hundreds of services and multiple engineering teams.
- Partner with platform, security, product engineering, product management, and other stakeholders on technical roadmaps and trade-offs.
- Own the infrastructure and reliability roadmap, milestones, delivery, and cross-team issue resolution.
- Drive GitOps, CI/CD, observability, load testing, automation, toil reduction, security, cost management, and infrastructure quality.
Requirements
- 15+ years of experience in infrastructure, platform, or site reliability engineering.
- 7+ years in engineering management or leadership roles.
- Experience managing geographically distributed teams across multiple time zones.
- Production experience with Kubernetes, Cluster API, Argo, Helm, and Crossplane at scale.
- Experience with service mesh technologies, including Istio.
- Deep experience with AWS and Azure cloud infrastructure.
- Experience building or leading an SRE function, including on-call, incident management, SLIs/SLOs, and error budgets.
- Experience designing and implementing cloud-agnostic infrastructure that enables engineering autonomy.
- Experience supporting multiple product teams, hundreds of services, high deployment frequency, and significant traffic scale.
- Excellent communication skills for technical and non-technical audiences.
Nice-to-Haves
- GitOps workflows and infrastructure-as-code at scale.
- FinOps and cloud cost optimization practices.
- Experience in a SaaS or API-focused product company.
- Passion for developer experience and enabling product engineers to work faster and more confidently.
Compensation & Benefits
- Pay-for-performance philosophy.
- Flexible schedule.
- Comprehensive benefits, including full medical coverage.
Skills
Kubernetes, Cluster Api, Argo, Helm, Crossplane, Istio, AWS, Azure, GitOps, Infrastructure As Code, CI/CD, Observability, Slis/Slos, Error Budgets, Capacity Planning
Similar jobs
Engineering Management jobsLeads the product and engineering strategy for AI-native enterprise applications across Finance, HR, and Legal, partnering with executives to transform workflows into intelligent products. Requires 15+ years in enterprise technology, product management, or software engineering and substantial multidisciplinary leadership experience.
Leads the organization responsible for deploying, sustaining, supporting, and improving integrated Hivemind software and hardware products in customer environments. Requires 15 years of technical operations or lifecycle leadership experience, systems engineering expertise, and experience building operational organizations.
Leads senior engineering teams responsible for ZoomInfo’s operational and analytical data platforms, including Bigtable, BigQuery, ingestion pipelines, CDC, reliability, and compliant data deletion. Requires 8+ years of software, platform, or data engineering experience and direct leadership of senior technical talent.
Leads the quality and reliability organization for mission-critical aerospace communication systems, establishing the QMS, driving AS9100 certification, and embedding quality and reliability across product development and production. Requires 8+ years in quality or reliability engineering and experience in high-reliability industries.
Own and improve technical integrations between Bilt and property partners, translating operational needs into requirements, configurations, testing, and scalable workflows. The role requires integration troubleshooting, cross-functional issue resolution, and clear communication of technical concepts to nontechnical stakeholders.