Principal Site Reliability Engineer
Leads reliability strategy, architecture, and operational excellence for large-scale cloud products, while building automation and internal platforms. The role requires deep Kubernetes, cloud infrastructure, distributed systems, observability, and reliability engineering expertise, plus strong cross-organizational technical leadership.
About the job
Responsibilities
Reliability Strategy & Architecture
- Define and drive the reliability strategy for critical product and platform services.
- Establish standards for availability, resilience, observability, incident management, and operational readiness.
- Lead architecture reviews for critical services and platform initiatives.
- Align reliability objectives with business priorities and customer expectations.
- Create frameworks, standards, and operational guardrails for safe operation at scale.
- Guide service architecture toward simplicity, scalability, resilience, and operational excellence.
- Drive initiatives that improve platform maturity and long-term sustainability.
Product & Platform Leadership
- Own reliability architecture and operational excellence for the Spera / Identity Security Posture Management (ISPM) product area.
- Establish reliability objectives and technical roadmaps with engineering leadership.
- Lead scalability, resiliency, and performance initiatives.
- Build self-service operational capabilities that improve developer productivity while strengthening reliability and security.
- Influence technical direction through data-driven recommendations, engineering expertise, and collaborative leadership.
- Support highly available, large-scale cloud environments as part of an on-call rotation.
Engineering & Automation
- Design, build, and operate large-scale cloud infrastructure and production services.
- Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
- Eliminate operational toil through automation, tooling, and platform engineering.
- Improve deployment safety and operational workflows through GitOps and infrastructure-as-code practices.
- Modernize existing workloads and align them with evolving platform capabilities.
- Lead engineering initiatives from conception through production rollout and long-term operational ownership.
Technical Leadership
- Mentor Staff and Senior engineers across multiple teams and organizations.
- Lead technical reviews, design reviews, and operational readiness assessments.
- Build engineering consensus across teams with differing priorities and objectives.
- Develop the next generation of technical leaders.
- Drive adoption of reliability engineering best practices across the Emerging Products Group.
- Share patterns, tooling, and operational practices across Workflows, Inbox, PAM, and ISPM teams.
- Influence technical direction through expertise, collaboration, and execution rather than organizational authority.
AI & Agentic Operations
- Explore and adopt AI-assisted reliability engineering practices.
- Design and champion agentic systems for troubleshooting, incident response, root-cause analysis, and operational decision-making.
- Evaluate emerging AI technologies and identify opportunities to improve reliability engineering workflows.
- Establish safe, effective, and measurable practices for AI in production operations.
- Reduce operational toil and improve engineering productivity through intelligent automation.
Technical Requirements
- Extensive experience designing and operating large-scale production systems in AWS and/or GCP.
- Deep expertise with Kubernetes in production environments.
- Experience designing reliability strategies for Kubernetes-based platforms.
- Expertise troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle challenges.
- Extensive experience with infrastructure-as-code technologies such as Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building internal platforms, developer tooling, and operational automation.
- Deep understanding of distributed systems architecture and cloud-native application design.
- Understanding of cloud networking fundamentals, including DNS, service discovery, ingress, load balancing, TLS, traffic management, and multi-region architectures.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Experience establishing observability standards, monitoring strategies, and operational best practices.
- Experience with or strong interest in AI-assisted engineering and operational automation.
- Strong expertise operating customer-facing production systems at scale.
- Deep understanding of reliability engineering principles, including SLIs, SLOs, error budgets, capacity planning, and resilience engineering.
- Experience leading major incident response efforts and driving long-term operational improvements.
- Strong understanding of CI/CD, GitOps, deployment strategies, and automation-first operations.
- Proven success driving reliability transformations and architectural modernization efforts.
- Ability to balance reliability, scalability, security, and delivery needs.
Skills
AWS, GCP, Kubernetes, Terraform, Helm, GitOps, Argo CD, Go, Python, Datadog, Splunk, Postgres, Redis, Opensearch, Distributed Systems
Similar jobs
DevOps / SRE jobsAs a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.