Lead Site Reliability Engineer
Leads the architecture, reliability, observability, and operational excellence of secure cloud and hybrid infrastructure for a mission-critical collaboration platform. Requires 5+ years in SRE, DevOps, or cloud infrastructure, with expertise in Kubernetes, Terraform, AWS, and regulated environments.
About the job
Responsibilities
- Define the strategy, architecture, and roadmap for the site reliability engineering function, aligning infrastructure initiatives with product and business goals.
- Lead the design, deployment, and optimization of production-grade containerized workloads, infrastructure as code, and compliant cloud environments for regulated domains such as FedRAMP and DoD.
- Establish and evolve observability, monitoring, and alerting frameworks for performance, reliability, and capacity planning at scale.
- Drive incident management processes, including on-call rotations, root cause analysis, and systemic reliability improvements.
- Partner with security and compliance teams to meet data sovereignty, security, and regulatory requirements.
- Champion automation and operational excellence to improve efficiency, reduce risk, and scale operations.
- Oversee cloud cost management and capacity planning while meeting performance targets.
- Build and maintain a developer platform that enables fast, secure software delivery and improves application stability in production.
- Mentor and coach SRE team members.
Requirements
- Bachelor's degree in Computer Science, Cybersecurity, Software Engineering, or a related technical field, or equivalent experience.
- 5+ years of relevant experience in site reliability engineering, DevOps, or cloud infrastructure roles.
- Expertise in container orchestration platforms, ideally Kubernetes.
- Extensive infrastructure-as-code experience, ideally Terraform.
- Strong cloud platform background, ideally AWS.
- Experience designing and implementing monitoring, alerting, and performance optimization strategies.
- Strong troubleshooting and incident management skills for distributed systems.
- Proficiency in at least one scripting or programming language for automation.
- Excellent communication and cross-functional influence skills.
- Experience leading globally distributed teams in a remote-first environment.
Nice-to-Haves
- Familiarity with observability stacks such as Grafana and Prometheus.
- Experience designing high-availability, disaster recovery, and scaling architectures.
- Exposure to Google Cloud and Azure.
- Leadership experience in defense, finance, or critical infrastructure.
- Experience with U.S. federal compliance frameworks, including FedRAMP, DoD ATO, and NIST 800-53.
- Experience preparing and supporting software offerings through AWS Marketplace, Azure Marketplace, or Google Cloud Marketplace.
- Open-source contributions in reliability, DevOps, or infrastructure tooling.
- Cloud infrastructure, reliability, or DevOps certifications such as CKA, CKAD, or AWS Certified Solutions Architect.
Compensation
- Salary range: $145,000–$200,000
Skills
Kubernetes, Terraform, AWS, Grafana, Prometheus, GCP, Azure, Infrastructure As Code, Observability, Incident Management, Disaster Recovery, Nist 800-53
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.
The Senior Site Reliability Engineer will build and operate secure, scalable infrastructure and Snowflake data systems, automate deployments and operational processes, and lead incident response. The role requires strong coding, Terraform, Kubernetes, CI/CD, and data-platform experience, plus U.S. Person status.
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.