Staff Site Reliability Engineer, Security - GCP
Hardens and operates large-scale GCP and AWS infrastructure, automating security remediation, incident response, IAM, and reliability improvements. The role requires deep cloud security and DevSecOps experience, strong infrastructure-as-code skills, and expertise in Kubernetes, Linux, and security automation.
About the job
Responsibilities
- Lead initiatives to strengthen security posture for critical infrastructure and promote best practices across engineering.
- Respond to production security incidents, perform root cause analysis, and build automated preventions for performance and reliability.
- Identify manual security processes and automate them with custom tooling and CI/CD integrations.
- Develop technical documentation, runbooks, and procedures for a 24x7 online environment.
- Evolve monitoring platforms from simple auditing to active, automated prevention.
- Design and maintain large-scale production IAM policies and secrets-management workflows.
- Implement and maintain Public Key Infrastructure (PKI), ensuring GCE and GKE environments meet strict compliance standards.
- Monitor system health and security telemetry using industry-standard tools.
- Lead phased transitions of security policies from audit/detection mode to blocking/prevention mode without affecting production uptime.
Requirements
- 8+ years of experience architecting and running complex cloud networking and infrastructure, including 7+ years specializing in DevSecOps or cloud security.
- At least 3+ years of deep, hands-on experience securing GCP, including GKE, GCE, and Shared VPC.
- 10+ years of experience using Terraform and Chef to manage cloud resources and OS hardening.
- Expert proficiency in Go, Python, or Ruby for building security tooling and automated remediation.
- Experience securing containerized workloads, including image scanning, Kubernetes RBAC, and runtime security tools.
- Ability to troubleshoot complex networking, IAM, and performance issues under pressure.
- Strong knowledge of Linux internals, OS hardening, CIS benchmarks, and IP protocols including TLS/SSL, DNSSEC, and BGP.
- Bachelor's degree in Computer Science or equivalent professional experience.
Nice to Have
- Experience designing unified IAM governance across AWS and GCP, including Workload Identity Federation, Workforce Identity Federation, SAML, OIDC, and automated least-privilege enforcement.
- Deep understanding of multi-cloud reliability patterns and maintaining high availability during security patching or infrastructure hardening.
- Advanced experience securing GKE, EKS, and kOps, including Pod Security Standards, network policies, and admission controllers for zero-trust environments.
- Experience with security reviews and threat modeling at the design and implementation levels.
Tools and Technologies
- OSQuery
- Splunk
- Chronicle
- Nessus
- Qualys
- CrowdStrike Falcon
- Falco
- gVisor
Skills
GCP, AWS, Terraform, Chef, Go, Python, Ruby, Kubernetes, GKE, Gce, IAM, Linux, Pki, Splunk, Crowdstrike Falcon
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.