Senior Site Reliability Engineer
Build and operate secure, scalable AWS infrastructure and developer enablement systems for internal engineering teams. The role requires 5+ years in SRE, DevOps, or systems engineering, with strong expertise in AWS governance, Terraform, Python automation, CI/CD, Kubernetes, and observability.
About the job
Responsibilities
- Design, build, and modernize scalable cloud environments and development tools while enforcing security policies and standards for regulated environments.
- Partner with software engineering teams to promote DevOps and SRE best practices, provide internal customer service, and contribute to Agile workflows such as demos and architecture sessions.
- Create and maintain technical documentation, including network diagrams, runbooks, and disaster recovery procedures.
Requirements
- 5+ years of experience in SRE, DevOps, or Systems Engineering, with a track record of delivering complex, large-scale infrastructure projects.
- Expertise building and managing AWS multi-account environments, including AWS Organizations, IAM, Identity Center, and StackSets.
- Strong skills in infrastructure as code with Terraform, secure automation with Python, and Git-based CI/CD workflows using GitLab or GitHub Actions.
- Experience managing Kubernetes environments and using observability tools such as Splunk, CloudWatch, and Grafana.
- Ability to access federal environments or protected federal data and provide documentation establishing U.S. Person status upon hire.
Nice-to-haves
- Networking experience, including BGP, IPsec, VPCs, Transit Gateways, and VPC endpoints.
- Linux systems administration experience.
- Experience in highly secure, regulated environments such as FedRAMP, with knowledge of FIPS, STIGs, and data boundary implementations.
Compensation and Benefits
- Base salary range of $165,000–$225,600 for candidates in the San Francisco Bay Area.
- Base salary range of $147,000–$202,000 for candidates in California excluding the San Francisco Bay Area, Colorado, Illinois, New York, and Washington.
- Equity where applicable, bonus, health, dental and vision insurance, 401(k), flexible spending account, and paid leave including PTO and parental leave.
Skills
AWS, Aws Organizations, IAM, Terraform, Python, GitLab, GitHub Actions, Kubernetes, Splunk, CloudWatch, Grafana, BGP, Ipsec, Linux, FedRAMP
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.
The Senior Site Reliability Engineer will build and operate secure, scalable infrastructure and Snowflake data systems, automate deployments and operational processes, and lead incident response. The role requires strong coding, Terraform, Kubernetes, CI/CD, and data-platform experience, plus U.S. Person status.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.