Senior / Staff Site Reliability Engineer
Senior/Staff SRE maintains high-availability production infrastructure on AWS EKS with Terraform, manages sharded MongoDB clusters, participates in on-call rotation, and engages with customers to enhance reliability for 1B+ daily API calls.
About the job
Responsibilities
- Work on core Radar infrastructure using Terraform all deployed to AWS via EKS
- Work on critical company-level initiatives, for example, 99.99+% availability, multi-region deployment, and cloud cost-saving activities
- Have your work be used by 100's of millions of devices
- Be part of the on-call rotation (one week rotation every 1-2 months)
- Talk to Radar customers and prospects, hear their feedback, incorporate it into your work and make them successful
Requirements
- Have experience managing a production AWS environment via Terraform
- Have experience with high availability multi-region production infrastructure running on Kubernetes
- Have experience managing large sharded Mongo clusters with heavy read-write workloads
- Have experience at a high growth startup
- Be interested in talking to customers or prospects and making them successful
Nice-to-haves
- Are a former technical co-founder
- Have experience with high throughput data intensive applications
Compensation
- Base salary range: $200,000 - $300,000/year
- Opportunity for performance bonuses and incentives
- Competitive equity plan with stock option grants
- 401(k) plan with 4% match
- Health, dental, and vision insurance with 100% coverage for employees
- 12 weeks of paid parental leave
- Commuter and fitness benefits
Skills
Terraform, AWS, EKS, Kubernetes, MongoDB, CircleCI, CloudWatch, Grafana, Pagerduty, Cloudflare
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.