Staff Site Reliability Engineer
Leads infrastructure evolution, CI/CD, observability, and reliability for scaling platform. Requires 10+ years SRE/infra experience, AWS expertise, and systems thinking. Onsite in NYC.
About the job
What You’ll Own
- AWS infrastructure direction and platform evolution, including migration from ECS/Fargate toward modern, scalable runtime
- CI/CD systems emphasizing developer experience, safety, and automation (GitHub Actions; maturing CD)
- Ephemeral environments and preview deploys
- Observability standards (metrics, logs, tracing, alert hygiene, dashboards, SLO development)
- Incident response, postmortems, and reliability culture
What You’ll Do
- Define and evolve reliability standards, SLIs, SLOs, and error budgets
- Improve observability, alerting, and incident processes
- Lead high-severity incidents and drive follow-ups
- Partner with teams to design resilient, scalable systems
- Build automation to reduce toil and risk
- Mentor engineers and influence best practices
Who You Are
- Run production systems on AWS and lead platform change
- Think in systems: risk, rollback, blast radius, feedback loops
- Treat CI/CD and environments as self-serve products
- Influence through trust and clarity
- Balance pragmatism with system health
- Value learning from failure
- Communicate clearly across teams
Experience
- 10+ years in SRE, infrastructure, or backend engineering
- Strong software engineering in modern languages
- Expertise in distributed systems at scale
- Deep AWS, observability tooling, CI/CD experience
- Comfortable with ambiguity
Perks and Benefits
- Competitive compensation and equity
- Unlimited PTO
- Up to 100% employer-covered healthcare
- Meals provided
- Parental leave, commuter benefits, 401k
Skills
AWS, Kubernetes, GitHub Actions, CI/CD, Slo, Sli, Observability, Distributed Systems, Incident Response, Terraform
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.