Staff Site Reliability Engineer
Leads infrastructure transformation from monoliths to scalable microservices at massive scale, architects observability/CI/CD systems, unifies complex stacks, and mentors engineers. Requires 10+ years coding internal tools, 5+ years cloud (GCP/AWS), Bachelor's in CS.
About the job
Responsibilities
- Execute on the transformation from monolith to scalable microservices (API/Platform focus).
- Drive initiatives to continually improve reliability, with a deep understanding of the implications of each “9.”
- Architect systems and write code that enables application teams to adopt best practices by default—not by instruction.
- Integrate and unify diverse infrastructure components into a cohesive, scalable platform within a massive tech stack.
- Design observability, reliability, and CI/CD frameworks to support growth and operational excellence at scale.
- Collaborate cross-functionally with product, application, and integration teams to align infrastructure direction with business goals.
- Provide technical leadership to shift the team from reactive support to a proactive, strategic function.
- Mentor and guide a team of 6 engineers while shaping the direction of infrastructure engineering.
Minimum Qualifications
- Bachelor's degree in Computer Science or related field of study.
- At least 10 years of hands-on coding experience in building internal platforms/tools to support developer experience and operational best practices.
- At least 5 years of experience in cloud platforms—GCP preferred, AWS acceptable; cloud engineering background required.
Preferred Qualifications
- Proven experience scaling infrastructure in environments with many thousands of nodes.
- Track record of leading architectural shifts from monolithic systems to microservices in large-scale environments.
- Deep knowledge of reliability engineering and high-availability systems; able to articulate the impact of increasing the number of 9s.
- Strong understanding of first-party infrastructure integration and unifying disparate systems.
- Familiarity with observability, CI/CD tooling, and infrastructure automation.
- Experience at large-scale tech companies (Google, Meta, Amazon, etc.) or equivalent environments highly preferred.
- Strong cross-functional collaboration skills and the ability to drive infrastructure alignment across engineering orgs.
Skills
GCP, AWS, Kubernetes, CI/CD, Observability, Microservices, Reliability Engineering, Infrastructure Automation, Cloud Engineering, Platform Engineering
Similar jobs
DevOps / SRE jobsLeads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.