Senior Site Reliability Engineer
Senior Site Reliability Engineer responsible for designing, implementing, and operating observability systems for complex cloud platforms. Focus on improving reliability, resilience, reducing toil through automation, and participating in on-call duties. Requires strong experience with IaC, cloud technologies, observability tools, and Linux engineering.
About the job
Responsibilities
- Engage with teams and improve service delivery and reliability across their entire lifecycle.
- Measure and monitor all production systems with an eye towards availability, latency and overall system health.
- Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence.
- Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability.
- Help identify and drive down toil with creative innovation and automation.
- This position will require stand-by, on-call, or off-hours duties.
Requirements
- Proven experience designing, implementing, and operating observability systems for complex cloud-based platforms, with deep knowledge of best practices.
- Experience with Configuration Management and Infrastructure as Code Tools like Terraform (preferred) or Ansible.
- Experience working with Cloud SDKs.
- Knowledge of cloud platforms (AWS and Azure preferred) and container + orchestration technologies.
- Experience with APM and Observability tools such as New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry.
- Extensive experience with enterprise scale continuous delivery environments.
- Development experience with JavaScript/Node.js/TypeScript in a Linux/Mac environment.
- Experience with sustainable incident response in a blameless environment.
- Background in Linux Systems Engineering.
- Experience with Incident response tools such as PagerDuty, FireHydrant, Blameless.
- Comfortable with a high level of autonomy and working with a distributed team.
- Knowledge of Cloud and application security best practices.
- Strong knowledge of cloud design patterns for scale, data management, resiliency, etc.
- A love for high quality and a knack for testing.
- Opinions about business metrics, and SLOs.
Benefits
- Base Salary Range: $141,800—$195,000 USD.
- In addition to a competitive salary, Cribl also offers a generous benefits package which includes health, dental, vision, short-term disability, and life insurance, paid holidays and paid time off, a fertility treatment benefit, 401(k), and equity.
- Employees are eligible to participate in the Cribl Corporate Bonus Program.
Skills
Terraform, Ansible, AWS, Azure, Kubernetes, Docker, Prometheus, Grafana, New Relic, Splunk, Pagerduty, Node.js, TypeScript, Linux
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.