Staff Software Engineer, Site Reliability Engineer (SRE)
Staff SRE ensures reliability, scalability, and performance of legal AI platform across global regions. Leads incident management, automates operations, optimizes infrastructure costs, and mentors teams. Requires 10+ years SRE experience, IaC, cloud platforms, and observability tools.
About the job
What You’ll Do
- Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
- Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
- Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
- Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle
- Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality
- Provide technical mentorship and leadership, promoting best practices and fostering team growth
What You Have
- 10+ years of experience in Site Reliability Engineering or similar roles supporting production environments, with proven ability to mentor and guide technical teams
- Expertise in infrastructure as code (IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
- Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
- Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
- Strong programming skills (Python, Bash, Go, or similar languages)
- Proven track record of diagnosing complex system problems and implementing durable solutions
- Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
- Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence
Compensation
$238,000 - $290,000 USD
Skills
Terraform, Pulumi, Kubernetes, Datadog, Pagerduty, Python, Go, AWS, GCP, Azure, CI/CD
Similar jobs
DevOps / SRE jobsOwns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.