Skip to content
Order.coOrder.co

Senior Site Reliability Engineer

Senior SRE responsible for building and operating reliable, scalable infrastructure on AWS with Kubernetes and Terraform. Focus on observability, incident response, automation, and mentoring engineers on SRE best practices.

About the job

Responsibilities

Reliability Engineering & Infrastructure Ownership

  • Design, build, and operate highly available, scalable, and fault-tolerant infrastructure and platform services
  • Own reliability, availability, latency, and operational excellence for critical production systems and services
  • Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets across platform systems
  • Lead incident response efforts for complex production outages; drive root-cause analysis and long-term remediation actions
  • Build resilient systems that gracefully handle failures, traffic spikes, dependency degradation, and regional outages
  • Continuously improve system reliability through automation, observability, performance tuning, and capacity planning

Automation & Platform Engineering

  • Develop infrastructure automation and self-service tooling to reduce operational toil and improve engineering velocity
  • Build and maintain CI/CD pipelines, deployment automation, and release engineering workflows
  • Implement infrastructure as code (IaC) practices using tools such as Terraform, CloudFormation, and container orchestration
  • Improve developer experience by building reliable internal platforms, operational tooling, and standardized deployment patterns
  • Drive adoption of GitOps, immutable infrastructure, and automated remediation patterns

Observability & Operational Excellence

  • Design and maintain comprehensive monitoring, logging, tracing, and alerting systems for distributed services
  • Establish actionable alerting standards that reduce noise while improving incident detection and response times
  • Analyze production trends, system bottlenecks, and failure patterns to proactively prevent incidents
  • Lead operational readiness reviews, disaster recovery planning, and game-day exercises
  • Improve mean time to detect (MTTD) and mean time to recovery (MTTR) through tooling, automation, and process refinement

Systems Architecture & Scalability

  • Participate actively in architecture and infrastructure design reviews
  • Propose scalable and reliable platform designs that account for multi-region deployment, redundancy, failover, and security considerations
  • Evaluate trade-offs between reliability, scalability, operational complexity, and engineering velocity
  • Identify systemic risks and operational gaps before they become production incidents
  • Partner with engineering teams to ensure services are designed with operability, observability, and resilience in mind from day one

Security & Compliance

  • Approach infrastructure and operational practices with a strong security mindset
  • Implement and maintain secure cloud networking, secrets management, IAM policies, and infrastructure hardening standards
  • Partner with Security and Compliance teams to ensure systems meet organizational and regulatory requirements
  • Drive operational best practices around vulnerability management, patching, and production access controls

End-to-End Ownership & Collaboration

  • Scope and estimate infrastructure and reliability initiatives accurately
  • Coordinate production rollouts, maintenance events, and reliability improvements across teams
  • Communicate operational risks, dependencies, and incident impacts clearly to technical and non-technical stakeholders
  • Collaborate closely with Software Engineering, Security, Product, and Operations teams to improve platform reliability and scalability
  • Serve as a trusted escalation point during critical production incidents

Mentorship & Technical Leadership

  • Mentor junior and mid-level engineers on reliability engineering principles, operational excellence, and infrastructure best practices
  • Raise the operational maturity of the engineering organization through documentation, reviews, and technical guidance
  • Drive improvements in team standards around observability, incident management, automation, and infrastructure design
  • Influence technical decisions through credibility, operational expertise, and strong engineering judgment

Requirements

  • Strong foundation in computer science fundamentals: data structures, algorithms, and system design
  • Familiarity with building production-grade applications and services using Ruby and Ruby on Rails
  • Deep expertise with Linux systems administration and production troubleshooting
  • Strong experience operating cloud infrastructure at scale, particularly within AWS environments
  • Experience with Kubernetes, container orchestration, and cloud-native infrastructure patterns
  • Proficiency with infrastructure as code tools such as Terraform or CloudFormation
  • Expertise designing and operating CI/CD pipelines and deployment automation systems
  • Deep understanding of observability tooling including Datadog, OpenTelemetry, or similar platforms
  • Strong knowledge of distributed systems reliability patterns including redundancy, failover, autoscaling, rate limiting, and graceful degradation
  • Experience building automation and operational tooling using languages such as Python, Go, Bash, or Ruby
  • Strong understanding of networking fundamentals including DNS, load balancing, TLS, VPNs, firewalls, and service discovery
  • Hands-on experience with incident response, root-cause analysis, and production operations in high-availability environments
  • Familiarity with SRE methodologies including SLOs, SLIs, error budgets, capacity planning, and operational maturity modeling

Skills

Ruby, Ruby on Rails, Linux, AWS, Kubernetes, Terraform, CloudFormation, CI/CD, Datadog, OpenTelemetry, Python, Go, Bash, SRE, SLOs

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Coinbase

Coinbase

United States

Senior Software Engineer, Core Infra Systems
$186k+/yrRemote5+ YOEDevOps / SRE

Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.

Shield AI

Shield AI

Seattle, WA
Senior Site Infrastructure Engineer
$110k+/yrOn-site5+ YOEDevOps / SRE

Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.