Staff Software Engineer, Reliability
Leads reliability engineering for a large-scale platform, designing multi-region failover, observability, disaster recovery, and resilient distributed systems. Requires 8+ years of engineering experience, expert Java or Scala proficiency, AWS operations expertise, and strong incident-management leadership.
About the job
Responsibilities
- Own the overall reliability posture for the Metropolis platform, establishing practices, metrics, and systems that ensure 99.9%+ uptime across all services.
- Design and implement automatic failover mechanisms for critical external dependencies such as Twilio for SMS/voice and Stripe for payments, including circuit breakers, retry policies, and degraded-mode operations.
- Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, DNS-based traffic routing, and disaster recovery planning and testing.
- Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation.
- Implement synthetic monitoring, SLO-based alerting, on-call rotations, escalation policies, and service-health dashboards showing customer impact.
- Own incident management workflows, tooling, post-mortem practices, runbook automation, and MTTR reduction initiatives.
- Drive adoption of resilience patterns across services, including health checks, graceful degradation, feature flags, rate limiting, backpressure mechanisms, and chaos engineering.
- Build and maintain local mirrors for critical dependencies with artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages.
Requirements
- 8+ years of engineering experience, including software engineering, reliability engineering, SRE practices, or production operations at scale.
- Expert-level reliability engineering experience with multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery.
- Deep production observability experience implementing monitoring, alerting, tracing, and logging systems at scale, specifically Datadog or similar APM platforms in high-load environments.
- Strong systems thinking and experience designing resilient distributed systems that handle failures, network partitions, and external dependency outages.
- Database and data-systems knowledge, including replication, backup and restore, connection pooling, query optimization, and relational and NoSQL databases.
- Production AWS experience, including multi-region deployments, load balancing, and DNS-based failover.
- Experience with AI-powered development tools such as Claude Code, GitHub Copilot, or similar agentic coding tools, particularly context engineering.
- Excellent technical communication skills, including influencing technical decisions, documenting complex systems, conducting post-mortems, and establishing organization-wide reliability standards.
- Expert-level Java and/or Scala proficiency, with strong knowledge of JVM performance, concurrency, and operational characteristics.
Nice to Have
- Scala experience.
- SRE or reliability engineering experience at companies known for operational excellence or high-growth startups where reliability practices were built from the ground up.
- Incident-response leadership, including incident management processes, blameless post-mortems, and MTTR reduction initiatives.
- Chaos engineering experience with tools such as Chaos Monkey, Gremlin, or similar, including game days and failure-injection testing.
- Performance optimization experience with profiling, benchmarking, capacity planning, and system tuning for high-throughput, low-latency systems.
- Open-source contributions or technical blog writing demonstrating expertise in reliability engineering, distributed systems, or production operations.
Technology Stack
- Languages and frameworks: TypeScript, React, Scala, Java
- Datastores: MySQL, PostgreSQL, Snowflake
- Cloud: AWS
- Version control: Git, GitHub
- AI tooling: GitHub Copilot
- Observability: Datadog
Work Arrangement
- Office-first model requiring employees to be on-site at least four days per week.
Skills
AWS, Datadog, Scala, Java, TypeScript, React, MySQL, Postgres, Snowflake, Git, GitHub, Github Copilot, Chaos Engineering, Distributed Systems, Dns Failover
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.