# Staff Software Engineer, Reliability

**Company:** [Metropolis](https://hotfix.jobs/companies/metropolis)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** AWS, Datadog, Scala, Java, TypeScript, React, MySQL, Postgres, Snowflake, Git, GitHub, Github Copilot, Chaos Engineering, Distributed Systems, Dns Failover
**Posted:** 2026-07-31

> Leads reliability engineering for a large-scale platform, designing multi-region failover, observability, disaster recovery, and resilient distributed systems. Requires 8+ years of engineering experience, expert Java or Scala proficiency, AWS operations expertise, and strong incident-management leadership.

## Job Description

## Responsibilities
- Own the overall reliability posture for the Metropolis platform, establishing practices, metrics, and systems that ensure 99.9%+ uptime across all services.
- Design and implement automatic failover mechanisms for critical external dependencies such as Twilio for SMS/voice and Stripe for payments, including circuit breakers, retry policies, and degraded-mode operations.
- Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, DNS-based traffic routing, and disaster recovery planning and testing.
- Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation.
- Implement synthetic monitoring, SLO-based alerting, on-call rotations, escalation policies, and service-health dashboards showing customer impact.
- Own incident management workflows, tooling, post-mortem practices, runbook automation, and MTTR reduction initiatives.
- Drive adoption of resilience patterns across services, including health checks, graceful degradation, feature flags, rate limiting, backpressure mechanisms, and chaos engineering.
- Build and maintain local mirrors for critical dependencies with artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages.

## Requirements
- 8+ years of engineering experience, including software engineering, reliability engineering, SRE practices, or production operations at scale.
- Expert-level reliability engineering experience with multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery.
- Deep production observability experience implementing monitoring, alerting, tracing, and logging systems at scale, specifically Datadog or similar APM platforms in high-load environments.
- Strong systems thinking and experience designing resilient distributed systems that handle failures, network partitions, and external dependency outages.
- Database and data-systems knowledge, including replication, backup and restore, connection pooling, query optimization, and relational and NoSQL databases.
- Production AWS experience, including multi-region deployments, load balancing, and DNS-based failover.
- Experience with AI-powered development tools such as Claude Code, GitHub Copilot, or similar agentic coding tools, particularly context engineering.
- Excellent technical communication skills, including influencing technical decisions, documenting complex systems, conducting post-mortems, and establishing organization-wide reliability standards.
- Expert-level Java and/or Scala proficiency, with strong knowledge of JVM performance, concurrency, and operational characteristics.

## Nice to Have
- Scala experience.
- SRE or reliability engineering experience at companies known for operational excellence or high-growth startups where reliability practices were built from the ground up.
- Incident-response leadership, including incident management processes, blameless post-mortems, and MTTR reduction initiatives.
- Chaos engineering experience with tools such as Chaos Monkey, Gremlin, or similar, including game days and failure-injection testing.
- Performance optimization experience with profiling, benchmarking, capacity planning, and system tuning for high-throughput, low-latency systems.
- Open-source contributions or technical blog writing demonstrating expertise in reliability engineering, distributed systems, or production operations.

## Technology Stack
- **Languages and frameworks:** TypeScript, React, Scala, Java
- **Datastores:** MySQL, PostgreSQL, Snowflake
- **Cloud:** AWS
- **Version control:** Git, GitHub
- **AI tooling:** GitHub Copilot
- **Observability:** Datadog

## Work Arrangement
- Office-first model requiring employees to be on-site at least four days per week.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc) - Together AI - Remote
- [Staff DevSecOps Engineer, Enterprise Technology](https://hotfix.jobs/jobs/85d2a18a-f6f6-4858-a074-0d6b16869dd1) - Okta - Bengaluru, India
- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/cf3c46c0-63dd-4681-b117-16200b400d89) - Together AI - London, United Kingdom
- [Staff Software Engineer](https://hotfix.jobs/jobs/17a92e38-a4f6-4c6b-a6cc-e3bf61bd9955) - Okta - Bengaluru, India
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/21252ce5-25cd-4de1-aa95-572af4c7e77e
**Canonical:** https://hotfix.jobs/jobs/21252ce5-25cd-4de1-aa95-572af4c7e77e