# Senior Engineer

**Company:** [Reltio](https://hotfix.jobs/companies/reltio)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 6+ years
**Skills:** Java, Spring Boot, Distributed Systems, Microservices, AWS, Azure, GCP, Kubernetes, Chaos Mesh, Gremlin, Harness, Jenkins, Grafana, Prometheus, Logdna
**Posted:** 2026-08-19

> Build and operate reliability and resilience capabilities for a multi-cloud platform, including chaos engineering, observability-driven validation, failover testing, and resilient distributed systems. The role requires 6–9 years of software engineering experience and hands-on expertise with Java, Kubernetes, cloud platforms, and CI/CD.

## Job Description

## Responsibilities

### Reliability & Resilience Engineering
- Design and implement resilience strategies across AWS, Azure, and Google Cloud environments.
- Support chaos engineering initiatives and integrate experiments into CI/CD pipelines.
- Validate RTO, RPO, SLO, and recovery objectives using metrics.
- Follow governance controls for chaos experiments.
- Identify and mitigate potential failure scenarios before production.

### Distributed Systems & Platform Engineering
- Improve resilience across databases, caching, messaging, and authentication services.
- Support multi-region and cross-cloud failover validation.
- Validate retry logic, circuit breakers, and graceful degradation.
- Reduce cascading failures and improve system stability.

### Observability & Continuous Validation
- Use Grafana, Prometheus, and LogDNA for validation.
- Integrate reliability checks into Jenkins CI/CD pipelines.
- Ensure production readiness before releases.
- Use telemetry to identify improvement areas.

### Collaboration & Engineering Excellence
- Work with SRE, platform, and development teams to improve resilience.
- Contribute to best practices and documentation.
- Participate in incident analysis and improvements.

## Requirements
- 6–9 years of software engineering experience.
- Experience with distributed systems and microservices.
- Exposure to AWS, Azure, and Google Cloud services.
- Hands-on Kubernetes experience, including EKS, AKS, or GKE.
- Strong Java and Spring Boot skills.
- Understanding of distributed-system failures and resilience patterns.
- Exposure to chaos tools such as Chaos Mesh, Gremlin, or Harness.
- CI/CD experience; Jenkins preferred.
- Understanding of SLOs and SLIs.
- Experience with Grafana, Prometheus, and LogDNA.

## Nice to Have
- Exposure to multi-region architecture.
- Familiarity with service meshes.
- Performance testing experience.
- Familiarity with security tools such as Snyk.
- Good communication skills.

## Success Measures
- Reliability checks are integrated into CI/CD.
- Chaos testing improves resilience.
- Failures are detected early.
- System stability improves.
- The team contributes to a resilience-first culture.

## Similar jobs

- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/0496c897-3cf6-419f-bea7-91e629b4307c) - Okta - Bengaluru, India
- [Senior Release Engineer](https://hotfix.jobs/jobs/a187d17c-c638-43f1-8376-209fe54f2503) - GitLab - Remote
- [Senior Site Reliability Engineer - Monitoring and Anomaly Detection](https://hotfix.jobs/jobs/4d131f28-278c-4b63-a411-189242898dcc) - GitLab - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/a90d1d14-3f9d-40d5-ae01-365e6600fcab) - ZoomInfo - Bengaluru, India

**Apply:** https://hotfix.jobs/jobs/6417eb9b-44bf-4990-9744-729aad2e9f20
**Canonical:** https://hotfix.jobs/jobs/6417eb9b-44bf-4990-9744-729aad2e9f20