Skip to content
ReltioReltio

Senior Engineer

Build and operate reliability and resilience capabilities for a multi-cloud platform, including chaos engineering, observability-driven validation, failover testing, and resilient distributed systems. The role requires 6–9 years of software engineering experience and hands-on expertise with Java, Kubernetes, cloud platforms, and CI/CD.

About the job

Responsibilities

Reliability & Resilience Engineering

  • Design and implement resilience strategies across AWS, Azure, and Google Cloud environments.
  • Support chaos engineering initiatives and integrate experiments into CI/CD pipelines.
  • Validate RTO, RPO, SLO, and recovery objectives using metrics.
  • Follow governance controls for chaos experiments.
  • Identify and mitigate potential failure scenarios before production.

Distributed Systems & Platform Engineering

  • Improve resilience across databases, caching, messaging, and authentication services.
  • Support multi-region and cross-cloud failover validation.
  • Validate retry logic, circuit breakers, and graceful degradation.
  • Reduce cascading failures and improve system stability.

Observability & Continuous Validation

  • Use Grafana, Prometheus, and LogDNA for validation.
  • Integrate reliability checks into Jenkins CI/CD pipelines.
  • Ensure production readiness before releases.
  • Use telemetry to identify improvement areas.

Collaboration & Engineering Excellence

  • Work with SRE, platform, and development teams to improve resilience.
  • Contribute to best practices and documentation.
  • Participate in incident analysis and improvements.

Requirements

  • 6–9 years of software engineering experience.
  • Experience with distributed systems and microservices.
  • Exposure to AWS, Azure, and Google Cloud services.
  • Hands-on Kubernetes experience, including EKS, AKS, or GKE.
  • Strong Java and Spring Boot skills.
  • Understanding of distributed-system failures and resilience patterns.
  • Exposure to chaos tools such as Chaos Mesh, Gremlin, or Harness.
  • CI/CD experience; Jenkins preferred.
  • Understanding of SLOs and SLIs.
  • Experience with Grafana, Prometheus, and LogDNA.

Nice to Have

  • Exposure to multi-region architecture.
  • Familiarity with service meshes.
  • Performance testing experience.
  • Familiarity with security tools such as Snyk.
  • Good communication skills.

Success Measures

  • Reliability checks are integrated into CI/CD.
  • Chaos testing improves resilience.
  • Failures are detected early.
  • System stability improves.
  • The team contributes to a resilience-first culture.

Skills

Java, Spring Boot, Distributed Systems, Microservices, AWS, Azure, GCP, Kubernetes, Chaos Mesh, Gremlin, Harness, Jenkins, Grafana, Prometheus, Logdna

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.