Skip to content

Senior Software Engineer - Observability

Develops scalable observability tooling and infrastructure for large-scale distributed systems, including logging, metrics, tracing, dashboards, and alerting. Requires 7+ years of production software experience and a bachelor’s degree or higher in computer science or a related field.

About the job

Responsibilities

  • Establish standards for logging, metrics, and tracing.
  • Collaborate with teams to identify metrics for observing system and component performance.
  • Build tooling and infrastructure for components to efficiently emit, aggregate, and store metrics for dashboards and alerting.
  • Contribute to and execute the technical roadmap for scalable, performant, and reliable systems.
  • Participate in on-call rotations and reduce incident response times.
  • Optimize platform and infrastructure costs by improving visibility, enforcing retention policies, streamlining queries, and right-sizing resources.

Requirements

  • Bachelor's degree or higher in Computer Science or a related field.
  • 7+ years of production-level experience with Python, Java, Scala, C++, or similar languages.
  • Experience developing software for large-scale distributed systems.
  • Familiarity with metrics collection, health monitoring, and observability tools.

Skills

Python, Java, Scala, C++, Distributed Systems, Metrics Collection, Health Monitoring, Observability, Logging, Tracing, Alerting

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

GitLab

GitLab

Bengaluru, India

Senior Site Reliability Engineer - Monitoring and Anomaly Detection
No salary listedRemote5+ YOEDevOps / SRE

Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

ZoomInfo

ZoomInfo

Bengaluru, India
Senior DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.