Skip to content
PostmanPostman

Staff Engineer – Observability Platform

Leads the technical vision and development of a scalable observability platform, improving reliability, performance, incident response, and engineering productivity across Postman. Requires 10+ years of software engineering experience with distributed systems, cloud-native architectures, and production operations.

About the job

Responsibilities

  • Drive the technical vision and architecture for Postman's observability platform.
  • Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics.
  • Investigate complex production issues, identify root causes, and drive long-term corrective actions.
  • Partner with engineering teams to improve service reliability, availability, performance, and operational maturity.
  • Establish observability standards, best practices, and instrumentation frameworks across the company.
  • Build tooling and automation that enable faster incident detection, diagnosis, and resolution.
  • Leverage telemetry data to uncover performance bottlenecks, capacity risks, and reliability gaps.
  • Drive cross-functional initiatives focused on platform health, operational excellence, and engineering productivity.
  • Mentor senior engineers and help raise the technical bar across the organization.

Requirements

  • 10+ years of software engineering experience with significant exposure to distributed systems and cloud-native architectures.
  • Strong expertise in observability domains, including monitoring, logging, distributed tracing, telemetry pipelines, and incident management.
  • Experience operating large-scale production systems with a strong focus on reliability, scalability, and performance.
  • Deep understanding of system debugging, root-cause analysis, performance optimization, and production operations.
  • Strong programming experience in one or more languages such as Go, Java, Python, or Node.js.
  • Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, or Honeycomb.
  • Ability to influence technical direction across teams without direct authority.
  • Strong communication skills and a data-driven approach to problem-solving.

Preferred Qualifications

  • Experience building internal developer platforms or observability platforms at scale.
  • Exposure to AIOps, intelligent alerting, anomaly detection, or AI-powered operational tooling.
  • Experience driving reliability initiatives across multiple engineering organizations.
  • Background in SRE, Platform Engineering, Infrastructure Engineering, or Developer Productivity.

Compensation and Benefits

  • Pay-on-performance philosophy.
  • Flexible schedule.
  • Full medical coverage.
  • Flexible PTO.
  • Wellness reimbursement.
  • Monthly lunch stipend.
  • Wellness programs.
  • Team-building events.
  • Donation-matching program.
  • Bangalore-based employees currently work in the office three days a week and are expected to transition to five days per week by the end of the year.

Skills

Distributed Systems, Cloud-Native Architecture, Observability, Monitoring, Logging, Distributed Tracing, Telemetry Pipelines, Incident Management, Go, Java, Python, Node.js, OpenTelemetry, Prometheus, Grafana

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Okta

Okta

Bengaluru, India

Staff DevSecOps Engineer, Enterprise Technology
No salary listedOn-site8+ YOEDevOps / SRE

Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Okta

Okta

Bengaluru, India

Staff Software Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.