Staff Engineer – Observability Platform
Leads the technical vision and development of a scalable observability platform, improving reliability, performance, incident response, and engineering productivity across Postman. Requires 10+ years of software engineering experience with distributed systems, cloud-native architectures, and production operations.
About the job
Responsibilities
- Drive the technical vision and architecture for Postman's observability platform.
- Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics.
- Investigate complex production issues, identify root causes, and drive long-term corrective actions.
- Partner with engineering teams to improve service reliability, availability, performance, and operational maturity.
- Establish observability standards, best practices, and instrumentation frameworks across the company.
- Build tooling and automation that enable faster incident detection, diagnosis, and resolution.
- Leverage telemetry data to uncover performance bottlenecks, capacity risks, and reliability gaps.
- Drive cross-functional initiatives focused on platform health, operational excellence, and engineering productivity.
- Mentor senior engineers and help raise the technical bar across the organization.
Requirements
- 10+ years of software engineering experience with significant exposure to distributed systems and cloud-native architectures.
- Strong expertise in observability domains, including monitoring, logging, distributed tracing, telemetry pipelines, and incident management.
- Experience operating large-scale production systems with a strong focus on reliability, scalability, and performance.
- Deep understanding of system debugging, root-cause analysis, performance optimization, and production operations.
- Strong programming experience in one or more languages such as Go, Java, Python, or Node.js.
- Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, or Honeycomb.
- Ability to influence technical direction across teams without direct authority.
- Strong communication skills and a data-driven approach to problem-solving.
Preferred Qualifications
- Experience building internal developer platforms or observability platforms at scale.
- Exposure to AIOps, intelligent alerting, anomaly detection, or AI-powered operational tooling.
- Experience driving reliability initiatives across multiple engineering organizations.
- Background in SRE, Platform Engineering, Infrastructure Engineering, or Developer Productivity.
Compensation and Benefits
- Pay-on-performance philosophy.
- Flexible schedule.
- Full medical coverage.
- Flexible PTO.
- Wellness reimbursement.
- Monthly lunch stipend.
- Wellness programs.
- Team-building events.
- Donation-matching program.
- Bangalore-based employees currently work in the office three days a week and are expected to transition to five days per week by the end of the year.
Skills
Distributed Systems, Cloud-Native Architecture, Observability, Monitoring, Logging, Distributed Tracing, Telemetry Pipelines, Incident Management, Go, Java, Python, Node.js, OpenTelemetry, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.