Skip to content
PinterestPinterest

Staff Software Engineer, Observability

Staff Software Engineer building and scaling Pinterest's observability platform (metrics, logs, traces) for massive distributed systems. Requires 7+ years distributed systems experience, strong data engineering skills, and expertise with modern observability tools.

About the job

What you'll do

  • Define and execute the observability roadmap, treating it as a product. Understand engineering team needs and translate them into technical solutions with measurable impact
  • Architect, build, and scale distributed observability infrastructure (metrics, logs, traces) to handle massive volumes across Pinterest's distributed systems
  • Build high-performance data pipelines and storage for real-time and historical telemetry analysis at Pinterest scale
  • Champion Best Practices: Establish observability standards and patterns across the organization, making it easy for teams to instrument their services and gain actionable insights
  • Technical Leadership: Mentor engineers, lead architectural reviews, and influence technical decisions across teams to improve overall system reliability and performance
  • Cross-functional Collaboration: Partner with SRE, Infrastructure, Product Engineering, and other teams to understand pain points and deliver solutions that improve developer productivity and system reliability
  • Innovation: Stay current with observability trends and technologies, evaluating and adopting cutting-edge tools and techniques to keep Pinterest at the forefront

What we're looking for

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience
  • Product Mindset: Demonstrated ability to work backwards from customer needs—understanding user needs, prioritizing features, measuring success, and iterating based on feedback. Experience building internal platforms or tools with strong adoption
  • Distributed Systems Expertise: 7+ years of experience designing and operating large-scale distributed systems with deep understanding of consistency, availability, scalability, and failure modes
  • Data Engineering Skills: Strong background in building data pipelines, working with time-series databases, columnar storage, stream processing (Kafka, Flink, etc.), and data modeling at scale
  • Observability Domain Knowledge: Hands-on experience with modern observability tools and practices including metrics, logging, tracing, and profiling. Familiarity with OpenTelemetry, Prometheus, Grafana, or similar technologies
  • Programming Proficiency: Expert-level coding skills in languages like Java, Python, Go, or Scala with ability to write production-quality code
  • Systems Thinking: Ability to see the big picture while managing complex technical details, balancing trade-offs between cost, performance, and reliability
  • Experience building observability platforms from the ground up or significantly scaling existing solutions
  • Familiarity with cloud-native architectures and technologies (Kubernetes, service mesh, etc.)
  • Track record of driving adoption of internal platforms through excellent documentation, UX, and developer advocacy
  • Experience with machine learning or anomaly detection applied to observability use cases
  • Strong communication skills with ability to influence stakeholders at all levels
  • Contributions to open-source observability projects, a plus

Skills

Java, Python, Go, Scala, OpenTelemetry, Prometheus, Grafana, Kafka, Flink, Kubernetes, Time-Series Databases, Distributed Systems, Data Pipelines, Stream Processing, Observability

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Attentive

Attentive

United States

Staff Site Reliability Engineer
$180k+/yrRemote7+ YOEDevOps / SRE

Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.

Shield AI

Shield AI

United States

Sr. Staff Platform/Data Reliability Engineer, Databricks
$180k+/yrRemote12+ YOEDevOps / SRE

Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.

Okta

Okta

Bellevue, WA
Staff Site Reliability Engineer - Kubernetes
$174k+/yrHybrid7+ YOEDevOps / SRE

Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.

Okta

Okta

Maryland
Staff Site Reliability Engineer, Kubernetes w/ active TS/SCI
$174k+/yrHybrid8+ YOEDevOps / SRE

Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.