Staff Software Engineer improving efficiency and performance of Pinterest's large-scale Kubernetes and distributed cloud infrastructure. Requires deep capacity/performance expertise, experience leading efficiency initiatives at scale, and strong AI collaboration skills.
177k – 365k/yr
Hybrid7+ YOEDevOps / SRE
About the role
What you’ll do
Improve the efficiency of large scale shared environments like Kubernetes
Improve the performance and efficiency of large scale distributed systems that drive Pinterest systems
Build, develop and mature profiling and optimization capabilities for Pinterest scale
Collaborate with Infrastructure Engineering and SRE teams in their mission to deliver highly available, resilient, secure and efficient foundations for Pinterest’s tech stack
Leverage AI to scale the impact of yourself and the team, including:
Accelerate performance investigations (e.g. quickly distill logs/metrics/traces and prior learnings) while verifying findings through measurement and testing
Build tooling and agents that allow users to self-serve efficiency insights and recommendations
Iterate faster on optimization approaches and rollout plans, then validate impact with experiments and production guardrails
What we’re looking for
Bachelor’s degree in computer science, a related field or equivalent experience
Deep understanding of infrastructure capacity and performance
Experience leading efficiency initiatives at scale on Kubernetes or other large scale shared infrastructure
Strong technical and performance engineering skills to collaborate with stakeholders on complex and ambiguous technical challenges
Experience building and managing highly available distributed applications at scale
Proficiency in software development languages such as Java, Python and C++
Excellent skills in communicating complex technical issues
Experience with AWS or similar cloud environments
Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables
Bonus points for the following
Hands-on experience with large, cloud-native multi-tenant platforms at Internet scale
Staff Software Engineer leading automation traffic management (bots/crawlers) and rate limiting at Pinterest's Edge (CDN, TLS, DNS, proxies). Design/implement Envoy-based L7 logic in C++/Go/Python; own roadmap, mentor, and drive reliability for 600M+ user platform.
177k – 365k/yr
Remote7+ YOEDevOps / SRE
Staff Software Engineer, Service Communications
PinterestSan Francisco, CA
Staff engineer leading Pinterest's service communications platform. Architect and scale Envoy-based service mesh, mTLS identity, traffic optimization, and multi-language RPC frameworks for reliable, secure, high-volume service-to-service communication.
177k – 365k/yr
Remote7+ YOEDevOps / SRE
Staff Software Engineer, Observability
PinterestSan Francisco, CA
Staff Software Engineer building and scaling Pinterest's observability platform (metrics, logs, traces) for massive distributed systems. Requires 7+ years distributed systems experience, strong data engineering skills, and expertise with modern observability tools.
177k – 365k/yr
Remote7+ YOEDevOps / SRE
Senior Staff Engineer, Cloud Site Operations
CrusoeSan Francisco, CA +1
Leads technical architecture for data center operations, overseeing global ticket queues, fleet supportability, power topology, resilience planning, and hardware failure escalations for AI infrastructure. Requires 10+ years in data center ops or HPC with deep NVIDIA GPU expertise.
179k – 218k/yr
On-site10+ YOEDevOps / SRE
Senior/Staff Site Reliability Engineer
SageNew York, NY
Leads design, operation, and evolution of highly reliable, scalable production infrastructure including cloud, databases, and observability. Drives incident response, SRE practices, automation, and capacity planning for large-scale distributed systems. Requires 7-12+ years in SRE/infrastructure engineering.