Staff Software Engineer, Capacity Engineering
Staff Software Engineer improving efficiency and performance of Pinterest's large-scale Kubernetes and distributed cloud infrastructure. Requires deep capacity/performance expertise, experience leading efficiency initiatives at scale, and strong AI collaboration skills.
About the job
What you’ll do
- Improve the efficiency of large scale shared environments like Kubernetes
- Improve the performance and efficiency of large scale distributed systems that drive Pinterest systems
- Build, develop and mature profiling and optimization capabilities for Pinterest scale
- Collaborate with Infrastructure Engineering and SRE teams in their mission to deliver highly available, resilient, secure and efficient foundations for Pinterest’s tech stack
- Leverage AI to scale the impact of yourself and the team, including:
- Accelerate performance investigations (e.g. quickly distill logs/metrics/traces and prior learnings) while verifying findings through measurement and testing
- Build tooling and agents that allow users to self-serve efficiency insights and recommendations
- Iterate faster on optimization approaches and rollout plans, then validate impact with experiments and production guardrails
What we’re looking for
- Bachelor’s degree in computer science, a related field or equivalent experience
- Deep understanding of infrastructure capacity and performance
- Experience leading efficiency initiatives at scale on Kubernetes or other large scale shared infrastructure
- Strong technical and performance engineering skills to collaborate with stakeholders on complex and ambiguous technical challenges
- Experience building and managing highly available distributed applications at scale
- Proficiency in software development languages such as Java, Python and C++
- Excellent skills in communicating complex technical issues
- Experience with AWS or similar cloud environments
- Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
- Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
- High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables
Bonus points for the following
- Hands-on experience with large, cloud-native multi-tenant platforms at Internet scale
Skills
Kubernetes, Java, Python, C++, AWS, Distributed Systems, Performance Engineering, Profiling, Optimization, AI Tools
Similar jobs
DevOps / SRE jobsBuild and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.