Staff Software Engineer, Service Communications
Staff engineer leading Pinterest's service communications platform. Architect and scale Envoy-based service mesh, mTLS identity, traffic optimization, and multi-language RPC frameworks for reliable, secure, high-volume service-to-service communication.
About the job
What you’ll do
- Architect and deploy advanced service mesh features, focusing on service discovery, traffic shaping, and deep observability using Envoy proxy.
- Lead the organization-wide adoption of service identity and mTLS to satisfy critical AAA security requirements for service-to-service paths.
- Design traffic optimization primitives like locality-aware routing to materially reduce data transfer costs for high-volume service traffic.
- Maintain and modernize service framework libraries in Java, Python, and C++, enhancing the developer experience and operational reliability.
- Collaborate with service owners across the company to drive adoption and refine infrastructure requirements for emerging feature needs.
- Partner with infrastructure peers on multi-region and multi-cloud strategies that rely on robust service communication primitives.
- Use AI to accelerate analysis and iteration, while applying judgment and verification to ensure correctness and quality.
- Join the team oncall rotation to manage incident response, perform post-mortems, and drive long-term reliability improvements.
What we’re looking for
- 6+ years of infrastructure or platform engineering experience, specifically within distributed systems or RPC framework development.
- Deep technical expertise in service mesh technologies such as Envoy or Istio, including hands-on experience with L7 proxying.
- Proficiency across multiple languages (Java, Python, C++) and a track record of building internal libraries or developer tools.
- Strong understanding of service security, including mTLS adoption and identity management via SPIFFE/SPIRE.
- Proven ability to design highly available and efficient distributed systems at massive scale.
- Demonstrated experience driving the adoption of complex platform capabilities across diverse cross-functional engineering teams.
- Demonstrated experience using AI to accelerate engineering workflows, with a clear approach to validating accuracy and quality.
- Bachelor’s/Master’s degree in Computer Science, a related field, or equivalent experience.
Skills
Envoy, Service Mesh, Mtls, Spiffe, Spire, Istio, Java, Python, C++, Distributed Systems, L7 Proxying, Traffic Shaping, Observability
Similar jobs
DevOps / SRE jobsBuild and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.