Staff Software Engineer, Traffic
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.
About the job
Responsibilities
Invent
- Lead the design and development of systems that optimize network traffic and scale for global expansion.
- Drive architectural decisions for high-impact projects, ensuring scalability and reliability.
- Co-author long-term technical roadmaps for network scalability, performance, and engineering velocity.
- Set the standard for engineering and operational excellence.
Own
- Build partnerships across engineering, product, and security teams to align on network infrastructure goals.
- Lead design reviews for critical projects, focusing on system-level tradeoffs and network scalability.
- Drive the design and implementation of secure-by-default network systems with security teams.
- Engage with customers and internal teams to understand business requirements and deliver solutions.
Learn
- Leverage Temporal software to build and scale networking infrastructure.
- Gather customer insights and incorporate them into technical decisions for traffic management and networking strategies.
- Stay current with advances in traffic management, networking, and cloud orchestration.
Collaborate
- Mentor and guide engineers on best practices and design principles for reliable, scalable networking systems.
- Drive alignment across teams and ensure roadmaps and deliverables remain on track.
- Foster a collaborative, growth-oriented environment.
Requirements
- Experience leading complex engineering efforts focused on network traffic management, network optimization, and cloud orchestration.
- Exceptional collaboration and communication skills, including cross-functional leadership.
- Experience contributing to long-term technical roadmaps and making system-level tradeoffs.
- At least 8+ years of coding experience in Go, Java, or similar languages.
- Strong expertise writing concurrent and distributed code.
- Extensive experience designing and building distributed systems.
- Experience with concurrency primitives and network performance optimization.
- Deep expertise in traffic and networking systems.
- Experience with cloud providers such as AWS, Google Cloud, or Azure.
- Ability to optimize cost, performance, and scalability.
- Strong ownership and ability to balance short-term priorities with long-term strategic goals.
Compensation and Benefits
- Estimated pay range: $212,000–$280,000.
- Eligibility to participate in Temporal's equity plan.
- Unlimited PTO, 12 holidays, and 2 floating holidays.
- 100% premium coverage for medical, dental, and vision insurance.
- AD&D, long-term and short-term disability, and life insurance.
- Empower 401(k) plan.
- Learning and development, lifestyle spending, home office setup, professional membership, work-from-home meals, internet stipend, and other perks.
- International benefits vary by country and are provided in partnership with Remote.com.
- Occasional travel may be required for company events, team offsites, and other in-person gatherings.
Skills
Go, Java, AWS, GCP, Microsoft Azure, Temporal, Distributed Systems, Concurrent Programming, Network Traffic Management, Network Optimization, Cloud Orchestration, Network Performance, Concurrency Primitives, Network Security
Similar jobs
DevOps / SRE jobsLeads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.