Leads a team building and operating a scalable SDN control plane for multi-tenant VPC networking across large GPU fleets. The role combines hands-on distributed-systems architecture, cloud networking expertise, production reliability, and engineering people leadership.
215k – 260k/yr
On-site8+ YOEEngineering Management
About the role
Responsibilities
Define the roadmap for the VPC control plane and evolve network virtualization systems such as OVN/OVS.
Build services that program multi-tenant virtual networks across large-scale GPU fleets, including VPCs, subnets, security groups, load balancing, NAT, and internet/hybrid connectivity.
Design distributed control-plane services for intent APIs, state reconciliation, network policy compilation, and southbound programming of hosts and DPUs.
Drive scalability, reliability engineering, convergence and API-latency benchmarking, regression prevention, and incident response.
Support operational excellence within 3–6 month execution cycles.
Mentor and grow mid-level to senior distributed-systems engineers while setting technical standards and fostering accountability.
Partner with data-plane and cloud product teams to deliver low-latency, highly available networking for multi-tenant GPU clusters.
Requirements
6–8+ years of experience in distributed systems or cloud networking engineering.
2–4+ years of experience managing engineering talent.
Strong knowledge of SDN and network virtualization, including overlay networking, VPC constructs, routing, and control-plane architectures such as OVN/OVS or equivalent.
Hands-on experience building large-scale control planes, including state reconciliation, consensus and consistency trade-offs, API design, and fleet-wide configuration propagation.
Experience with Go, Kubernetes-style controllers, or similar technologies.
Understanding of operating multi-tenant cloud services, including SLOs, convergence-time and scale benchmarking, graceful degradation, and blast-radius containment.
Ability to resolve complex technical challenges in a fast-moving, execution-focused environment.
Nice-to-haves
Experience with DPU- or SmartNIC-programmed data planes.
Experience with cloud load balancing or NAT at scale.
Experience with network security policy engines.
Open-source contributions to OVN/OVS or SDN projects.
Compensation and Benefits
Compensation range of $215,000–$260,000 plus bonus.
Restricted Stock Units included in all offers.
Paid time off, paid holidays, and leave programs.
Comprehensive health, dental, and vision insurance.
Employer HSA contributions.
Paid parental leave.
Paid life insurance and short- and long-term disability coverage.
Professional development and tuition reimbursement.
Mental health and wellness support.
Commuter benefits, cell phone stipend, and 401(k) plan with company match up to 4% of salary.
Volunteer time off, global travel insurance, emergency assistance, daily meals allowance, and location-specific perks.
Skills
sdnovn/ovsvxlangenevevpcBGPevpnGoKubernetesebpfdpdkdpusmartnicDistributed Systems
Leads and scales an engineering team building a fault-tolerant managed AI platform for LLM workloads, including task queues, model management, scheduling, and agentic execution infrastructure. Requires 5+ years leading engineering teams plus depth in distributed systems, cloud-native platforms, and AI infrastructure.
215k – 260k/yrOn-site8+ YOEEngineering Management
Engineering Manager, Data Platform
CrusoeSan Francisco, CA
Lead and grow a team of data engineers building Crusoe's scalable data platform for AI and cloud intelligence. Define roadmap with cross-functional partners, ensure operational excellence, and drive technical architecture for data lakes, warehouses, and ETL.
215k – 260k/yrOn-site7+ YOEEngineering Management
Senior Engineering Manager
CreditgenieNew York, NY +3
Lead a backend engineering team at a fintech startup as a hands-on Senior Engineering Manager. Set technical direction for scalable financial systems, write production code, mentor engineers, and drive operational excellence in a fast-paced environment.
215k – 275k/yrOn-site8+ YOEEngineering Management
Engineering Manager, Telemetry Agent and Edge
CrusoeSan Francisco, CA
Lead a team of 4-6 engineers building and operating Crusoe's telemetry agent for metrics/logs from hosts and GPUs. Own delivery of next-gen agent releases against fixed deadlines while hiring, coaching staff-level engineers, and maintaining high operational standards for fleet-wide software.
215k – 260k/yrOn-site7+ YOEEngineering Management
Software Engineering Manager, Database
LangChainSan Francisco, CA
Hands-on Engineering Manager to lead a small systems team building SmithDB, LangChain's purpose-built storage and query layer for AI observability at massive scale. Write production Rust, drive architecture and performance, manage the team, and own the technical roadmap.