Skip to content
AlembicAlembic

Senior Network & Site Reliability Engineer

Design, operate, and automate the global network and reliability layer for a high-performance NVIDIA DGX SuperPOD supporting ML workloads. Own architecture, observability, incident response, and security for mission-critical infrastructure.

About the job

What You'll Do

  • Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.
  • Own network device configuration management end to end, ensuring consistency and reliability across the fleet.
  • Improve system and network reliability and performance through automation, observability, and proactive capacity planning.
  • Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.
  • Build and maintain comprehensive monitoring, alerting, and incident response — SLOs, runbooks, and on-call rotations — and drive post-incident analysis and continuous improvement.
  • Ensure security, compliance, and operational readiness across our network and cloud infrastructure.
  • Partner across engineering and data science to drive a culture of performance and reliability.

What Will Help You Succeed

  • 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.
  • A strong background in network security, architecture, design, and operations.
  • Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs.
  • Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP).
  • Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).
  • WAN engineering — carrier circuit provisioning and external network peering.
  • Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.
  • Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).
  • Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication.

Also Helpful

  • NVIDIA networking technologies — Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs.
  • Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).
  • Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments.

Skills

BGP, Vpn, Mpls, Ipsec, Ansible, Terraform, Kubernetes, Python, Prometheus, Grafana

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

Descript

Descript

San Francisco, CA

Software Engineer, Infrastructure
$220k+/yrRemote8+ YOEDevOps / SRE

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.

Onos Health

Onos Health

San Francisco, CA

Lead Infrastructure Engineer
$200k+/yrHybrid7+ YOEDevOps / SRE

Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.