Senior Network & Site Reliability Engineer
Design, operate, and automate the global network and reliability layer for a high-performance NVIDIA DGX SuperPOD supporting ML workloads. Own architecture, observability, incident response, and security for mission-critical infrastructure.
About the job
What You'll Do
- Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.
- Own network device configuration management end to end, ensuring consistency and reliability across the fleet.
- Improve system and network reliability and performance through automation, observability, and proactive capacity planning.
- Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.
- Build and maintain comprehensive monitoring, alerting, and incident response — SLOs, runbooks, and on-call rotations — and drive post-incident analysis and continuous improvement.
- Ensure security, compliance, and operational readiness across our network and cloud infrastructure.
- Partner across engineering and data science to drive a culture of performance and reliability.
What Will Help You Succeed
- 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.
- A strong background in network security, architecture, design, and operations.
- Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs.
- Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP).
- Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).
- WAN engineering — carrier circuit provisioning and external network peering.
- Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.
- Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).
- Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication.
Also Helpful
- NVIDIA networking technologies — Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs.
- Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).
- Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments.
Skills
BGP, Vpn, Mpls, Ipsec, Ansible, Terraform, Kubernetes, Python, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsSenior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.