Software Engineer, Frontier Systems
Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.
About the job
Responsibilities
- Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
- Build and evolve health checks that detect, remediate, and verify failures at scale.
- Ensure critical health checks execute with minimal latency to maximize workload uptime.
- Investigate hardware failures and system-level issues across large-scale compute environments.
- Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes.
- Build automation and tooling that enables global cluster management with minimal manual intervention.
- Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.
Requirements
- 7+ years of industry experience in software or infrastructure engineering.
- Strong proficiency with Python and shell scripting.
- Experience building large-scale distributed systems or infrastructure platforms.
- Comfort digging into noisy operational data using SQL, PromQL, or similar tooling.
- Experience building reproducible analyses and operational tooling.
- Strong systems debugging and operational instincts with an ownership mindset.
Nice-to-Haves
- Experience with low-level hardware systems and Linux tooling (e.g. PCIe, InfiniBand, RoCE, networking, power management, kernel performance tuning, FW/SW debugging).
- Experience operating or debugging large-scale GPU or accelerator clusters.
- Expertise in network operations, observability, or systems telemetry.
- Experience with automated remediation systems or fleet lifecycle management.
- Experience improving reliability, utilization, or workload uptime in distributed compute environments.
Skills
Python, SQL, Promql, Kubernetes, Linux, Pcie, InfiniBand, Roce, GPU, Distributed Systems
Similar jobs
DevOps / SRE jobsOwn and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.