Software Engineer - Reliability
Builds, maintains, and scales multi-cloud GPU infrastructure for AI training/inference, focusing on reliability, performance tuning, automation, and security in a fast-paced startup. Requires 8+ years SRE experience with deep Linux, cloud, and high-performance networking expertise.
About the job
What You’ll Do
- Architect for Reliability & Scale: Participate in critical re-architecture sessions to redesign our systems for higher efficiency and scale.
- Own Multi-Cloud GPU Clusters: Take end-to-end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance.
- Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices.
- Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level.
- Build Robust Automation: Write high-quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure.
- Debug Complex Hardware/Software Failures: Serve as the final escalation point for the most challenging GPU, networking (InfiniBand/RDMA), and system-level issues.
Who You Are
- 8+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
- Deep Linux mastery: hands-on expertise in Linux, containerized systems, and debugging low-level system performance.
- Cloud infrastructure expert: strong experience with providers like AWS or OCI.
- Tenacious troubleshooter for hardware/software intersections.
- Security-minded with knowledge of SOC 2 and ISO compliance.
- Expert in high-performance networking: InfiniBand, RDMA, or RoCE.
What Sets You Apart (Bonus Points)
- Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm.
- Experience managing large-scale GPU clusters for AI/ML workloads.
- Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray.
Skills
Linux, AWS, Oci, Kubernetes, Python, Go, Bash, InfiniBand, Rdma, Nvidia, Amd, Dcgm, Rocm, SOC 2
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.