Skip to content
Luma AILuma AI

Software Engineer - Reliability

Builds, maintains, and scales multi-cloud GPU infrastructure for AI training/inference, focusing on reliability, performance tuning, automation, and security in a fast-paced startup. Requires 8+ years SRE experience with deep Linux, cloud, and high-performance networking expertise.

About the job

What You’ll Do

  • Architect for Reliability & Scale: Participate in critical re-architecture sessions to redesign our systems for higher efficiency and scale.
  • Own Multi-Cloud GPU Clusters: Take end-to-end ownership of our production clusters for training and inference across AWS and OCI, ensuring high availability and peak performance.
  • Drive Security & Compliance: Assist in achieving and maintaining security certifications (SOC 2 Type 1 & 2, ISO standards) by implementing robust infrastructure security practices.
  • Deep Linux Performance Tuning: Use your mastery of Linux systems to troubleshoot and optimize performance at the OS and kernel level.
  • Build Robust Automation: Write high-quality tools and automation in Python, Go, or Bash to manage, monitor, and heal our infrastructure.
  • Debug Complex Hardware/Software Failures: Serve as the final escalation point for the most challenging GPU, networking (InfiniBand/RDMA), and system-level issues.

Who You Are

  • 8+ years of experience as an SRE, production engineer, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep Linux mastery: hands-on expertise in Linux, containerized systems, and debugging low-level system performance.
  • Cloud infrastructure expert: strong experience with providers like AWS or OCI.
  • Tenacious troubleshooter for hardware/software intersections.
  • Security-minded with knowledge of SOC 2 and ISO compliance.
  • Expert in high-performance networking: InfiniBand, RDMA, or RoCE.

What Sets You Apart (Bonus Points)

  • Deep expertise with GPU tooling for NVIDIA and AMD GPUs like DCGM or ROCm.
  • Experience managing large-scale GPU clusters for AI/ML workloads.
  • Familiarity with job management systems based on Kubernetes or orchestration frameworks like Ray.

Skills

Linux, AWS, Oci, Kubernetes, Python, Go, Bash, InfiniBand, Rdma, Nvidia, Amd, Dcgm, Rocm, SOC 2

Commure

Commure

Mountain View, CA
Senior Software Engineer, Infrastructure
$170k+/yrHybrid6+ YOEDevOps / SRE

Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.

Kindred

Kindred

United States
Senior Infrastructure Engineer
$170k+/yrRemote5+ YOEDevOps / SRE

Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.

Creditgenie

Creditgenie

Plymouth Meeting, PA
Senior DevOps/SRE Engineer
$170k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.

Reality Defender

Reality Defender

New York, NY

Senior Dev Ops Engineer
$170k+/yrOn-site7+ YOEDevOps / SRE

Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.

VSCO

VSCO

San Francisco, CA

Senior Software Engineer, Infrastructure
$165k+/yrHybrid5+ YOEDevOps / SRE

Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.