Senior Production Engineer, Operational Excellence
Senior Production Engineer ensures reliability, scalability, and performance of GPU cloud infrastructure powering AI workloads. Drives observability, incident response, automation, and operational improvements in large-scale distributed systems.
About the job
Responsibilities
- Collaborate with cross-functional teams to define and evolve availability metrics for Crusoe’s cloud platform, including establishing, measuring, and improving SLIs and SLOs
- Participate in production incident response, diagnosing and resolving service disruptions while contributing to post-incident reviews and root cause analysis
- Build, operate, and improve observability across Crusoe’s infrastructure using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry
- Identify reliability risks, performance bottlenecks, and early indicators of potential production issues across distributed systems
- Develop automation and tooling that reduces operational toil, improves recovery times, and enables self-healing infrastructure
- Partner with compute, networking, storage, and platform teams to strengthen service resilience and disaster recovery capabilities
- Contribute to improving operational processes, knowledge sharing, and reliability best practices across the engineering organization
Requirements
- 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
- Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems
- Strong knowledge of Linux/Unix systems, including debugging complex issues across kernel and user space
- Previous experience in Infrastructure roles building or managing compute, storage or networking platforms
- Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms (AWS/GCP)
- Familiarity with incident management practices and reliability frameworks (SRE, ITIL, or similar)
- Experience with monitoring and observability tools such as Prometheus and Grafana
- Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Ansible
- Scripting or programming experience with languages such as Go, Python, C, or C++
- Strong communication skills and the ability to collaborate across engineering teams
- Ability to remain calm and effective while troubleshooting complex issues in high-impact production environments
Nice-to-Haves
- Experience working with Kubernetes or container orchestration platforms at scale
- Exposure to change management processes, operational readiness reviews, or structured root cause analysis
- Experience designing self-healing systems, automated remediation, or event-driven operational tooling
- Interest in scaling AI or HPC infrastructure and solving reliability challenges in GPU-heavy environments
Compensation
- Base salary range: $172,000 – $209,000 + Bonus
- Restricted Stock Units included
Skills
Prometheus, Grafana, Alertmanager, OpenTelemetry, Kubernetes, Linux, Terraform, Ansible, Python, Go, AWS, GCP, SRE, Gpu Workloads, Hpc
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Leads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.