# Manager, Production Engineering

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** Tel Aviv, IL
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** Go, Python, C++, Linux, Container Orchestration, Temporal, On-Call Operations, Slis, SLOs, Error Budgets, InfiniBand, Rocev2, Bmc, Firmware, Dcgm
**Posted:** 2026-07-20

> Leads the founding Tel Aviv Production Engineering team, combining people management with hands-on reliability engineering, incident response, automation, and firmware optimization. Requires 8+ years in infrastructure, SRE, or production engineering and 2+ years of direct engineering leadership.

## Job Description

## Responsibilities
- Recruit, mentor, and establish a high-performing Production Engineering team in Tel Aviv.
- Partner with US and Dublin teams to run a follow-the-sun global on-call rotation.
- Champion blameless post-mortems focused on systemic failures.
- Drive alert-reduction initiatives and improve fleet signal-to-noise ratios.
- Automate manual workflows using runbook automation such as Temporal.
- Build predictive monitoring to identify SEV1/SEV2 events before customer impact.
- Govern Production Readiness Reviews and change control across compute, storage, networking, and platform teams.
- Ensure at least 30% of team bandwidth is dedicated to strategic automation, tooling, and firmware optimization.
- Convert recurring physical interventions into software-defined auto-remediation.

## Requirements
- 8+ years of experience in infrastructure, SRE, or production engineering environments.
- 2+ years of direct leadership experience with first-line engineering teams in a high-growth neocloud, hyperscaler, or large-scale distributed environment.
- Hands-on software engineering proficiency in Go, Python, C++, or a comparable systems language.
- Expert knowledge of Linux internals, container orchestration at scale, and root-cause analysis across physical-to-virtual boundaries.
- Experience running tiered on-call models and establishing SLIs, SLOs, and error budgets.
- Proven ability to reduce paging fatigue.

## Nice-to-haves
- Experience at a neocloud or AI infrastructure company operating large GPU clusters.
- Exposure to InfiniBand, RoCEv2, BMC, firmware qualification, or attestation.
- Familiarity with NVIDIA or AMD accelerator failure modes and DCGM counters.
- Experience managing major scaling milestones and systems designed for 10x fleet expansions.

## Benefits
- Benefits package supporting financial security, health, and well-being.
- Pension contributions and additional perks aligned with local market standards.

## Similar jobs

- [Senior Platform Engineer](https://hotfix.jobs/jobs/71245322-afd8-4019-8525-65088fed1493) - Shield AI - San Diego, CA - $141k – $212k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/d4ecdaa5-c0ab-49f3-baed-0b13deaaf6c0) - Shield AI - San Mateo, CA - $140k – $211k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/3317571d-7a67-4059-b23e-af1a9033cf96) - Astra - Remote - $190k – $230k/yr
- [Senior Software Engineer, Core Infra Systems](https://hotfix.jobs/jobs/6261dc33-aecd-4734-b00d-46f60f4f38c9) - Coinbase - Remote - $186k – $219k/yr
- [Senior Site Infrastructure Engineer](https://hotfix.jobs/jobs/847213bd-91b8-44f9-b450-477e4aa3714e) - Shield AI - Seattle, WA - $110k – $210k/yr

**Apply:** https://hotfix.jobs/jobs/3c6e082a-f3d5-4b5d-b5e5-7558069a8a6b
**Canonical:** https://hotfix.jobs/jobs/3c6e082a-f3d5-4b5d-b5e5-7558069a8a6b