Manager, Production Engineering
Leads the founding Tel Aviv Production Engineering team, combining people management with hands-on reliability engineering, incident response, automation, and firmware optimization. Requires 8+ years in infrastructure, SRE, or production engineering and 2+ years of direct engineering leadership.
About the job
Responsibilities
- Recruit, mentor, and establish a high-performing Production Engineering team in Tel Aviv.
- Partner with US and Dublin teams to run a follow-the-sun global on-call rotation.
- Champion blameless post-mortems focused on systemic failures.
- Drive alert-reduction initiatives and improve fleet signal-to-noise ratios.
- Automate manual workflows using runbook automation such as Temporal.
- Build predictive monitoring to identify SEV1/SEV2 events before customer impact.
- Govern Production Readiness Reviews and change control across compute, storage, networking, and platform teams.
- Ensure at least 30% of team bandwidth is dedicated to strategic automation, tooling, and firmware optimization.
- Convert recurring physical interventions into software-defined auto-remediation.
Requirements
- 8+ years of experience in infrastructure, SRE, or production engineering environments.
- 2+ years of direct leadership experience with first-line engineering teams in a high-growth neocloud, hyperscaler, or large-scale distributed environment.
- Hands-on software engineering proficiency in Go, Python, C++, or a comparable systems language.
- Expert knowledge of Linux internals, container orchestration at scale, and root-cause analysis across physical-to-virtual boundaries.
- Experience running tiered on-call models and establishing SLIs, SLOs, and error budgets.
- Proven ability to reduce paging fatigue.
Nice-to-haves
- Experience at a neocloud or AI infrastructure company operating large GPU clusters.
- Exposure to InfiniBand, RoCEv2, BMC, firmware qualification, or attestation.
- Familiarity with NVIDIA or AMD accelerator failure modes and DCGM counters.
- Experience managing major scaling milestones and systems designed for 10x fleet expansions.
Benefits
- Benefits package supporting financial security, health, and well-being.
- Pension contributions and additional perks aligned with local market standards.
Skills
Go, Python, C++, Linux, Container Orchestration, Temporal, On-Call Operations, Slis, SLOs, Error Budgets, InfiniBand, Rocev2, Bmc, Firmware, Dcgm
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.