# Senior SRE AI Engineer

**Company:** [Navan](https://hotfix.jobs/companies/navan)
**Location:** Tel Aviv, Israel
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Python, Go, Java, Cloud Infrastructure, Kubernetes, Linux, Networking, CI/CD, Terraform, CloudFormation, Grafana, Prometheus, New Relic, Datadog, Splunk
**Posted:** 2026-07-18

> Senior SRE AI Engineer responsible for operating reliable production infrastructure and AI-powered applications, with ownership of automation, observability, incident response, and provider integrations. Requires 5+ years of SRE or infrastructure engineering experience and hands-on cloud, platform, and AI operations expertise.

## Job Description

## Responsibilities
- Support AI-based application solutions where reliability matters, partnering with development teams on development and production operations.
- Work with AI solutions, providers, and APIs, focusing on API reliability, authentication, quotas, rate limits, latency, and provider-specific operational constraints.
- Troubleshoot AI tools and provider issues across workflows, APIs, configuration, permissions, degraded responses, and related areas.
- Operate reliable production platforms and cloud infrastructure while helping product teams move quickly without compromising reliability.
- Improve observability by building dashboards, alerts, traces, logs, and runbooks tied to SLOs and customer impact.
- Apply AI to SRE workflows by prototyping and productionizing AI-assisted operational systems.
- Automate operational toil through tools, workflows, and automation that reduce repetitive manual work.

## Requirements
- 5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
- 3+ years of experience operating production, 24x7 customer-facing systems.
- Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
- Strong software engineering skills in Python, Go, Java, or a similar language, with emphasis on production-quality code, testing, monitoring, and documentation.
- Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
- Experience building, tuning, and automating observability systems.
- Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
- Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
- Ability to troubleshoot AI tools and provider/API issues, including rate limits, quotas, authentication, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations.
- Excellent communication skills and ability to work with stakeholders and domain experts across the company.

## Benefits and Compensation
- This position is based out of the Tel Aviv office.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior DevOps Engineer](https://hotfix.jobs/jobs/0899d52e-7a00-40c7-ade1-f757d30d835e) - Viz.ai - Tel Aviv, Israel
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/6d7b2812-de3f-4dfa-8586-9010e1e594e7) - Clickhouse - Remote
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/56b82d25-159d-41f5-94fb-fb1f71dd6a5c) - Clickhouse - Remote

**Apply:** https://hotfix.jobs/jobs/df414b01-4b54-48aa-a8ac-edcdc8432b6f
**Canonical:** https://hotfix.jobs/jobs/df414b01-4b54-48aa-a8ac-edcdc8432b6f