Skip to content
NavanNavan

Senior SRE AI Engineer

Senior SRE AI Engineer responsible for operating reliable production infrastructure and AI-powered applications, with ownership of automation, observability, incident response, and provider integrations. Requires 5+ years of SRE or infrastructure engineering experience and hands-on cloud, platform, and AI operations expertise.

About the job

Responsibilities

  • Support AI-based application solutions where reliability matters, partnering with development teams on development and production operations.
  • Work with AI solutions, providers, and APIs, focusing on API reliability, authentication, quotas, rate limits, latency, and provider-specific operational constraints.
  • Troubleshoot AI tools and provider issues across workflows, APIs, configuration, permissions, degraded responses, and related areas.
  • Operate reliable production platforms and cloud infrastructure while helping product teams move quickly without compromising reliability.
  • Improve observability by building dashboards, alerts, traces, logs, and runbooks tied to SLOs and customer impact.
  • Apply AI to SRE workflows by prototyping and productionizing AI-assisted operational systems.
  • Automate operational toil through tools, workflows, and automation that reduce repetitive manual work.

Requirements

  • 5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
  • 3+ years of experience operating production, 24x7 customer-facing systems.
  • Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
  • Strong software engineering skills in Python, Go, Java, or a similar language, with emphasis on production-quality code, testing, monitoring, and documentation.
  • Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
  • Experience building, tuning, and automating observability systems.
  • Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
  • Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
  • Ability to troubleshoot AI tools and provider/API issues, including rate limits, quotas, authentication, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations.
  • Excellent communication skills and ability to work with stakeholders and domain experts across the company.

Benefits and Compensation

  • This position is based out of the Tel Aviv office.

Skills

Python, Go, Java, Cloud Infrastructure, Kubernetes, Linux, Networking, CI/CD, Terraform, CloudFormation, Grafana, Prometheus, New Relic, Datadog, Splunk

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Viz.ai

Viz.ai

Tel Aviv, Israel

Senior DevOps Engineer
No salary listedHybrid6+ YOEDevOps / SRE

Owns and improves cloud infrastructure, CI/CD, Kubernetes, observability, security, scalability, and developer experience. The role requires at least six years of DevOps experience, strong AWS and automation expertise, and fluent Hebrew and English communication.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.