Senior SRE AI Engineer
Senior SRE AI Engineer responsible for operating reliable production infrastructure and AI-powered applications, with ownership of automation, observability, incident response, and provider integrations. Requires 5+ years of SRE or infrastructure engineering experience and hands-on cloud, platform, and AI operations expertise.
About the job
Responsibilities
- Support AI-based application solutions where reliability matters, partnering with development teams on development and production operations.
- Work with AI solutions, providers, and APIs, focusing on API reliability, authentication, quotas, rate limits, latency, and provider-specific operational constraints.
- Troubleshoot AI tools and provider issues across workflows, APIs, configuration, permissions, degraded responses, and related areas.
- Operate reliable production platforms and cloud infrastructure while helping product teams move quickly without compromising reliability.
- Improve observability by building dashboards, alerts, traces, logs, and runbooks tied to SLOs and customer impact.
- Apply AI to SRE workflows by prototyping and productionizing AI-assisted operational systems.
- Automate operational toil through tools, workflows, and automation that reduce repetitive manual work.
Requirements
- 5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
- 3+ years of experience operating production, 24x7 customer-facing systems.
- Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
- Strong software engineering skills in Python, Go, Java, or a similar language, with emphasis on production-quality code, testing, monitoring, and documentation.
- Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and infrastructure as code such as Terraform or CloudFormation.
- Experience building, tuning, and automating observability systems.
- Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
- Practical experience or strong interest in AI solutions, AI providers, agents, AI APIs, provider integrations, or AI-assisted internal tools.
- Ability to troubleshoot AI tools and provider/API issues, including rate limits, quotas, authentication, permission errors, latency, SDK or API contract changes, content quality issues, and service degradations.
- Excellent communication skills and ability to work with stakeholders and domain experts across the company.
Benefits and Compensation
- This position is based out of the Tel Aviv office.
Skills
Python, Go, Java, Cloud Infrastructure, Kubernetes, Linux, Networking, CI/CD, Terraform, CloudFormation, Grafana, Prometheus, New Relic, Datadog, Splunk
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Owns and improves cloud infrastructure, CI/CD, Kubernetes, observability, security, scalability, and developer experience. The role requires at least six years of DevOps experience, strong AWS and automation expertise, and fluent Hebrew and English communication.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.