Latest DevOps / SRE jobs
Job results
Leads a global 24×7 DevOps and Customer Enablement organization supporting highly available SaaS operations. The role requires 12+ years in DevOps, SRE, cloud operations, or production engineering, including substantial experience managing engineering teams and driving reliability, incident response, automation, and customer outcomes.
Senior platform engineer responsible for shipping and operating scalable infrastructure and services supporting financial products. Requires at least five years building non-trivial products or services, with backend experience and strong collaboration, ownership, and communication skills.
Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.
Staff SIP Engineer responsible for growing, optimizing, and troubleshooting a distributed cloud-based VoIP platform built on Kamailio, FreeSWITCH, and Kubernetes. Requires deep expertise in SIP/RTP protocols, VoIP troubleshooting, and DevOps practices to ensure high reliability and call quality.
Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.
Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.
Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.
Staff SRE building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring in isolated environments to support national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI clearance.
Build and operate scalable infrastructure platforms that improve reliability, developer velocity, security, and operational efficiency across R&D. The role requires senior-level experience with cloud-based distributed systems, automation, incident response, and AI-assisted engineering.
Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.
Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.
DevOps Engineer embedded in platform teams to build and operate scalable AI infrastructure on GCP/AWS. Own CI/CD, Kubernetes, observability, reliability (SLIs/SLOs), automation, and performance at petabyte scale. Requires 4-5 years production DevOps/SRE experience.
Leads the founding Tel Aviv Production Engineering team, combining people management with hands-on reliability engineering, incident response, automation, and firmware optimization. Requires 8+ years in infrastructure, SRE, or production engineering and 2+ years of direct engineering leadership.
Infrastructure Engineer building low-latency, high-reliability APIs, streaming gateways, and observability for Arena's real-world AI model evaluation platform. Requires 4+ years backend/distributed systems experience with Go/Rust, LLM APIs, and cloud infra (K8s/Terraform).
Senior SRE AI Engineer responsible for operating reliable production infrastructure and AI-powered applications, with ownership of automation, observability, incident response, and provider integrations. Requires 5+ years of SRE or infrastructure engineering experience and hands-on cloud, platform, and AI operations expertise.
Build and ship software that automates data center commissioning, including test orchestration, data capture, pass-fail analysis, and live reporting by integrating with BMS, EPMS, and test equipment. Requires experience building test automation or orchestration systems for hardware/infrastructure and replacing manual processes with trusted software.
Founding Platform Engineer owning compute, orchestration, and infrastructure for CodeRabbit's GenAI code review platform on multi-region GCP. Build from 0-to-1 with deep Kubernetes, distributed systems, and cloud infrastructure expertise.
Build and standardize AI-powered coding tools, agents, and dev environments to accelerate internal software development while maintaining security and quality. Requires experience with productivity tooling for large codebases, container/CI platforms, and AI model APIs.
Site Reliability Engineer II responsible for designing, deploying, and maintaining multi-cloud infrastructure (Azure primary, AWS/GCP) for Illumio's SaaS products. Focus on IaC, CI/CD pipelines, monitoring, incident response, automation, and improving reliability/scalability in collaboration with engineering and security teams. Requires 2+ years SRE/DevOps experience with Azure.
Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.
Site Reliability Engineer responsible for deploying, monitoring, and maintaining highly available infrastructure on GCP using Terraform, Kubernetes, and various cloud services. Requires 5+ years experience (3+ in production), strong Kubernetes/IaC background, and on-call participation to ensure reliable systems for millions of users.
Senior Site Reliability Engineer responsible for monitoring, incident response, and optimizing the reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure and SaaS services. Requires 5+ years SRE experience with strong cloud platform expertise.
Senior Software Engineer responsible for production platform reliability, data pipeline troubleshooting, incident management, operational tooling, and automation. Requires 4+ years of engineering experience, strong Java and SQL skills, distributed data pipeline expertise, and cloud platform experience.
Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.
Lead technical direction for facilities production engineering, architecting telemetry pipelines from OT systems (BMS/EPMS/SCADA) into modern data stacks and setting controls integration standards across massive AI data center fleet.
Build and operate the scalable, low-latency infrastructure powering Arena's real-world AI model evaluation platform, including API gateways, observability, and enterprise features for frontier model routing and evaluation.
Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.
Staff Software Engineer improving efficiency and performance of Pinterest's large-scale Kubernetes and distributed cloud infrastructure. Requires deep capacity/performance expertise, experience leading efficiency initiatives at scale, and strong AI collaboration skills.
Build, scale, and operate OpenAI's global compute infrastructure for frontier AI models like GPT-5.6. Solve complex cross-disciplinary problems spanning distributed systems, hardware, ML infrastructure, power/cooling, manufacturing, supply chain, and data center development at unprecedented scale.
Infrastructure Engineer responsible for hands-on installation, provisioning, maintenance, and troubleshooting of high-performance on-premise server hardware, Linux systems, and high-speed networking (100G/400G) in a data center environment. Requires 3+ years experience with Linux admin, x86 hardware, and network configuration.
DevOps Engineer building and enhancing CI/CD pipelines for Scale's lowside and highside products in classified environments. Integrate ML tasks into automated SDLC, incorporate security best practices, and collaborate across teams. Requires active TS/SCI clearance.
Senior Software Engineer building and scaling the AI-native web platform layer at Databricks, including CI/CD, deployment automation, consent management, accessibility testing, and tech stack unification. Requires 8+ years experience in production web infrastructure.
This role builds and operates scalable platform infrastructure, automates operational and database workflows, and improves reliability, observability, and deployment practices. It requires at least eight years of platform experience, production systems expertise, Kubernetes, cloud, Linux, and SQL datastore experience.
Senior frontend platform engineer strengthening React, TypeScript, and Vite foundations, optimizing testing, CI/CD, automation, and observability to enable fast, high-quality frontend development across the organization. Requires 6-8+ years experience, leadership of complex projects, and deep frontend tooling expertise.
Cloud Operations Engineer on 2nd shift weekends responsible for monitoring Atlas platform, diagnosing incidents, on-call rotations, automation, and ensuring uptime for MongoDB customers in FedRamp environments. Requires 2+ years DevOps/SRE experience, Linux expertise, cloud familiarity, and scripting skills.
Sr/Staff SRE building and automating infrastructure for a fintech platform using AI agents, Terraform, Kubernetes, and GCP services. Focus on eliminating toil, improving observability, and ensuring reliability at scale for consumer apps processing billions in transactions.
Build the Site Reliability Engineering function from the ground up at Forward, defining SLOs, building observability infrastructure, leading incident response, and embedding reliability into the SDLC for their complex SaaS platform. Requires 6+ years SRE/DevOps experience, strong networking and Kubernetes skills, and a track record maturing SRE practices.
Platform engineer responsible for designing, building, and operating cloud infrastructure on AWS and Kubernetes. Focus on developer tooling, automation with AI, performance tuning, security, observability, and on-call support for a SaaS platform serving private capital markets. Requires 5+ years in DevOps/SRE/Platform roles.
Senior Software Engineer building Snowflake's Snowpark Container Services, a managed Kubernetes-based platform for running containerized applications inside Snowflake's Data Cloud. Lead engineering efforts on highly scalable, reliable, multi-tenant container compute infrastructure.
Software Engineer L2 responsible for evolving and maintaining Twilio's Compute infrastructure, including VM orchestration, AWS Auto Scaling Groups, hardened AMIs, secure container images, and automation of operational tasks in a remote-first environment.
Staff Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents that support 911 emergency call centers. Requires 6+ years in infrastructure/platform/backend roles with experience in reliability and scale.
Senior Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents used in 911 emergency response centers. Requires 4+ years in infrastructure/platform/backend roles with experience in reliability and scale.
Software Engineer owning Forus' compute platform (EKS/Kubernetes), data layer (Postgres, OpenSearch, BigQuery), cloud cost optimization, reliability (SLOs, observability), and IaC primitives in a regulated healthcare environment. Requires production Kubernetes at scale, deep AWS/Terraform expertise, and database migration experience.
Software Engineer owning CI/CD, build systems, and infrastructure for AI coding agents to accelerate the entire engineering organization's velocity at an AI-powered healthcare network company. Requires strong infrastructure instincts, Python/Bash expertise, and experience driving developer productivity and adoption of agentic workflows.
Own reliability and infrastructure for large-scale video and social systems, including on-call response, postmortems, database scaling, infrastructure as code, and CI/CD. The role requires deep GCP, Kubernetes, Terraform, Elasticsearch, and production database experience.
Site Reliability Engineer responsible for the reliability, observability, performance, and security of a core AI agent platform including code sandboxes. Requires 5+ years software engineering experience (3+ in SRE/DevOps), strong CS fundamentals, expertise in containers, cloud IaC, monitoring, and distributed systems.
Staff Engineer owning end-to-end delivery of Stripe's deployment platform, including orchestrator evolution, Kubernetes fleet migration, anomaly detection, and reliability. Requires 10+ years experience leading large-scale infrastructure projects with deep expertise in distributed systems, containers, and operational excellence.