Latest DevOps / SRE jobs
Job results
Staff Platform Engineer building and scaling the internal DevSecOps platform with IaC, Kubernetes, and multi-cloud automation. Requires 8+ years cloud/infra experience to drive reliability, cost optimization, zero-downtime strategies, and AI-enabled tooling.
Owns and automates multi-cloud, multi-region SaaS infrastructure, focusing on reliability, observability, performance, and incident response. Requires 5–7+ years of DevOps or infrastructure engineering experience, cloud tooling expertise, and business-level English and Japanese.
Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.
Staff engineer leading Pinterest's service communications platform. Architect and scale Envoy-based service mesh, mTLS identity, traffic optimization, and multi-language RPC frameworks for reliable, secure, high-volume service-to-service communication.
Build platform primitives for service provisioning, deploy tooling, workflow orchestration, and service ownership at a fast-scaling AI coding tool company. Requires experience with durable workflows like Temporal, internal dev platforms, and strong focus on developer experience and reliability.
Design, build, and operate scalable distributed systems and edge networks on AWS to handle Figma's growing customer traffic and services. Requires 4+ years building infrastructure at scale, experience with TypeScript or Go, and distributed/traffic systems.
Design, build, and operate petabyte-scale Product Metrics systems that process massive event volumes with strong reliability, performance, and availability. Requires 5+ years of distributed-systems experience, Golang expertise, cloud-platform experience, and familiarity with Kubernetes and infrastructure as code.
Senior Cloud Engineer responsible for designing, operating, and improving petabyte-scale Product Metrics systems built with Go, Kubernetes, and ClickHouse. The role requires 5+ years of experience with scalable distributed systems and strong reliability, performance, and production debugging skills.
Own stability and deployment of PostgreSQL products. Package software with Nix, manage upgrades, optimize CI/CD, and resolve production issues. Requires 3+ years PostgreSQL experience and Nix proficiency.
Design, build, and maintain developer infrastructure and tooling for Android Automotive OS, including Gerrit, static analysis, and build systems. Requires 2+ years of software engineering experience and automotive domain exposure.
Senior software engineer on the Production SRE team building reliability, monitoring, alerting, and incident-management tools for large-scale distributed systems. The role requires 5+ years of software engineering or SRE experience, strong programming skills, cloud and containerization expertise, and incident leadership.
Owns the technical direction of an AI-first platform, defining architecture, standards, and self-service capabilities that help engineering teams ship securely and efficiently. The role requires multi-team technical leadership, strong judgment on foundational trade-offs, and the ability to connect platform strategy to business outcomes.
Design and build AI automations, integrations, and agents that connect company systems and eliminate manual work. Requires 4+ years building production automations and LLMs, strong API skills, and experience with RAG, agents, and observability.
Ensure reliability of large GPU supercomputing clusters by diagnosing hardware/firmware/OS issues, automating monitoring, driving firmware rollouts, and working directly with vendors.
Own and debug multi-thousand-GPU network fabric (RDMA/RoCE, NVLink/NVSwitch) for large-scale AI training and inference. Requires backend language proficiency, large-scale cluster experience, and cross-stack ownership.
Build internal developer platform, tooling, and automation to accelerate engineering velocity. Focus on CI/CD pipelines, test infrastructure, build systems, and metrics to help engineers ship faster and more reliably.
Senior engineer building and scaling internal developer platforms with strong focus on AI tooling, reliability, and developer experience. Requires 4+ years in backend/infrastructure and proven project leadership.
Build AI-powered developer experience systems spanning bug resolution, CI and build reliability, development environments, testing, and internal tooling. The role requires 10+ years of software engineering experience, strong systems thinking, technical leadership, and expertise in reliable developer infrastructure.
Lead the rollout of Go as a fully supported, production-grade platform at Notion. Own service patterns, tooling, and guardrails while tackling high-leverage developer experience challenges across AI workflows, CI, and reliability.
Build and operate backend services and infrastructure for model customization and evaluation at Together AI. Requires 3+ years building production infrastructure, strong Python/Go skills, and deep experience with Kubernetes, Linux, and cloud platforms.
The Senior Site Reliability Engineer will build and operate secure, observable AI/ML infrastructure across cloud platforms. The role requires at least five years of production SRE or infrastructure experience, strong Terraform and observability expertise, and hands-on incident response and automation skills.
Own and scale ML infrastructure for robotics AI: backend services, GPU orchestration, storage, and internal developer platforms across cloud and on-prem.
Build and operate the cloud platform powering Zilliz Cloud and Vector Lakebase across multi-cloud environments, integrating control plane, scheduling, and database runtime for scalable AI workloads. Requires 3+ years building production systems, strong Kubernetes and cloud experience, and a bachelor's degree or equivalent.
Build software, services, and frameworks for network management, automation, and monitoring of large-scale GPU supercomputing fabrics. Requires deep network protocol knowledge and experience orchestrating tens of thousands of devices.
Software engineer building and operating the orchestration layer for a globally distributed, high-performance AI inference platform on custom wafer-scale hardware.
Design, deploy, and operate enterprise network infrastructure for corporate facilities and hybrid cloud environments with zero-trust architecture and compliance requirements. Requires 5+ years enterprise networking experience and ability to obtain TS/SCI clearance.
Hands-on technical lead building and operating the orchestration layer for a globally distributed, high-performance AI inference platform on custom wafer-scale hardware.
Staff SRE on the Release Engineering team defining and scaling reliability practices, architecting SLO/error-budget programs, and driving progressive delivery and automated safety gates across product engineering.
Builds automation and testing infrastructure for a platform, enabling reliable validation across environments and configurations. The role requires 5+ years of software engineering experience, strong Python and TypeScript skills, and expertise in CI/CD, containers, cloud platforms, and developer tooling.
Build and operate multi-tenant, distributed infrastructure for an enterprise AI platform across Kubernetes and multiple clouds. The role requires 5+ years of platform, infrastructure, or backend engineering experience, strong Golang and Python skills, and deep expertise in reliability, databases, and distributed systems.
Builds and operates multi-tenant, cloud-native platform infrastructure for an agentic AI product. The role requires 5+ years of platform, infrastructure, or backend engineering experience, strong Golang and Python skills, Kubernetes expertise, and deep distributed-systems knowledge.
SRE embedded in Service Operations to establish reliability practices, frameworks, and feedback loops across engineering teams. Focus on SLOs/SLIs, ORR processes, incident-to-improvement pipelines, and influencing without authority in a distributed environment.
Senior SRE responsible for production infrastructure reliability, incident response, deployment automation, and scaling SaaS systems on Kubernetes and major cloud platforms.
Operate and scale a cloud-native CTV advertising platform on AWS and Kubernetes. Focus on reliability, GitOps workflows, infrastructure automation, observability, and incident response.
Own internal developer platforms, CI/CD pipelines, AI-assisted tooling, and local dev environments to make Cape engineers faster and more confident. Report to engineering leadership and partner across platform, product, and security teams.
Senior Software Engineer on the DevOps and Tooling team building internal tools. Requires 3-5+ years experience, Rust or strong systems background, TypeScript/React, Linux, Docker, and CI/CD.
Senior engineer building and operating Astronomer's high-scale PaaS platform. Owns testing, deployment, reliability, and observability for Astro and related products.
Own technical strategy and roadmap for node lifecycle management, health automation, and scaling AI clusters across clouds and accelerators. Requires deep distributed systems expertise, ML accelerator experience, and 12+ years leading complex multi-team infrastructure initiatives.
Senior-level engineer to own and scale Anthropic's massive Kubernetes control plane and scheduler for training frontier AI models across hundreds of thousands of nodes. Requires deep Kubernetes internals experience and 12+ years building production distributed systems.
Own technical strategy and roadmap for agent-driven cluster lifecycle management across cloud providers and datacenters. Lead complex multi-quarter infrastructure initiatives and mentor engineers on large-scale compute systems.
Senior or Staff Site Reliability Engineer focused on continuous delivery infrastructure using Argo Workflows, ArgoCD, and Kubernetes. Owns deployment tooling, onboarding flows, and participates in 24/7 on-call. Requires 6+ years building and operating distributed systems.
Build and operate Argo's network performance and reliability platform powering Cloudflare products. Requires systems programming (Go/Rust/C/C++), deep networking knowledge (L3/L4, HTTP, TLS), and distributed systems experience.
Operates and improves Crusoe’s Kubernetes and virtual machine infrastructure, focusing on reliability, observability, incident response, and automation. Requires 3–6 years of production engineering experience, Kubernetes platform expertise, programming ability, and strong Linux and distributed-systems fundamentals.
Senior Platform Reliability Engineer establishing reliability standards, observability, and incident response practices across engineering teams. Requires 6+ years operating production systems at scale with AWS, Kubernetes, Terraform, and modern observability tooling.
Design, build, and operate AI-powered engineering tools and developer productivity platforms. Focus on AI pairing pipelines, automated workflows, and internal tooling to accelerate engineering velocity.
This senior DevOps role builds and operates AWS infrastructure for a SaaS platform and agentic AI services, with responsibility for reliability, observability, security, compliance, and disaster recovery. It requires 6+ years of platform experience, AI/ML infrastructure expertise, Terraform proficiency, and technical leadership.
Lead technical direction for Komodo's core control plane (KMC/PSS, identity, subscriptions) and App Builder/Connector. Architect platform primitives, APIs, and AI tooling in a multi-tenant SaaS environment.
Lead bringup, administration, and operations for a large-scale anime AI training GPU cluster. Bridge researchers and bare-metal systems using SLURM, filesystems, networking, and Linux sysadmin skills.
Staff-level engineer to own internal infrastructure and agentic productivity tools that accelerate SDLC, platform operations, and developer experience for a small AI underwriting startup.
Build and own CI/CD systems, agentic AI tooling, and developer platforms that power engineering velocity at a fast-growing healthcare AI company. Requires strong experience with modern build systems, Kubernetes, and AI-assisted development workflows.