Latest DevOps / SRE jobs
Job results
Staff Production Engineer responsible for developing automation/observability, scaling virtualization (KVM/QEMU), optimizing Linux kernel performance, and supporting AI/HPC workloads on CPU/GPU/DPU hardware. Requires 8+ years in Linux systems engineering, kernel internals, and virtualization.
Site Reliability Engineer building and operating scalable, reliable infrastructure for an AI-powered accounting platform at a fast-growing startup. Requires 5+ years production infrastructure experience, strong software engineering skills, IaC, observability, and on-call ownership in an onsite NYC role.
Build and maintain CI/CD pipelines, artifact management, cloud infrastructure, and developer productivity tooling at Cerebras to accelerate AI hardware and software engineering. Requires 2-5 years DevOps/infrastructure experience, Kubernetes, AWS, and strong troubleshooting skills.
Design, build, and operate secure cloud infrastructure and identity platforms in AWS and on-prem data centers. Implement IAM, automation with Terraform/Python/Go, security controls for AI systems, and Zero Trust principles while participating in on-call.
Build and evolve Flexport’s internal developer platform, CI/CD systems, build tooling, service communication, IAM, orchestration, and observability capabilities. The role requires 5+ years of software engineering experience, strong full-stack instincts, and a product mindset for improving developer experience.
Principal DevOps Engineer responsible for platform engineering across bare-metal and Azure environments, including Kubernetes, CI/CD automation, Jenkins lifecycle management, Infrastructure as Code, networking, reliability, and disaster recovery. Requires 8+ years of core DevOps experience and deep hands-on expertise.
Builds and manages scalable infrastructure, CI/CD pipelines, containerized deployments, and monitoring systems. The role requires 4+ years of DevOps experience, strong Kubernetes and scripting skills, cloud platform knowledge, and familiarity with AI/ML deployment workflows.
Staff Platform Engineer owning service infrastructure strategy for Together AI's Product Foundations (API, UI, Billing, IAM). Lead Kubernetes, AWS, Terraform, networking, and reusable primitives to improve reliability, consistency, and scalability across teams.
Staff Platform Engineer building self-service infrastructure tools and platforms on Kubernetes and GCP. Requires 10+ years cloud/software experience, strong Go and Kubernetes expertise, and mentoring skills.
Senior Platform Engineer building developer and application platforms at Orb, including CI/CD, IaC, sandboxes, preview environments, observability, and AI-assisted development tooling. Requires 5+ years owning platforms at scale with deep AWS, Terraform, containers, and CI/CD experience.
Build and operate scalable cloud infrastructure platforms, abstractions over Kubernetes and major cloud providers, and secure enterprise deployments for Decagon's agentic AI systems. Requires 4+ years in infrastructure/DevOps with deep Terraform, Kubernetes, and cloud networking experience.
Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.
Software Engineer building scalable control and data plane infrastructure for Anyscale's Ray platform. Design and optimize cluster orchestration, scheduling, Kubernetes deployments, and accelerator support for distributed AI/ML workloads. Requires 3+ years production experience with cloud-native tech, Go/Python, and distributed systems.
Build and operate foundational data infrastructure at Chime, owning deployment platforms (Airflow, Flink) and core storage (DynamoDB, RDS). Requires 2+ years infrastructure/backend experience, Terraform, Kubernetes, AWS, and Python.
Senior Software Engineer owning and evolving Zoox's Bazel-based build infrastructure for a large multi-language monorepo. Optimize build speeds, develop custom rules/tooling, integrate quality tools, and accelerate developer velocity for safety-critical autonomous vehicle software. Requires 8+ years experience including 4+ years with Bazel.
Own and scale AWS infrastructure, CI/CD pipelines, GitOps automation, observability, and cloud security for high-throughput payment systems. The role requires 5+ years of DevOps or infrastructure experience plus strong Kubernetes, automation, and SRE expertise.
Senior Platform Engineer owning AWS/EKS infrastructure, IaC with Terraform, GitOps, security, observability, and data systems for a fast-growing expert network marketplace. Requires 5+ years production Kubernetes/AWS experience, strong IaC and security skills.
Build and operate Radar’s high-scale infrastructure, developer platform, and data systems to support 1B daily API calls. Generalist engineer focused on availability, self-serve capabilities, automation, and customer feedback.
Builds and improves platform infrastructure and developer experience for AI, data, and product teams. The role requires 6+ years of backend or platform engineering experience, with expertise in cloud infrastructure, microservices, CI/CD, containers, monitoring, and high-level programming.
Build and operate the infrastructure foundation for CrewAI's multi-agent AI platform across AWS, Azure, and GCP. Own CI/CD, observability, reliability, security, and self-hosted deployment tooling for cloud and enterprise environments.
Staff Engineer building scalable, resilient cloud architecture and distributed systems on AWS/Azure/GCP with Kubernetes for Illumio's Zero Trust cybersecurity platform. Requires strong cloud programming experience, distributed systems expertise, and 7+ years overall experience.
Lead the design and evolution of Crusoe Cloud's large-scale telemetry and observability systems for metrics and logs. Own high-throughput distributed pipelines from edge collection through ingestion, storage, and low-latency querying while ensuring scalability, reliability, and multi-tenancy.
Operate and expand banking/financial data integrations at Stripe while building and reviewing LLM-powered agentic automation. Requires strong technical problem-solving across code, data, logs, and partner coordination plus experience building with AI tooling.
Infrastructure Software Engineer building scalable backend systems, observability, and developer tools on the Foundation team. Lead projects on database sharding, event bus, search scaling, and latency; mentor engineers. Requires 5+ years with distributed systems, NoSQL (MongoDB), IaC (Terraform), and architecture ownership.
Senior Site Reliability Engineer responsible for designing, implementing, and operating observability systems for complex cloud platforms. Focus on improving reliability, resilience, reducing toil through automation, and participating in on-call duties. Requires strong experience with IaC, cloud technologies, observability tools, and Linux engineering.
Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.
Site Reliability Engineer owning reliability, automation, and upgrades across Retool Cloud, BYOC, and self-hosted Kubernetes deployments for enterprise customers. Requires deep AWS, Kubernetes, Terraform, and Postgres experience plus strong automation and customer collaboration skills.
Build and lead Anthropic's managed caching infrastructure as a foundational service, including a scalable Redis fleet, client libraries, and CDC-driven invalidation. Requires deep distributed systems and caching expertise to optimize latency and consistency across hot paths for Claude.
Senior Network Engineer responsible for designing, implementing, and maintaining high-performance compute network infrastructure for AI systems. Requires 8+ years experience with large-scale data center networks, deep expertise in routing/switching protocols, automation, and multi-vendor hardware.
Senior Cloud Infrastructure Engineer owning design, implementation, and management of Microsoft 365, Entra ID, Azure, and Azure Arc hybrid infrastructure. Requires 8+ years infrastructure experience with deep Azure/M365 expertise, IaC, Windows Server admin, and hands-on data center hardware work.
Principal SRE to architect self-service reliability platforms, capacity orchestration, and production control planes for Cerebras' ultra-high-speed AI inference infrastructure at massive scale. Requires 15+ years in SRE/platform engineering with deep large-scale fleet and observability experience.
Senior Site Reliability Engineer building and scaling Teleport's secure SaaS cloud infrastructure. Focus on global scaling, observability, automation, incident response, and on-call in a security-first environment using Go and Kubernetes.
Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.
Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.
Build and own the observability platform, control plane APIs, and fleet state management for a hyperscale GPU infrastructure powering AI compute at 10-100s of GW scale. Requires production service ownership at scale, comfort with AI coding tools, and on-call incident response.
Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure at Fluidstack. Requires systems thinking, automation-first mindset, on-call ownership, and daily use of AI coding tools like Claude/Cursor alongside Go/Python and network protocols.
Staff+ Software Engineer owning the strategy, architecture, and development of Anthropic's configuration management, feature flagging, and large-scale experimentation platforms to enable safe, data-driven changes and boost developer productivity.
Staff+ Software Engineer building Anthropic's Agent Runtime Platform and knowledge infrastructure to enable thousands of employees to be highly productive with AI agents. Requires 10+ years large-scale distributed systems experience and agent expertise to define agentic productivity, build runtimes, write evals, and drive operational excellence.
Senior Site Reliability Engineer partnering with development teams to design, build, and operate scalable GCP and Kubernetes infrastructure at high request volumes (75k RPS). Focus on observability, reliability, CI/CD, edge modernization, on-call, and mentoring to maintain platform performance and uptime.
Senior Software Engineer building and operating Brex's release infrastructure, CI/CD pipelines, observability, and incident management systems. Requires 7+ years experience with backend languages, Kubernetes, cloud platforms, and SRE practices to ensure safe, scalable deployments.
Staff SRE embedding early with product/platform teams to design reliability, observability, and scalability into systems across multi-cloud environments. Define SLIs/SLOs, build golden paths with Terraform, lead incidents/postmortems, influence org-wide standards, and share real on-call rotation.
Senior Cloud Engineer responsible for designing, operating, and improving petabyte-scale Product Metrics systems built with Golang, Kubernetes, and ClickHouse. The role requires 5+ years of experience with scalable distributed systems and production ownership.
Infrastructure Engineer building and operating scalable cloud and self-hosted platforms on AWS and Kubernetes. Own core systems for graph, search, and runtime while ensuring enterprise-grade reliability, security, and deployment for autonomous AI agents.
Build and operate shared Kubernetes (EKS) and AWS cloud infrastructure powering Upstart's product and ML workloads. Requires 3+ years Kubernetes production experience plus strong AWS, IaC, and GitOps skills.
Principal IC owning technical coherence of Cloudflare's Intent Management platform for safe, health-mediated configuration changes and deployments across global distributed storage, progressive releases, and testing systems. Requires 10+ years experience leading large-scale distributed systems initiatives, strong technical leadership, LLM integration, and expertise in at least one modern strongly-typed language.
Build and evolve continuous deployment, progressive delivery, and AI-augmented release platforms for safe, large-scale multi-cloud rollouts on Kubernetes at Snowflake. Requires strong systems programming (Golang/Java/C++), Kubernetes, observability, and DevOps experience.
Staff Engineer building automated testing frameworks, log analysis tools, and scenario-generation scripts on the Hivemind Platform SDK to enable rapid verification of AI Pilot behaviors. Requires strong Python expertise, simulation/testing experience, and ability to obtain SECRET clearance.
Staff Software Engineer leading automation traffic management (bots/crawlers) and rate limiting at Pinterest's Edge (CDN, TLS, DNS, proxies). Design/implement Envoy-based L7 logic in C++/Go/Python; own roadmap, mentor, and drive reliability for 600M+ user platform.
Senior Platform Engineer building CI/CD pipelines, developer productivity tools, and Kubernetes infrastructure at Mux to accelerate engineering velocity from local dev to production. Requires strong Go engineering, production Kubernetes experience, networking fundamentals, and cross-team partnership.
Senior software engineer owning a multi-environment compute platform for simulation orchestration and workload scheduling at scale. Requires 5+ years backend/distributed systems experience and deep Kubernetes expertise.