Skip to content
1,067 jobs

Job results

Crusoe

Crusoe

San Francisco, CA

Staff Production Engineer, Compute
$209k+/yrOn-site8+ YOEDevOps / SRE

Staff Production Engineer responsible for developing automation/observability, scaling virtualization (KVM/QEMU), optimizing Linux kernel performance, and supporting AI/HPC workloads on CPU/GPU/DPU hardware. Requires 8+ years in Linux systems engineering, kernel internals, and virtualization.

Basis

Basis

New York, NY

Member of Technical Staff
$100k+/yrOn-site5+ YOEDevOps / SRE

Site Reliability Engineer building and operating scalable, reliable infrastructure for an AI-powered accounting platform at a fast-growing startup. Requires 5+ years production infrastructure experience, strong software engineering skills, IaC, observability, and on-call ownership in an onsite NYC role.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA
Software Engineer
No salary listedOn-site2+ YOEDevOps / SRE

Build and maintain CI/CD pipelines, artifact management, cloud infrastructure, and developer productivity tooling at Cerebras to accelerate AI hardware and software engineering. Requires 2-5 years DevOps/infrastructure experience, Kubernetes, AWS, and strong troubleshooting skills.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA
Cloud Infrastructure Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Design, build, and operate secure cloud infrastructure and identity platforms in AWS and on-prem data centers. Implement IAM, automation with Terraform/Python/Go, security controls for AI systems, and Zero Trust principles while participating in on-call.

Flexport

Flexport

Amsterdam, Netherlands

Senior Software Engineer: Platform
No salary listedOn-site5+ YOEDevOps / SRE

Build and evolve Flexport’s internal developer platform, CI/CD systems, build tooling, service communication, IAM, orchestration, and observability capabilities. The role requires 5+ years of software engineering experience, strong full-stack instincts, and a product mindset for improving developer experience.

RSA

RSA

Bengaluru, India

Principal DevOps Engineer
No salary listedOn-site8+ YOEDevOps / SRE

Principal DevOps Engineer responsible for platform engineering across bare-metal and Azure environments, including Kubernetes, CI/CD automation, Jenkins lifecycle management, Infrastructure as Code, networking, reliability, and disaster recovery. Requires 8+ years of core DevOps experience and deep hands-on expertise.

Storable

Storable

Hyderabad, India

DevOps II
No salary listedOn-site4+ YOEDevOps / SRE

Builds and manages scalable infrastructure, CI/CD pipelines, containerized deployments, and monitoring systems. The role requires 4+ years of DevOps experience, strong Kubernetes and scripting skills, cloud platform knowledge, and familiarity with AI/ML deployment workflows.

Together AI

Together AI

San Francisco, CA

Staff Platform Engineer, Service Infrastructure
$240k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer owning service infrastructure strategy for Together AI's Product Foundations (API, UI, Billing, IAM). Lead Kubernetes, AWS, Terraform, networking, and reusable primitives to improve reliability, consistency, and scalability across teams.

Calendly

Calendly

California

Staff Platform Engineer
No salary listedRemote10+ YOEDevOps / SRE

Staff Platform Engineer building self-service infrastructure tools and platforms on Kubernetes and GCP. Requires 10+ years cloud/software experience, strong Go and Kubernetes expertise, and mentoring skills.

Orb

Orb

Senior Platform Engineer
$195k+/yrHybrid5+ YOEDevOps / SRE

Senior Platform Engineer building developer and application platforms at Orb, including CI/CD, IaC, sandboxes, preview environments, observability, and AI-assisted development tooling. Requires 5+ years owning platforms at scale with deep AWS, Terraform, containers, and CI/CD experience.

Decagon

Decagon

San Francisco, CA
Senior Software Engineer, Cloud Infrastructure
$200k+/yrOn-site4+ YOEDevOps / SRE

Build and operate scalable cloud infrastructure platforms, abstractions over Kubernetes and major cloud providers, and secure enterprise deployments for Decagon's agentic AI systems. Requires 4+ years in infrastructure/DevOps with deep Terraform, Kubernetes, and cloud networking experience.

Runpod

Runpod

United States

Site Reliability Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.

Anyscale

Anyscale

San Francisco, CA

Software Engineer, Infrastructure
$202k+/yrOn-site5+ YOEDevOps / SRE

Software Engineer building scalable control and data plane infrastructure for Anyscale's Ray platform. Design and optimize cluster orchestration, scheduling, Kubernetes deployments, and accelerator support for distributed AI/ML workloads. Requires 3+ years production experience with cloud-native tech, Go/Python, and distributed systems.

Chime

Chime

San Francisco, CA

Software Engineer, Infrastructure
$133k+/yrHybrid2+ YOEDevOps / SRE

Build and operate foundational data infrastructure at Chime, owning deployment platforms (Airflow, Flink) and core storage (DynamoDB, RDS). Requires 2+ years infrastructure/backend experience, Terraform, Kubernetes, AWS, and Python.

Zoox

Zoox

Foster City, CA
Senior Software Engineer, Build Infrastructure
$208k+/yrHybrid8+ YOEDevOps / SRE

Senior Software Engineer owning and evolving Zoox's Bazel-based build infrastructure for a large multi-language monorepo. Optimize build speeds, develop custom rules/tooling, integrate quality tools, and accelerate developer velocity for safety-critical autonomous vehicle software. Requires 8+ years experience including 4+ years with Bazel.

VGS

VGS

Greece

Senior Infrastructure Engineer, DevOps
$80k+/yrRemote5+ YOEDevOps / SRE

Own and scale AWS infrastructure, CI/CD pipelines, GitOps automation, observability, and cloud security for high-throughput payment systems. The role requires 5+ years of DevOps or infrastructure experience plus strong Kubernetes, automation, and SRE expertise.

Office Hours

Office Hours

San Francisco, CA
Platform Engineer
$160k+/yrRemote5+ YOEDevOps / SRE

Senior Platform Engineer owning AWS/EKS infrastructure, IaC with Terraform, GitOps, security, observability, and data systems for a fast-growing expert network marketplace. Requires 5+ years production Kubernetes/AWS experience, strong IaC and security skills.

Radar Labs

Radar Labs

New York, NY

Senior / Staff Platform Engineer
$200k+/yrOn-site7+ YOEDevOps / SRE

Build and operate Radar’s high-scale infrastructure, developer platform, and data systems to support 1B daily API calls. Generalist engineer focused on availability, self-serve capabilities, automation, and customer feedback.

Viz.ai

Viz.ai

Tel Aviv, Israel

Senior Platform Engineer
No salary listedHybrid6+ YOEDevOps / SRE

Builds and improves platform infrastructure and developer experience for AI, data, and product teams. The role requires 6+ years of backend or platform engineering experience, with expertise in cloud infrastructure, microservices, CI/CD, containers, monitoring, and high-level programming.

CrewAI

CrewAI

San Francisco, CA

Software Engineer, Infrastructure & Reliability
No salary listedOn-site5+ YOEDevOps / SRE

Build and operate the infrastructure foundation for CrewAI's multi-agent AI platform across AWS, Azure, and GCP. Own CI/CD, observability, reliability, security, and self-hosted deployment tooling for cloud and enterprise environments.

Illumio

Illumio

Sunnyvale, CA

Staff Engineer
$194k+/yrOn-site7+ YOEDevOps / SRE

Staff Engineer building scalable, resilient cloud architecture and distributed systems on AWS/Azure/GCP with Kubernetes for Illumio's Zero Trust cybersecurity platform. Requires strong cloud programming experience, distributed systems expertise, and 7+ years overall experience.

Crusoe

Crusoe

San Francisco, CA

Staff Software Engineer, Cloud Monitoring Service
$215k+/yrOn-site8+ YOEDevOps / SRE

Lead the design and evolution of Crusoe Cloud's large-scale telemetry and observability systems for metrics and logs. Own high-throughput distributed pipelines from edge collection through ingestion, storage, and low-latency querying while ensuring scalability, reliability, and multi-tenancy.

Stripe

Stripe

United States

Financial Connections TechOps Integration Reliability Engineer
No salary listedRemoteDevOps / SRE

Operate and expand banking/financial data integrations at Stripe while building and reviewing LLM-powered agentic automation. Requires strong technical problem-solving across code, data, logs, and partner coordination plus experience building with AI tooling.

Kustomer

Kustomer

New York, NY

Software Engineer, Infrastructure
$130k+/yrHybrid5+ YOEDevOps / SRE

Infrastructure Software Engineer building scalable backend systems, observability, and developer tools on the Foundation team. Lead projects on database sharding, event bus, search scaling, and latency; mentor engineers. Requires 5+ years with distributed systems, NoSQL (MongoDB), IaC (Terraform), and architecture ownership.

Cribl

Cribl

United States

Senior Site Reliability Engineer
$142k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for designing, implementing, and operating observability systems for complex cloud platforms. Focus on improving reliability, resilience, reducing toil through automation, and participating in on-call duties. Requires strong experience with IaC, cloud technologies, observability tools, and Linux engineering.

Together AI

Together AI

San Francisco, CA

Senior Software Engineer, Observability
$200k+/yrOn-site5+ YOEDevOps / SRE

Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.

Retool

Retool

San Francisco, CA
Site Reliability Engineer
$164k+/yrHybrid5+ YOEDevOps / SRE

Site Reliability Engineer owning reliability, automation, and upgrades across Retool Cloud, BYOC, and self-hosted Kubernetes deployments for enterprise customers. Requires deep AWS, Kubernetes, Terraform, and Postgres experience plus strong automation and customer collaboration skills.

Anthropic

Anthropic

San Francisco, CA
Staff+ Software Engineer, Caching
$320k+/yrHybrid10+ YOEDevOps / SRE

Build and lead Anthropic's managed caching infrastructure as a foundational service, including a scalable Redis fleet, client libraries, and CDC-driven invalidation. Requires deep distributed systems and caching expertise to optimize latency and consistency across hot paths for Claude.

Together AI

Together AI

San Francisco, CA
Senior Network Engineer
$190k+/yrOn-site8+ YOEDevOps / SRE

Senior Network Engineer responsible for designing, implementing, and maintaining high-performance compute network infrastructure for AI systems. Requires 8+ years experience with large-scale data center networks, deep expertise in routing/switching protocols, automation, and multi-vendor hardware.

Crusoe

Crusoe

San Francisco, CA

Senior Microsoft Cloud Infrastructure Engineer
$160k+/yrOn-site8+ YOEDevOps / SRE

Senior Cloud Infrastructure Engineer owning design, implementation, and management of Microsoft 365, Entra ID, Azure, and Azure Arc hybrid infrastructure. Requires 8+ years infrastructure experience with deep Azure/M365 expertise, IaC, Windows Server admin, and hands-on data center hardware work.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Principal Site Reliability Engineer
No salary listedOn-site15+ YOEDevOps / SRE

Principal SRE to architect self-service reliability platforms, capacity orchestration, and production control planes for Cerebras' ultra-high-speed AI inference infrastructure at massive scale. Requires 15+ years in SRE/platform engineering with deep large-scale fleet and observability experience.

Teleport

Teleport

United States

Senior Site Reliability Engineer
$222k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer building and scaling Teleport's secure SaaS cloud infrastructure. Focus on global scaling, observability, automation, incident response, and on-call in a security-first environment using Go and Kubernetes.

Fluidstack

Fluidstack

San Francisco, CA
Distributed Systems Engineer
$208k+/yrOn-site5+ YOEDevOps / SRE

Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.

Fluidstack

Fluidstack

San Francisco, CA
Backend Engineer
$208k+/yrOn-site5+ YOEDevOps / SRE

Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.

Fluidstack

Fluidstack

San Francisco, CA
Software Engineer, Cloud Infrastructure
$175k+/yrOn-site5+ YOEDevOps / SRE

Build and own the observability platform, control plane APIs, and fleet state management for a hyperscale GPU infrastructure powering AI compute at 10-100s of GW scale. Requires production service ownership at scale, comfort with AI coding tools, and on-call incident response.

Fluidstack

Fluidstack

San Francisco, CA
Production Engineer, Network
$175k+/yrOn-site5+ YOEDevOps / SRE

Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure at Fluidstack. Requires systems thinking, automation-first mindset, on-call ownership, and daily use of AI coding tools like Claude/Cursor alongside Go/Python and network protocols.

Anthropic

Anthropic

San Francisco, CA
[Pipeline] Staff+ Software Engineer, Experimentation
$405k+/yrHybrid7+ YOEDevOps / SRE

Staff+ Software Engineer owning the strategy, architecture, and development of Anthropic's configuration management, feature flagging, and large-scale experimentation platforms to enable safe, data-driven changes and boost developer productivity.

Anthropic

Anthropic

San Francisco, CA
[Pipeline] Staff+ Software Engineer, Developer Acceleration
$405k+/yrHybrid10+ YOEDevOps / SRE

Staff+ Software Engineer building Anthropic's Agent Runtime Platform and knowledge infrastructure to enable thousands of employees to be highly productive with AI agents. Requires 10+ years large-scale distributed systems experience and agent expertise to define agentic productivity, build runtimes, write evals, and drive operational excellence.

Sanity

Sanity

United States
Senior Site Reliability Engineer
No salary listedRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer partnering with development teams to design, build, and operate scalable GCP and Kubernetes infrastructure at high request volumes (75k RPS). Focus on observability, reliability, CI/CD, edge modernization, on-call, and mentoring to maintain platform performance and uptime.

Brex

Brex

San Francisco, CA
Senior Software Engineer, Release Infra
$192k+/yrHybrid7+ YOEDevOps / SRE

Senior Software Engineer building and operating Brex's release infrastructure, CI/CD pipelines, observability, and incident management systems. Requires 7+ years experience with backend languages, Kubernetes, cloud platforms, and SRE practices to ensure safe, scalable deployments.

Stackblitz

Stackblitz

United States

Staff Site Reliability Engineer
No salary listedRemote7+ YOEDevOps / SRE

Staff SRE embedding early with product/platform teams to design reliability, observability, and scalability into systems across multi-cloud environments. Define SLIs/SLOs, build golden paths with Terraform, lead incidents/postmortems, influence org-wide standards, and share real on-call rotation.

Clickhouse

Clickhouse

United States

Senior Cloud Engineer - Product Metrics
No salary listedRemote5+ YOEDevOps / SRE

Senior Cloud Engineer responsible for designing, operating, and improving petabyte-scale Product Metrics systems built with Golang, Kubernetes, and ClickHouse. The role requires 5+ years of experience with scalable distributed systems and production ownership.

Console

Console

San Francisco, CA

Infrastructure Engineer
$200k+/yrOn-site5+ YOEDevOps / SRE

Infrastructure Engineer building and operating scalable cloud and self-hosted platforms on AWS and Kubernetes. Own core systems for graph, search, and runtime while ensuring enterprise-grade reliability, security, and deployment for autonomous AI agents.

Upstart

Upstart

United States

DevOps Engineer, Cloud Platform
$136k+/yrRemote5+ YOEDevOps / SRE

Build and operate shared Kubernetes (EKS) and AWS cloud infrastructure powering Upstart's product and ML workloads. Requires 3+ years Kubernetes production experience plus strong AWS, IaC, and GitOps skills.

Cloudflare

Cloudflare

San Francisco, CA
Principal Software Engineer: Distributed Systems
$200k+/yrHybrid10+ YOEDevOps / SRE

Principal IC owning technical coherence of Cloudflare's Intent Management platform for safe, health-mediated configuration changes and deployments across global distributed storage, progressive releases, and testing systems. Requires 10+ years experience leading large-scale distributed systems initiatives, strong technical leadership, LLM integration, and expertise in at least one modern strongly-typed language.

Snowflake

Snowflake

Senior Software Engineer, Accelerated Delivery
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and evolve continuous deployment, progressive delivery, and AI-augmented release platforms for safe, large-scale multi-cloud rollouts on Kubernetes at Snowflake. Requires strong systems programming (Golang/Java/C++), Kubernetes, observability, and DevOps experience.

Shield AI

Shield AI

Washington, DC

Staff Engineer, HMS Factory
$150k+/yrOn-site7+ YOEDevOps / SRE

Staff Engineer building automated testing frameworks, log analysis tools, and scenario-generation scripts on the Hivemind Platform SDK to enable rapid verification of AI Pilot behaviors. Requires strong Python expertise, simulation/testing experience, and ability to obtain SECRET clearance.

Pinterest

Pinterest

Washington, DC

Staff Software Engineer, Edge
$177k+/yrRemote7+ YOEDevOps / SRE

Staff Software Engineer leading automation traffic management (bots/crawlers) and rate limiting at Pinterest's Edge (CDN, TLS, DNS, proxies). Design/implement Envoy-based L7 logic in C++/Go/Python; own roadmap, mentor, and drive reliability for 600M+ user platform.

Mux

Mux

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Senior Platform Engineer building CI/CD pipelines, developer productivity tools, and Kubernetes infrastructure at Mux to accelerate engineering velocity from local dev to production. Requires strong Go engineering, production Kubernetes experience, networking fundamentals, and cross-team partnership.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Senior Software Engineer
$153k+/yrOn-site5+ YOEDevOps / SRE

Senior software engineer owning a multi-environment compute platform for simulation orchestration and workload scheduling at scale. Requires 5+ years backend/distributed systems experience and deep Kubernetes expertise.