Skip to content
1,067 jobs

Job results

The Voleon Group

The Voleon Group

Berkeley, CA

Senior Cluster Site Reliability Engineer
$205k+/yrRemote5+ YOEDevOps / SRE

Senior SRE scales and maintains high-performance research compute clusters for machine learning in finance, ensuring uptime, reliability, and observability across on-prem and cloud infrastructure. Requires 5+ years SRE/DevOps experience with HPC frameworks, IaC, and cloud services.

Seeq

Seeq

Oregon

Staff Software Engineer - Developer Experience
No salary listedRemote8+ YOEDevOps / SRE

Designs and builds developer tools, SDKs, APIs, and CI/CD pipelines to boost engineering productivity. Requires 8+ years experience in platform tooling, proficiency in Java/Python/TypeScript/Go, and cloud platforms like AWS/GCP.

The Voleon Group

The Voleon Group

United States

Senior Linux Infrastructure Engineer
$215k+/yrRemote7+ YOEDevOps / SRE

Builds, automates, and maintains scalable, secure Linux infrastructure using modern tooling. Requires 7+ years Linux sysadmin experience, deep OS/network knowledge, Python/Ruby scripting, config management (Ansible/Puppet/Chef), and monitoring (Prometheus/Grafana).

Crusoe

Crusoe

Sunnyvale, CA
Senior API Integration Engineer
$165k+/yrOn-site7+ YOEDevOps / SRE

Leads design and delivery of enterprise API integrations using Workato for People Tech ecosystem, automating workflows across ERP, CRM, HCM systems. Requires 7+ years experience with Workato recipes, iPaaS patterns, API security, and stakeholder collaboration.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Staff Software Engineer, Inference Cloud
No salary listedOn-site8+ YOEDevOps / SRE

Staff engineer owns architecture of Inference Cloud Platform, building distributed systems for high-QPS AI workloads with focus on availability, latency, reliability, and global scale. Requires 8+ years experience in large-scale cloud systems and backend languages like Go, C++, Python.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Principal Engineer, Inference Cloud
No salary listedOn-site10+ YOEDevOps / SRE

Principal Engineer leads Inference Cloud Platform, defining architecture for multi-region, high-QPS AI inference systems. Focuses on reliability, performance optimization, production code, and cross-team technical strategy. Requires 10+ years in distributed systems.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA
Staff Site Reliability Engineer – Automation and Platform
No salary listedRemote8+ YOEDevOps / SRE

Leads automation and platform engineering for ultra-reliable AI inference infrastructure, architecting self-service GitOps pipelines, observability, and tooling to eliminate toil across datacenters. Requires 8+ years SRE experience with large-scale clusters and tools like Argo CD and Prometheus.

Supabase

Supabase

Remote

Multigres Deployment Engineer
No salary listedRemoteDevOps / SRE

Owns deployment and operational infrastructure for Multigres distributed Postgres platform. Builds and maintains Go-based Kubernetes operator, architects cloud deployments on EKS/other K8s, manages storage/networking, and develops tooling for reliable scaling.

Twenty

Twenty

Arlington, VA

Senior / Staff DevSecOps Engineer
No salary listedOn-site8+ YOEDevOps / SRE

Builds and owns security infrastructure including runtime security, IAM, secrets management, CI/CD hardening, and compliance for cloud/container environments. Embeds with engineering teams to enable secure-by-default development using AWS, Terraform, and policy-as-code tools. Requires 8+ years DevSecOps experience.

Crusoe

Crusoe

San Francisco, CA

Senior Staff Software Engineer, Managed Orchestration
$238k+/yrOn-site10+ YOEDevOps / SRE

Leads architecture and development of scalable managed Kubernetes and AI orchestration systems, providing technical direction for cloud infrastructure reliability and performance. Requires 10+ years in software engineering with deep expertise in Go, Kubernetes, and large-scale systems.

OpenAI

OpenAI

San Francisco, CA
ChatGPT Performance Engineer
$325k+/yrRemote7+ YOEDevOps / SRE

Performance Engineer optimizes infrastructure and application performance for ChatGPT and OpenAI API, focusing on latency, throughput, and efficiency at scale. Requires 7+ years in high-scale systems with expertise in profiling, tracing, and cross-layer optimizations.

Shield AI

Shield AI

San Diego, CA
Staff Engineer, DevOps (4797)
$150k+/yrOn-site7+ YOEDevOps / SRE

Owns and maintains C++ build systems for autonomous aircraft software, improves developer velocity by optimizing CI/CD pipelines, integrates testing with simulations, and implements monitoring to resolve issues quickly. Requires 7+ years experience with deep expertise in build tools and DevOps practices.

Alchemy

Alchemy

San Francisco, CA
Cloud Infrastructure Engineer
$135k+/yrHybrid5+ YOEDevOps / SRE

Designs, deploys, and improves scalable blockchain infrastructure using Kubernetes, Terraform, and cloud tools. Drives AI enablement, builds observability with Prometheus/Grafana, manages multi-cloud networks, and leads incident response. Requires 5+ years in SRE/infrastructure with strong automation focus.

Forterra

Forterra

Clarksburg, MD

Senior AWS Cloud Engineer
$145k+/yrOn-site3+ YOEDevOps / SRE

Designs, manages, and secures AWS cloud infrastructure across multi-account environments using Terraform for autonomous vehicle platforms. Requires 3+ years AWS experience, strong IaC skills, and security focus; GovCloud/FedRAMP preferred.

Polymath

Polymath

San Francisco, CA

Member of Technical Staff - Engineering
$200k+/yrOn-siteDevOps / SRE

Builds infrastructure and simulation engines for training autonomous AI agents in complex environments using reinforcement learning. Requires strong engineering skills in distributed systems, containerization, networking, and data systems.

Pylon

Pylon

Palo Alto, CA

Infrastructure Engineer, Foundation
$140k+/yrHybridDevOps / SRE

Infrastructure Engineer on the Foundation team builds and maintains highly available systems and developer tooling to ensure platform stability and productivity for processing mortgage transactions. Requires deep curiosity, full ownership from design to maintenance, and ability to solve hard problems under pressure.

Sfcompute

Sfcompute

San Francisco, CA

HPC/ GPU Cluster Architect
$220k+/yrHybrid5+ YOEDevOps / SRE

Designs, architects, and scales production GPU/HPC clusters globally. Debugs hardware/software issues, automates operations, and mentors juniors. Requires 5+ years experience and hybrid SF presence.

NexHealth

NexHealth

San Francisco, CA

Senior Software Engineer, Platform
$165k+/yrOn-site5+ YOEDevOps / SRE

Build and scale backend platform systems including real-time EHR integrations, data lakes from TB to PB scale, and core infrastructure for 100x growth. Requires 5+ years in scalable backends, Kubernetes, AWS, PostgreSQL, and DevOps practices.

Vannevar

Vannevar

United States

DevOps Engineer
No salary listedRemote5+ YOEDevOps / SRE

Owns build and deployment automation, platform health, scaling, and self-service tools in a multi-account AWS environment. Requires 5+ years DevOps experience, deep AWS knowledge, IaC tools like Terraform/Pulumi, Docker, Python/Bash scripting.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Infrastructure - Analytics Platform
$230k+/yrHybridDevOps / SRE

Owns end-to-end production-critical infrastructure for analytics platform, building performant backend systems in Rust or C++ and operating distributed services at scale on Kubernetes. Requires strong systems experience in performance optimization, debugging, and on-call reliability.

Crusoe

Crusoe

San Francisco, CA

Staff Software Engineer, Systems Engineering Focus
$210k+/yrOn-siteDevOps / SRE

Designs, builds, and scales customer-facing managed services with a focus on edge agents running on customer infrastructure. Provides technical oversight for high-reliability systems using eBPF, Kubernetes, and low-level Linux metrics; leads cross-team collaboration and mentors engineers.

PointOne

PointOne

New York, NY

Product Reliability Engineer
$100k+/yrOn-site2+ YOEDevOps / SRE

Owns end-to-end system reliability, incident response, observability, and proactive stability improvements in a serverless AWS environment. Requires 2+ years software engineering with production-facing experience, strong debugging, and hands-on AWS/Go/TypeScript skills.

Square

Square

San Francisco, CA
Senior Site Reliability Engineer
$161k+/yrOn-site5+ YOEDevOps / SRE

Senior SRE improves platform reliability using AI-driven automation, leads incident response and oncall for critical services, and implements safe deployment practices. Requires 5+ years experience in backend/platform engineering, observability, and high-availability systems.

Phonely

Phonely

San Francisco, CA

DevOps Engineer
$180k+/yrOn-site5+ YOEDevOps / SRE

Designs, builds, and operates reliable cloud infrastructure for real-time voice AI systems. Owns Kubernetes clusters, CI/CD pipelines, observability, and security using AWS and IaC tools. Requires 5+ years DevOps experience with strong Python and async programming skills.

MongoDB

MongoDB

Boston, MA
Senior Site Reliability Engineer
$127k+/yrRemote6+ YOEDevOps / SRE

Senior SRE ensures reliability of MongoDB's multi-tenant cloud storage layer by defining SLOs, building resilient infrastructure, optimizing performance, and participating in on-call. Requires 6+ years experience with distributed systems, Python/Go, Kubernetes, and cloud platforms.

Sola

Sola

Remote

Software Engineer, Desktop Automation
$160k+/yrRemoteDevOps / SRE

Build and own desktop automation execution platform integrating AI agents with Windows sessions via remote protocols like VNC/RDP. Requires deep systems knowledge in OS APIs, accessibility frameworks, and automation infrastructure for reliable enterprise workflows.

MongoDB

MongoDB

Toronto, Canada

Staff Site Reliability Engineer, Fabric
$127k+/yrHybrid10+ YOEDevOps / SRE

Staff SRE on the Fabric team builds and maintains secure multi-cloud networking infrastructure for service communication, leveraging deep networking expertise to ensure resilience and scalability. Requires 10+ years experience in distributed systems and networking fundamentals.

Anthropic

Anthropic

San Francisco, CA
Staff+ Software Engineer, Platform
$405k+/yrHybrid8+ YOEDevOps / SRE

Staff-level software engineer builds and scales platform infrastructure across teams, including dev tools, service infra, multicloud, auth, connectivity, API distributability, and ML adaptation systems. Requires 8+ years full-stack experience with Staff leadership, focusing on robust, scalable solutions in fast-paced AI environment.

Zocks

Zocks

Budapest, Hungary

Senior DevSecOps Engineer
No salary listedHybrid6+ YOEDevOps / SRE

Own and secure multi-region cloud infrastructure, delivery pipelines, and network boundaries while driving vulnerability remediation, incident response, and compliance readiness. The role requires 6+ years of DevOps or security experience and deep AWS security expertise.

Together AI

Together AI

San Francisco, CA

Director, Data Center Operations
$250k+/yrOn-siteDevOps / SRE

Lead design, fit-out, and commissioning of data center sites focused on power, cooling, and IT infrastructure for high-density GPU workloads. Build and manage a 20-person operations team, oversee multi-site portfolio, and establish processes from scratch.

Eve

Eve

San Mateo, CA

Platform Engineer
$195k+/yrHybridDevOps / SRE

Build scalable infrastructure and data pipelines for AI/ML applications in legal tech. Collaborate with ML and dev teams to optimize performance and enhance developer productivity; requires cloud expertise and bachelor's/master's in CS.

Nuro

Nuro

Mountain View, CA

Senior Software Engineer, Performance Tooling and Infrastructure
$183k+/yrOn-site5+ YOEDevOps / SRE

Owns performance simulation platform infrastructure for validating autonomy code changes on robot hardware, including benchmarking orchestration, data pipelines, observability, and fleet management. Requires 5+ years experience in Python/C++, Linux systems, data engineering, with technical leadership.

Nuro

Nuro

Mountain View, CA

Software Engineer, Performance Tooling and Infrastructure
$152k+/yrOn-site3+ YOEDevOps / SRE

Builds and maintains performance simulation platform with bench-top rigs, cloud orchestration, and data pipelines to validate autonomy code changes for real-time performance on robot hardware. Requires 3+ years experience in Python/C++, Linux systems, data engineering, with technical leadership.

Cursor

Cursor

San Francisco, CA
Software Engineer, Infrastructure
No salary listedOn-siteDevOps / SRE

Builds and operates core infrastructure including Kubernetes clusters, geo-deployments, edge security, and cost optimization to support rapid growth and team productivity. Requires deep AWS/K8s experience and strong engineering fundamentals.

Cursor

Cursor

San Francisco, CA
Engineering Manager, Infrastructure
No salary listedOn-siteDevOps / SRE

Leads infrastructure team managing cloud, networking, storage, and compute for high-scale developer tool. Sets technical direction, codes, hires, and optimizes costs/regional deployments with deep Kubernetes/AWS expertise.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Engineer, Cloud Site Operations
$179k+/yrOn-site10+ YOEDevOps / SRE

Leads technical architecture for data center operations, overseeing global ticket queues, fleet supportability, power topology, resilience planning, and hardware failure escalations for AI infrastructure. Requires 10+ years in data center ops or HPC with deep NVIDIA GPU expertise.

Thatch

Thatch

San Francisco, CA

Software Engineer: Infrastructure
$180k+/yrRemoteDevOps / SRE

Infrastructure Engineer owns platform reliability, security, and performance in a HIPAA-compliant environment handling healthcare data. Designs infrastructure with Terraform, improves CI/CD workflows, and ensures resilient systems using PostgreSQL and modern cloud tools.

Crusoe

Crusoe

Denver, CO

Data Center Systems Engineer, R&D
No salary listedOn-site15+ YOEDevOps / SRE

Leads R&D for next-generation data center architectures, evaluating emerging technologies across power, cooling, compute, and facilities. Defines reference designs, guides pilot-to-hyperscale progression, and influences cross-functional teams. Requires 15+ years experience and systems thinking across engineering domains.

Deepgram

Deepgram

United States

Systems Architect AI/ML Infrastructure
$160k+/yrRemote7+ YOEDevOps / SRE

Designs end-to-end infrastructure architecture for AI/ML workloads, including multi-cloud strategies, GPU orchestration, storage for massive datasets, and capacity planning to support production inference and research training at scale. Requires 7+ years in systems architecture with deep expertise in Kubernetes, AWS, and GPU infrastructure.

Phylo

Phylo

South San Francisco, CA
Member of Technical Staff - System Engineering
$200k+/yrOn-siteDevOps / SRE

Builds and operates core distributed systems, infrastructure, and service architecture for Phylo's agentic AI platform, ensuring reliable scaled execution across cloud and enterprise environments. Requires 3+ years in backend/infrastructure engineering with Kubernetes and cloud expertise.

Docker

Docker

Seattle, WA

Senior Software Engineer, Infrastructure
$161k+/yrRemote6+ YOEDevOps / SRE

Senior engineer building self-service internal platform capabilities at Docker, focusing on multi-region networking, continuous deployment, EKS foundations, and AI-assisted operations to enable faster, safer provisioning for engineering teams.

Krea

Krea

San Francisco, CA

Engineer, Supercomputing & Distributed Systems
No salary listedOn-siteDevOps / SRE

Builds and operates supercomputing infrastructure for AI research including 1000+ GPU Kubernetes clusters, distributed data pipelines processing petabytes, and fault-tolerant training systems. Requires strong distributed systems intuition and experience with Python, PyTorch, and large-scale infrastructure.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Software Engineer, Supercomputing
$350k+/yrOn-siteDevOps / SRE

Designs, builds, and operates GPU supercomputing environments for large-scale AI training and inference. Automates cluster management, extends orchestration systems, and optimizes performance metrics in collaboration with researchers.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Software Engineer, Systems Generalist
$350k+/yrOn-siteDevOps / SRE

Builds and scales core infrastructure for AI model training, data systems, and developer tools in a high-impact team. Requires backend proficiency (Python/Rust), experience with large-scale clusters like Kubernetes, and end-to-end project ownership.

Clickhouse

Clickhouse

Netherlands

Database Reliability Engineer - Core Team
No salary listedRemote5+ YOEDevOps / SRE

Build and lead reliability engineering practices for ClickHouse Core, improving production performance, observability, incident response, and database operations. The role requires at least five years of reliability, QA, or customer-facing engineering experience, plus production SQL database and cloud expertise.

Clickhouse

Clickhouse

EMEA

Database Reliability Engineer - Core Team
No salary listedRemote5+ YOEDevOps / SRE

The Database Reliability Engineer will improve the reliability, scalability, performance, and incident response practices for ClickHouse Core. The role requires at least five years of reliability, QA, or customer-facing engineering experience, production database operations expertise, and strong scripting and debugging skills.

Clickhouse

Clickhouse

United Kingdom

Database Reliability Engineer - Core Team
No salary listedRemote5+ YOEDevOps / SRE

Build and lead reliability practices for ClickHouse Core, improving database performance, observability, incident response, and scalability. The role requires at least five years of reliability, QA, or customer-facing engineering experience plus production SQL database operations and cloud expertise.

Clickhouse

Clickhouse

EMEA

Database Reliability Engineer - Core Team
No salary listedRemote5+ YOEDevOps / SRE

The Database Reliability Engineer will improve ClickHouse Core’s reliability, scalability, performance, and production operations through observability, incident response, chaos initiatives, and database debugging. The role requires at least five years of reliability, QA, or customer-facing engineering experience and production SQL database expertise.

Clickhouse

Clickhouse

Netherlands

Database Reliability Engineer - Core Team
No salary listedRemote5+ YOEDevOps / SRE

Build and lead reliability practices for ClickHouse Core, improving production performance, observability, incident response, and scalability. The role requires at least five years in reliability, QA, or customer-facing engineering plus production database operations experience.

Propel

Propel

Brooklyn, NY
Senior Platform Engineer
$170k+/yrRemote5+ YOEDevOps / SRE

Senior Platform Engineer owns cloud services, Kubernetes clusters, and developer tooling on AWS to enable engineering teams. Requires 5+ years DevOps/SRE experience, strong Kubernetes/Terraform/AWS skills, and on-call participation.