Skip to content
1,066 jobs

Job results

Eloquent AI

Eloquent AI

San Francisco, CA

AI Engineer, AIOps & Infrastructure
No salary listedOn-site5+ YOEDevOps / SRE

Designs and builds scalable AI infrastructure for deploying enterprise AI agents, automates LLMOps/MLOps workflows, optimizes GPU workloads, and ensures production reliability using Kubernetes and cloud platforms. Requires 5+ years in MLOps/infrastructure and deep expertise in Python and distributed systems.

Zuma

Zuma

United States

Senior Engineer
No salary listedRemote3+ YOEDevOps / SRE

Builds and scales backend systems, cloud infrastructure, and deployment pipelines for an AI platform serving property managers. Requires 3+ years experience with Node.js, IaC tools like Terraform, and DevOps practices.

Palantir

Palantir

London, United Kingdom

Forward Deployed Infrastructure Engineer - UK Government
No salary listedHybridDevOps / SRE

Operates, deploys, and improves reliable Palantir infrastructure and services for government environments. The role emphasizes production troubleshooting, automation, scalable systems, cross-functional collaboration, and security clearance eligibility.

Palantir

Palantir

New York, NY
Backend Software Engineer - Infrastructure
No salary listedHybridDevOps / SRE

Builds performant, scalable infrastructure for Palantir's Foundry platform, including distributed systems, search/indexing, and container orchestration. Requires strong coding skills in backend languages like Java/Rust/Go and familiarity with storage/data systems; engineering degree preferred.

Orchestra Bio

Orchestra Bio

San Francisco, CA

Senior Platform Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Builds and maintains scalable cloud infrastructure powering an AI-driven R&D platform for biotech, focusing on reliability, security, AI workloads, and developer tooling. Requires 5+ years in platform engineering with expertise in cloud IaC, CI/CD, and observability.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Cloud Infrastructure
$230k+/yrOn-site5+ YOEDevOps / SRE

Builds and maintains cloud infrastructure abstractions for scalable, reliable product platforms like ChatGPT. Requires 5+ years in core infrastructure, Kubernetes at scale, and cloud abstractions; onsite in San Francisco with on-call duties.

Railway

Railway

Remote

Senior Platform Engineer: Storage
No salary listedRemoteDevOps / SRE

Designs and evolves production Ceph storage clusters, builds APIs and orchestration services for block/object storage using Go and gRPC. Requires experience with distributed systems, filesystems like ZFS/BTRFS, and building scalable infrastructure.

Pulse

Pulse

San Francisco, CA

DevOps Engineer
$150k+/yrOn-site4+ YOEDevOps / SRE

Designs and operates secure multi-tenant platforms using Kubernetes and AWS, managing provisioning, monitoring, CI/CD with Terraform and GitHub Actions, and ensuring security/compliance. Requires 4+ years DevOps experience with strong fundamentals in listed technologies.

Convex

Convex

San Francisco, CA

Software Engineer, Infra/Systems
$200k+/yrHybrid5+ YOEDevOps / SRE

Designs, builds, and maintains global cloud infrastructure for Convex's backend platform, focusing on performance, reliability, and scalability. Requires 5+ years backend experience with systems at scale, strong architecture skills, and on-call ownership in a startup environment.

ElevenLabs

ElevenLabs

Dublin, Ireland
Automations Engineer
No salary listedRemoteDevOps / SRE

Designs and builds internal tools, AI agents, and automations for GTM, Operations, and Finance teams using full-stack skills. Partners cross-functionally to optimize sales lifecycle and drive business impact with Python, SQL, and AI expertise.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Caching Infrastructure
$230k+/yrOn-site5+ YOEDevOps / SRE

Design, build, and operate a multi-tenant caching platform powering OpenAI's inference, identity, and products. Requires 5+ years in distributed systems with deep Redis/Memcached expertise and Kubernetes experience.

Perplexity

Perplexity

San Francisco, CA
AI Infra Engineer
$220k+/yrHybrid3+ YOEDevOps / SRE

Builds, deploys, and optimizes large-scale Kubernetes and Slurm clusters for AI training and inference. Requires 3+ years in ML infrastructure, expert Kubernetes/Slurm admin, Python/C++, and PyTorch experience.

Hebbia

Hebbia

New York, NY
Platform Engineer, Document Intelligence
$160k+/yrOn-site5+ YOEDevOps / SRE

Platform Engineer building high-scale distributed document indexing and search systems for an AI platform serving top financial institutions. Requires 5+ years experience in backend/distributed systems with Python/Java/Go and cloud infrastructure.

Zuma

Zuma

San Francisco, CA

Staff Engineer (Backend, DevOps, Infrastructure)
No salary listedHybrid5+ YOEDevOps / SRE

Staff Engineer leads architecture rewrites, builds scalable cloud infrastructure with IaC, optimizes data systems, and establishes DevOps practices for AI-powered property management platform. Requires 5+ years backend experience with Node.js and cloud expertise.

Vooma

Vooma

San Francisco, CA

Software Engineer: Backend & Infrastructure
No salary listedOn-siteDevOps / SRE

Builds and maintains secure, scalable infrastructure for AI-powered agents, deploys ML workloads, manages cloud environments, IaC, CI/CD, and ensures observability and security on AWS/Azure/GCP/Kubernetes.

Palantir

Palantir

London, United Kingdom

Software Engineer, New Grad - Infrastructure
No salary listedHybridDevOps / SRE

Build and improve large-scale infrastructure platforms supporting Palantir products and developer workflows. The role is suited to a new graduate with an engineering background, programming experience, and familiarity with cloud, storage, frontend, and infrastructure technologies.

Phantom

Phantom

Remote

Staff Software Engineer (SRE)
$200k+/yrRemote5+ YOEDevOps / SRE

Staff SRE manages and scales Kubernetes clusters on AWS EKS, automates infrastructure with IaC tools, optimizes performance, maintains blockchain nodes and databases, and improves system reliability using monitoring tools. Requires 5+ years SRE experience with strong Kubernetes and AWS proficiency.

Crusoe

Crusoe

San Francisco, CA
Senior Software Engineer, Managed Orchestration (Managed Kubernetes)
$180k+/yrOn-site5+ YOEDevOps / SRE

Build and scale managed Kubernetes and AI training clusters, developing operators, controllers, and infrastructure using Go, Terraform, and GCP. Design reliable, high-performance systems competing with GKE/EKS, with 5+ years experience required.

Palantir

Palantir

Washington, DC

Forward Deployed Site Reliability Engineer
No salary listedHybrid4+ YOEDevOps / SRE

Forward Deployed Site Reliability Engineer responsible for building, operating, and maintaining scalable infrastructure in air-gapped on-prem environments for US Government customers. Requires 4+ years Linux admin experience, hardware/networking knowledge, scripting skills, 50% travel availability, and active Top Secret clearance.

Abridge

Abridge

San Francisco, CA
Senior Platform Engineer
$179k+/yrHybrid8+ YOEDevOps / SRE

Senior Platform Engineer designs, builds, and scales cloud infrastructure (80% focus) including Kubernetes and GCP, implements CI/CD pipelines, security practices, and developer tools to boost engineering velocity and maintain compliance at hyperscale.

Zilliz

Zilliz

Redwood City, CA

Senior Site Reliability Engineer Cloud Platform
$175k+/yrHybrid4+ YOEDevOps / SRE

Senior SRE focuses on ensuring reliability, availability, and performance of distributed database systems in cloud-native environments. Requires 4+ years experience with Kubernetes, Docker, cloud platforms (AWS/GCP/Azure), IaC tools, and scripting in Python/Go/Java.

Loti AI

Loti AI

United States

Web3 Infrastructure DevOps Engineer
No salary listedRemoteDevOps / SRE

DevOps engineer building and maintaining decentralized Web3 infrastructure, including IPFS clusters, Arweave redundancy, cross-chain bridges, blockchain nodes, and Graph Protocol subgraphs. Requires proficiency in IPFS, Arweave, The Graph, Docker, and Kubernetes.

AngelList

AngelList

San Francisco, CA

Senior Infrastructure Engineer
No salary listedHybridDevOps / SRE

Builds and owns scalable platform infrastructure on AWS and Kubernetes, including deployment pipelines, automation, and observability tools. Provides technical leadership to enable efficient engineering workflows and production reliability.

Loti AI

Loti AI

United States

Staff Platform Engineer
No salary listedRemote8+ YOEDevOps / SRE

Leads architecture, development, and scaling of high-throughput API-based transaction processing platforms handling massive volumes. Mentors engineers and requires 8+ years experience in distributed systems, cloud platforms like AWS, and languages like Go/Python/C++.

Blacksmith

Blacksmith

New York, NY

Systems Engineer
$200k+/yrOn-site2+ YOEDevOps / SRE

Builds and scales distributed systems infrastructure for high-performance CI workloads, focusing on performance/reliability in filesystem, network, and storage. Requires 2+ years experience with C/C++/Go/Rust/Zig and onsite work in NYC.

Sieve

Sieve

San Francisco, CA

Distributed Systems Engineer
$150k+/yrOn-site3+ YOEDevOps / SRE

Designs and engineers distributed systems for compute, scheduling, and orchestration of ML and ETL pipelines handling large-scale video data. Requires 3+ years building data infrastructure, proficiency in Go/Python, and experience with petabyte-scale pipelines and CI/CD.

Inngest

Inngest

San Francisco, CA

Distributed Systems Engineer - Platform
No salary listedRemote2+ YOEDevOps / SRE

Builds scalable distributed systems including queueing, state stores, and execution layers for developer tools platform. Requires experience with Go, distributed systems at scale, and strong engineering judgment. Works with US PST overlap.

Braintrust

Braintrust

San Francisco, CA
Cloud Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

Builds and maintains scalable cloud infrastructure using Terraform and Kubernetes, owns CI/CD pipelines and observability, supports multi-cloud deployments, and assists customers with self-hosting. Requires 5+ years in DevOps/SRE, deep AWS experience, and programming skills.

Zocks

Zocks

Budapest, Hungary

Senior DevOps Engineer
No salary listedHybrid5+ YOEDevOps / SRE

The Senior DevOps Engineer will build and operate scalable, secure multi-region cloud infrastructure, networking, monitoring, and delivery pipelines across cloud and private environments. The role requires 5+ years of DevOps or cloud engineering experience, strong Linux and AWS expertise, and participation in production on-call support.

Radar Labs

Radar Labs

New York, NY

Senior / Staff Site Reliability Engineer
$200k+/yrOn-siteDevOps / SRE

Senior/Staff SRE maintains high-availability production infrastructure on AWS EKS with Terraform, manages sharded MongoDB clusters, participates in on-call rotation, and engages with customers to enhance reliability for 1B+ daily API calls.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Developer Productivity
$210k+/yrOn-site5+ YOEDevOps / SRE

Builds and maintains foundational systems, tools, and processes to boost developer productivity and engineering velocity at OpenAI. Requires 5+ years engineering experience, including infrastructure tooling, with core tech like Kubernetes, Python, and Terraform. Onsite in SF HQ.

Datology AI

Datology AI

Redwood City, CA

Software Engineer, Infrastructure
$180k+/yrOn-siteDevOps / SRE

Builds scalable data and ML infrastructure supporting multi-cloud and deployment models. Partners with founders and engineers on core platforms, tooling, and research features for reliable production systems.

Baseten

Baseten

San Francisco, CA
Software Engineer - Infrastructure
$165k+/yrHybridDevOps / SRE

Builds and maintains infrastructure components for ML inference platform using Python and Go. Implements Kubernetes deployments, monitoring systems, and resource management for efficient model serving, requiring Kubernetes knowledge and ML basics.

Perplexity

Perplexity

Belgrade, Serbia
Member Of Technical Staff
No salary listedOn-siteDevOps / SRE

Build and operate low-latency infrastructure and distributed backend systems powering a high-throughput search stack. The role requires strong cloud, Linux, automation, observability, and systems-language experience across AWS, Kubernetes, Rust, and Go.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Fleet Infrastructure
$230k+/yrHybridDevOps / SRE

Designs, implements, and operates infrastructure systems for model training and deployment on a massive GPU fleet. Requires experience with hyperscale compute, Kubernetes, public clouds like Azure, and strong programming skills.

Palantir

Palantir

Singapore

Production Engineer - Database Operations
No salary listedHybridDevOps / SRE

This engineer automates deployment and operation of production databases, manages migrations, improves reliability, and participates in incident response and on-call rotations. The role requires programming experience, cloud and orchestration expertise, database knowledge, and a strong Linux foundation.

Railway

Railway

United States

Infra Engineer - Datacenters
No salary listedRemoteDevOps / SRE

Builds and operates global datacenters bridging physical infrastructure to software, focusing on performance, resilience, and reliability. Requires experience with Golang, gRPC, Postgres, OS primitives, oncall rotations, and delivering 0→N projects remotely.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Frontier Clusters Infrastructure
$230k+/yrOn-siteDevOps / SRE

Builds and scales massive Kubernetes clusters for OpenAI's frontier supercomputers, automates bare-metal provisioning, and ensures reliability across data centers for AI model training. Requires expertise in distributed systems, Kubernetes operations, and infrastructure automation.

Factory AI

Factory AI

San Francisco, CA

Platform Engineer
No salary listedOn-site3+ YOEDevOps / SRE

Builds and optimizes core platform services including search, indexing, infrastructure, auth, billing, and developer tooling. Collaborates with AI teams and requires 3+ years full stack experience with strong systems programming skills.

Modal

Modal

New York, NY
Member of Technical Staff - Systems
$200k+/yrOn-site5+ YOEDevOps / SRE

Designs, builds, and maintains high-performance distributed systems for a serverless AI platform. Requires 5+ years experience in production code, large-scale systems, cloud, OS foundations, and performance optimization; onsite in NYC.

Palantir

Palantir

Seattle, WA
Software Engineer - Environment Platform
No salary listedHybrid3+ YOEDevOps / SRE

Develops APIs and Kubernetes controllers to enable containerization across Palantir platforms like Foundry, Gotham, Apollo for diverse infrastructures. Requires 3+ years software experience, Go/Java proficiency, system design, and bachelor's in CS.

DISQO

DISQO

Los Angeles, CA

Senior Site Reliability Engineer
$170k+/yrHybrid6+ YOEDevOps / SRE

Senior Site Reliability Engineer building agentic AI platforms, MCP servers, and automation tools to achieve zero KTLO. Drives SecDevOps culture with focus on security, reliability, observability while partnering on product roadmap and incident response. Requires 6+ years SRE/DevOps experience plus deep expertise in AWS, EKS, Terraform, Kubernetes.

Railway

Railway

Remote

Infrastructure Engineer
No salary listedRemoteDevOps / SRE

Builds low-level infrastructure software from first principles to power a cloud platform, focusing on OS primitives, distributed systems, and scalable gRPC services in Golang/Rust for high developer leverage.

Suno

Suno

Boston, MA

Staff / Senior Software Engineer, Infrastructure
$220k+/yrOn-site5+ YOEDevOps / SRE

Builds and operates scalable infrastructure systems including Kubernetes clusters, distributed databases, and cloud services to support AI music platform at consumer scale. Requires 5+ years experience in infrastructure engineering with strong ownership and scaling expertise.

Braintrust

Braintrust

San Francisco, CA
Software Engineer, Systems
No salary listedOn-siteDevOps / SRE

Builds high-performance data processing systems for real-time analysis of massive unstructured data at enterprise scale, focusing on storage, indexing, query optimization, and Braintrust's btql language. Requires expert systems programming in C++ or Rust, concurrency, databases, and OS knowledge.

Palantir

Palantir

Seattle, WA
Senior Software Engineer, Network Infrastructure
No salary listedHybrid3+ YOEDevOps / SRE

Senior Software Engineer builds and maintains APIs for Kubernetes-based network infrastructure, managing traffic flows across zero-trust clusters. Requires 3+ years experience, Kubernetes expertise, Go/Java skills, and networking knowledge.

Spellbrush

Spellbrush

San Francisco, CA

AI Infrastructure Engineer
No salary listedOn-siteDevOps / SRE

Builds and runs end-to-end ML inference infrastructure for generative AI image models across mobile and web platforms. Requires expertise in distributed systems, Kubernetes, Kafka, Redis, GPUs, and multi-cloud environments.

Palantir

Palantir

New York, NY

Site Reliability Operations Analyst - Commercial
No salary listedHybrid3+ YOEDevOps / SRE

Site Reliability Operations Analyst streamlines workflows, stabilizes projects, and acts as first responder for Palantir deployments to free engineers for technical work. Requires 3+ years project management, travel willingness, and strong judgment under pressure.

Collectly

Collectly

Belgrade, Serbia

Senior DevOps/DevSecOps Engineer
$85k+/yrRemote5+ YOEDevOps / SRE

Builds and secures scalable infrastructure, CI/CD pipelines, monitoring, and incident-response processes. The role requires DevOps automation experience, cloud expertise, Terraform and Ansible proficiency, and a bachelor's degree or equivalent experience.

Level AI

Level AI

Noida, India
Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will optimize Kubernetes and GPU infrastructure for cost, throughput, and reliability while enabling backend teams through tooling and instrumentation. The role requires 4–5 years of systems experience, backend development depth, Kubernetes expertise, GCP and Terraform fluency, and hybrid on-premises infrastructure experience.