Skip to content
1,066 jobs

Job results

Nuro

Nuro

Mountain View, CA

Staff/Senior Software Engineer, Offboard Infrastructure
$194k+/yrOn-site5+ YOEDevOps / SRE

Staff/Senior Software Engineer building data platforms, simulation systems, or technical infrastructure for autonomous driving technology. Requires 5+ years experience, strong Python/C++/Go skills, and expertise in distributed systems or related areas.

Palantir

Palantir

New York, NY
Forward Deployed Reliability Engineer
No salary listedHybridDevOps / SRE

Forward Deployed Reliability Engineer ensures stability of Palantir's mission-critical workflows by handling on-call incidents, automating solutions, and driving product improvements. Requires proficiency in Python, Java, SQL, and a technical background.

Composio

Composio

San Francisco, CA

Member of Technical Staff - Platform Eng
No salary listedOn-siteDevOps / SRE

Builds and scales platform infrastructure enabling AI agents to integrate with tools like GitHub and Salesforce. Evolves APIs, manages code runtimes, optimizes performance, and collaborates with product teams on backend distributed systems.

Cerebras Systems

Cerebras Systems

United States
AI Infrastructure Operations Engineer
No salary listedOn-siteDevOps / SRE

Entry-level SiteOps engineer supporting deployment, validation, monitoring, and first-line troubleshooting of AI clusters in data center environments. Requires a relevant engineering degree or equivalent experience, familiarity with server hardware, networking, and Linux, and readiness to work hands-on in data centers.

Stripe

Stripe

New York, NY

Staff Infrastructure Engineer, Privy
No salary listedOn-site8+ YOEDevOps / SRE

Architects and manages multi-tenant infrastructure handling billions of requests monthly, designs build/deploy systems, drives observability, and ensures high availability. Requires 8+ years experience with TypeScript/Python, Terraform, Cloudflare, and AWS.

Retell AI

Retell AI

San Francisco, CA

Staff Engineer, Platform & Systems
$225k+/yrOn-site9+ YOEDevOps / SRE

Staff engineer owns core platform and systems architecture, leads technical initiatives, makes high-leverage decisions for speed/reliability/scale, and partners with product/leadership. Requires 9+ years experience in production systems at staff/principal level.

Rad AI

Rad AI

San Francisco, CA

Software Engineer, Infrastructure (All Levels)
$160k+/yrOn-site4+ YOEDevOps / SRE

Designs, builds, and operates scalable cloud infrastructure on AWS with Kubernetes and serverless tech to support AI healthcare products. Requires 4+ years experience in cloud-native platforms, IaC, automation, and reliability practices for regulated environments.

Zoox

Zoox

Foster City, CA

Senior Systems Engineer
$187k+/yrOn-site8+ YOEDevOps / SRE

Senior Systems Engineer owns end-to-end feature domains across vehicle, autonomy, and operations for mobility-as-a-service. Leads cross-functional teams through full product lifecycle, creates requirements and MBSE models, requiring 8+ years experience and bachelor's in engineering/CS/physics.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, GPU Infrastructure - HPC
$230k+/yrOn-siteDevOps / SRE

Software engineer focused on reliability and uptime of OpenAI's GPU/HPC compute fleet through automation, monitoring tools, and performance optimization. Requires proficiency in Python/Go, Linux, networking, and data analysis skills.

Joyful Health

Joyful Health

New York, NY

Founding Infra Engineer
$150k+/yrHybrid5+ YOEDevOps / SRE

Founding Infra Engineer owns Kubernetes-based platform, CI/CD pipelines, observability, security, and PostgreSQL optimization for a healthcare financial OS. Requires 5+ years Infra/SRE experience; NYC-based hybrid role.

Sieve

Sieve

San Francisco, CA

Reliability Engineer
$150k+/yrOn-site3+ YOEDevOps / SRE

Designs and maintains reliable infrastructure for petabyte-scale video processing across multi-cloud environments. Owns incident response, observability (Prometheus, OpenTelemetry), security, and CI/CD tooling with 3+ years experience in scalable systems.

Postman

Postman

San Francisco, CA

Member of Technical Staff, AI Reliability & Monitoring Engineering Lead
$256k+/yrHybridDevOps / SRE

Lead AI reliability engineering for Postman's API and agentic systems, building monitoring, observability, and automation for high availability. Requires strong SRE/DevOps background in large-scale AI infrastructure and cloud platforms.

Clickhouse

Clickhouse

Remote

Cloud Database Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

Build and operate scalable, highly available cloud-native database infrastructure for ClickHouse’s serverless platform. The role requires 5+ years of distributed-systems experience, production programming in Go, C++, or Java, cloud infrastructure expertise, and operational on-call experience.

Maven Robotics

Maven Robotics

San Francisco, CA

Software Engineer - DevOps and MLOps
No salary listedOn-siteDevOps / SRE

Builds and maintains infrastructure for software development, ML models, and AI operations including CI/CD pipelines, cloud platforms, and automation tools. Requires BS/MS in CS, experience with containerization, IaC, and programming in Python/Go/Java.

Shield AI

Shield AI

San Diego, CA
Staff Engineer, Software Integration (R4483)
$150k+/yrOn-site7+ YOEDevOps / SRE

Integrates autonomy software stack for AI robotics platforms, including multi-agent systems, sensor processing, and hardware deployment across simulation, HIL, and flight environments. Requires 7+ years experience, Python/C++, CI/CD expertise, and strong systems integration skills.

Shield AI

Shield AI

San Diego, CA

Senior Engineer, Autonomy Systems Integration and Test (R4472)
$110k+/yrOn-site4+ YOEDevOps / SRE

Leads hands-on integration and testing of autonomy systems on small unmanned aircraft, debugging issues across hardware, software, and flight environments. Requires 4-7 years experience in robotics/aerospace testing, Python/C++ proficiency, and bachelor's in engineering or related field.

Crusoe

Crusoe

Denver, CO

Data Center Systems Engineer (MDC)
$149k+/yrOn-site6+ YOEDevOps / SRE

Leads integration of electrical, thermal, mechanical, and networking systems in modular data centers for high-density AI compute. Ensures compatibility with power sources, optimizes thermal management, and supports transition to liquid cooling with 6+ years systems engineering experience.

Databricks

Databricks

Bellevue, WA

Principal Engineer, Compute Fleet Management
$264k+/yrOn-siteDevOps / SRE

Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.

Databricks

Databricks

San Francisco, CA

Senior Software Engineer - Infrastructure and Tools
$166k+/yrHybrid5+ YOEDevOps / SRE

Build and extend scalable infrastructure for Databricks' data and AI platform, including multi-cloud systems and Kubernetes at massive scale. Requires 5+ years experience in Java/Scala/Go/C++/Python, distributed systems, and cloud technologies.

Atticus

Atticus

Remote

Lead Infrastructure Engineer
$170k+/yrRemote5+ YOEDevOps / SRE

Lead Infrastructure Engineer shapes Atticus's infrastructure roadmap, builds security/reliability foundations, and empowers product teams with self-service tools. Requires 5+ years infra/SRE experience, GCP/Terraform expertise, and broad technical knowledge across networking, observability, and CI/CD.

Astera

Astera

Site Reliability Engineer
$100k+/yrHybridDevOps / SRE

Owns digital infrastructure for AI research, managing compute access, auto-scaling, resource visibility, and reproducibility using Kubernetes and observability tools. Requires systems intuition, operational rigor, and pragmatism for experimental workloads.

Vercel

Vercel

San Francisco, CA
Software Engineer, Deployment Infrastructure
$180k+/yrHybrid6+ YOEDevOps / SRE

Owns and evolves Vercel's CI/CD infrastructure, designing scalable microservices, ensuring reliability for millions of daily builds, and leading end-to-end projects. Requires 6+ years experience with JavaScript/TypeScript, AWS, Terraform, and distributed systems.

Rad AI

Rad AI

United States

Staff Software Engineer, Infrastructure
$175k+/yrRemote8+ YOEDevOps / SRE

Designs and operates scalable cloud infrastructure on AWS, focusing on Kubernetes orchestration, reliability practices, and observability for AI healthcare products. Requires 8+ years experience with IaC, containerization, and cross-team leadership.

Datology AI

Datology AI

Redwood City, CA

Software Engineer, Cloud Infrastructure
$180k+/yrHybridDevOps / SRE

Designs, builds, and operates scalable multi-cloud infrastructure powering ML training, inference, and data curation pipelines. Collaborates with teams on AWS-focused systems using Kubernetes and IaC tools like Terraform.

Bretton AI

Bretton AI

San Francisco, CA

Software Engineer, Infrastructure
$168k+/yrOn-site8+ YOEDevOps / SRE

Owns and evolves Kubernetes-based infrastructure for secure, compliant AI deployments in financial services, including observability with Datadog, IaC with Terraform, and incident response. Requires 8+ years experience with Docker, K8s, AWS, and Python at scale.

Superblocks

Superblocks

New York, NY

Infrastructure Engineer / SRE
$175k+/yrOn-siteDevOps / SRE

Designs and operates large-scale infrastructure for secure, scalable AI agent runtimes, untrusted code execution, and multi-cloud deployments. Requires strong expertise in distributed systems, containers, Kubernetes, and security.

FlexAI

FlexAI

Bengaluru, India

Senior DevOps Engineer/SRE
No salary listedOn-site5+ YOEDevOps / SRE

Build and operate scalable infrastructure powering FlexAI’s AI and PaaS platform. The role focuses on Kubernetes, infrastructure as code, CI/CD, observability, incident response, and reliability practices, requiring 4+ years of DevOps, SRE, or infrastructure engineering experience.

EngFlow

EngFlow

New York, NY
Software Engineer - Build Systems, Compilers and Languages
No salary listedRemoteDevOps / SRE

Develops core features for cloud-based build systems and compilers, optimizing scalability and performance. Requires expertise in build tools like Bazel/CMake, Linux/cloud infrastructure, and languages like Java/C++/Rust.

Palantir

Palantir

New York, NY
Product Reliability Engineer - Defense
No salary listedHybridDevOps / SRE

Product Reliability Engineers own end-to-end service health for Palantir's critical platforms, combining on-call incident response with forward-looking work on observability, resilience, code improvements, and infrastructure migrations. Requires strong backend coding (Java), troubleshooting skills, and ownership in dynamic environments; US security clearance eligibility needed.

Poetic

Poetic

San Francisco, CA

Platform Engineer
$175k+/yrOn-siteDevOps / SRE

Owns and builds core parts of the Forge platform end-to-end, including IDE, compiler, runtime, and infra for AI-powered English programming language. Requires staff+ engineer experience building products 0-1 with deep technical craft.

Tigerdata

Tigerdata

Remote

Senior Platform Engineer - Postgres Specialist
No salary listedRemoteDevOps / SRE

Designs and implements cloud database features for Postgres platform, ensures operational excellence, performance, and observability for large-scale Postgres/Timescale instances on Kubernetes. Requires deep Postgres expertise, Golang, and experience managing stateful workloads at scale.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Platform Systems
$310k+/yrOn-siteDevOps / SRE

Designs and builds distributed failure detection, tracing, and observability systems for large-scale AI training jobs. Requires deep expertise in performance, distributed systems, hardware, networking, and low-level software engineering.

Harvey

Harvey

San Francisco, CA

Senior Software Engineer, Core Infrastructure
$200k+/yrHybrid4+ YOEDevOps / SRE

Designs, builds, and scales core infrastructure for Harvey's AI platform, focusing on multi-cloud systems (Azure, GCP), Kubernetes, and observability. Requires 4+ years in infrastructure engineering with strong IaC and distributed systems expertise.

Rippling

Rippling

San Francisco, CA

Staff Software Engineer - Platform
No salary listedHybrid8+ YOEDevOps / SRE

Staff Software Engineer on the Platform team leads technical roadmap, scales distributed systems for high QPS and throughput, prototypes AI-powered primitives, and owns core infrastructure to enable Rippling's product ecosystem. Requires 8+ years experience with deep platform expertise.

Runloop

Runloop

San Francisco, CA

Software Engineer, Infrastructure
No salary listedHybrid2+ YOEDevOps / SRE

Builds and scales core infrastructure using MicroVMs, Kubernetes, AWS, and GCP to support AI agent development sandboxes. Requires 2+ years experience in infrastructure engineering and systems programming.

Flux Auto

Flux Auto

San Francisco, CA

DevOps Engineer
No salary listedOn-site5+ YOEDevOps / SRE

DevOps Engineer builds and maintains full-stack systems for billing, authentication, onboarding, and integrations on Flux's AI hardware platform. Requires 5+ years SRE/DevOps experience, TypeScript, observability tools like Datadog/Sentry, and IaC with Pulumi across GCP/AWS/Firebase.

Modal

Modal

New York, NY
Member of Technical Staff - Reliability Engineering
$150k+/yrOn-site5+ YOEDevOps / SRE

Define and implement reliability systems for a growing AI cloud infrastructure platform, including architectural improvements, operational processes, monitoring, and incident response. Requires 5+ years production coding and 2+ years on-call experience with strong cloud skills.

Zoox

Zoox

Foster City, CA

Platform Engineer
$165k+/yrHybrid5+ YOEDevOps / SRE

Build and operate reliable infrastructure and testing services for autonomous-vehicle development. The role requires 5+ years supporting production services and SRE responsibilities, plus proficiency in Python or Golang and experience with automation, observability, CI/CD, and resilient infrastructure.

EliseAI

EliseAI

New York, NY
Senior DevOps Engineer
$230k+/yrOn-site3+ YOEDevOps / SRE

Owns infrastructure, deployment pipelines, monitoring, and security using AWS and IaC tools. Requires 3+ years DevOps experience, strong AWS and Terraform proficiency, and onsite work 4-5 days/week in NYC or SF.

CodeRabbit

CodeRabbit

San Francisco, CA
Velocity Solutions Engineer
$175k+/yrHybrid2+ YOEDevOps / SRE

Technical advisor partnering with sales to demo CodeRabbit's AI code review platform, develop custom solutions, and support pre-sales PoVs. Requires 2+ years customer-facing experience, cloud knowledge (AWS/GCP/Azure), and command line comfort.

Cohere

Cohere

San Francisco, CA
Site Reliability Engineer, Inference Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Operates and scales reliable infrastructure for serving Cohere’s language models, building Kubernetes automation, observability, resilience, and customized production deployments. Requires 5+ years running large-scale production infrastructure and experience with distributed systems, cloud platforms, Linux, GPU workloads, and high-performance server development.

Moment

Moment

New York, NY

Platform Engineer
$250k+/yrOn-siteDevOps / SRE

Builds and maintains shared libraries, SDKs, and AWS infrastructure (ECS, RDS, MSK) to enable faster, safer development. Architects golden paths, optimizes observability, and mentors engineers with deep Go backend and cloud expertise.

SentiLink

SentiLink

United States

Infrastructure Engineer (Experienced/Senior+ Levels)
$145k+/yrRemote4+ YOEDevOps / SRE

Senior Infrastructure Engineer builds and maintains scalable cloud infrastructure, CI/CD pipelines, and monitoring systems using AWS, Kubernetes, and Terraform. Collaborates with engineering teams to enhance reliability, efficiency, and security; requires 4+ years experience.

Together AI

Together AI

San Francisco, CA

Senior Software Engineer - Together Cloud Infrastructure
$160k+/yrRemote5+ YOEDevOps / SRE

Build and maintain highly available AI cloud infrastructure virtualizing ML hardware like GB200 GPUs and BlueField DPUs, enabling self-serve Kubernetes/Slurm clusters for internal and external customers. Requires 5+ years experience with distributed systems, backend development (Golang preferred), and cloud providers.

Together AI

Together AI

San Francisco, CA

Senior Developer Productivity Engineer
$150k+/yrOn-site5+ YOEDevOps / SRE

Owns systems and tooling to optimize developer workflows, CI/CD pipelines, and local environments for faster software delivery. Requires 5+ years in DevOps/SRE, proficiency in Python/Go/JS, and CI/CD expertise.

Retool

Retool

San Francisco, CA

Software Engineer, Core Infrastructure
$164k+/yrOn-siteDevOps / SRE

Builds and scales core cloud infrastructure for high availability, performance, and developer productivity. Collaborates across teams to evolve backend architecture, automate workflows, and maintain production systems using Kubernetes, Docker, Postgres, and Node.

Sesame

Sesame

San Francisco, CA
Backend Infrastructure Engineer
$175k+/yrOn-site5+ YOEDevOps / SRE

Builds developer infrastructure tools including CI/CD pipelines, build systems, and self-serve frameworks to enable fast, reliable engineering workflows. Requires 5+ years experience with Python/TypeScript, automation focus, and dev tooling expertise.

HappyRobot

HappyRobot

San Francisco, CA

Infrastructure Engineer
$200k+/yrOn-site3+ YOEDevOps / SRE

Infrastructure Engineer owns system stability, observability, and debugging at scale. Requires 3+ years experience with Go, Kubernetes, and tools like Datadog/Prometheus for production incident response.

Reducto

Reducto

San Francisco, CA

Senior Software Engineer, Platform
$150k+/yrOn-site5+ YOEDevOps / SRE

Senior Software Engineer on the Platform team builds and optimizes core APIs and pipelines for document parsing using LLMs, integrates cutting-edge models, reduces latency/costs, and collaborates with ML engineers. Requires 5+ years experience with 2+ years in production LLMs and Python expertise.

Tempo

Tempo

United States

Infrastructure Engineer, Blockchain (Remote - USA)
No salary listedRemoteDevOps / SRE

Builds and maintains blockchain infrastructure including validators, Kubernetes deployments, and observability systems to enable fast engineering velocity. Requires expertise in Rust/Go/Python, Terraform, Prometheus/Grafana, Linux, and Ethereum ecosystem.