Skip to content
1,067 jobs

Job results

Alpaca

Alpaca

United States

Senior AI Platform Engineer
No salary listedRemote8+ YOEDevOps / SRE

Builds and maintains AI platform infrastructure for agentic systems, including connectors, execution environments, governance, and self-service tools to enable safe, scalable AI use across engineering and business teams. Requires 8+ years experience with LLM agents, GCP, and cloud-native tech.

Stellar Cyber

Stellar Cyber

Spain
Staff SRE Engineer
No salary listedRemote7+ YOEDevOps / SRE

The Staff SRE Engineer will drive reliability, scalability, observability, and operational efficiency across highly available cloud and distributed systems. The role requires 5+ years of SRE, DevOps, or platform experience, advanced Kubernetes expertise, strong automation skills, and leadership in incident management.

Baseten

Baseten

San Francisco, CA
Forward Deployed SRE
$135k+/yrHybridDevOps / SRE

Site Reliability Engineer owns reliability of multi-cloud Kubernetes infrastructure for AI/ML platform, builds observability tooling as code, automates mitigations, leads incident response, and defines SLOs/SLIs. Requires extensive Kubernetes and observability experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Frontier Systems
$250k+/yrOn-site7+ YOEDevOps / SRE

Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.

Cerebras Systems

Cerebras Systems

Toronto, Canada

DevOps Engineer - New Grad 2026
No salary listedHybridDevOps / SRE

Build and maintain scalable infrastructure, CI/CD workflows, and cloud or datacenter automation for Cerebras’s AI software stack. The role requires a current university student or new graduate with software development experience and proficiency in Python, shell scripting, containers, Jenkins, and cloud platforms.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Productivity - Inference Runtime
$230k+/yrOn-siteDevOps / SRE

Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.

Railway

Railway

Remote

Senior Infra Engineer: Baremetal Orchestration
No salary listedRemoteDevOps / SRE

Builds and maintains bare metal provisioning, orchestration engine, and internal tools for Railway's infrastructure platform. Optimizes fleet efficiency and develops resilient services using Golang/Rust, Ansible, and Terraform for distributed systems.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Member of Technical Staff (Software Engineer)
$170k+/yrRemoteDevOps / SRE

Develops and optimizes Kubernetes-based infrastructure for high-performance AI inference services, including deployment, scaling, debugging, and integration with ML workflows. Requires Master's in CS and 1+ year experience with Docker, Kubernetes, Python, and related tools.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Sr. Member of Technical Staff
$230k+/yrHybridDevOps / SRE

Develops resilient, high-availability software for AI inference on AWS, including deployment workflows, container orchestration with Docker/Kubernetes, monitoring, and debugging. Requires Master's in CS and 18 months experience with AWS services, IaC tools, and Python.

Retool

Retool

San Francisco, CA

Software Engineer, Governance
No salary listedHybrid2+ YOEDevOps / SRE

Builds governance infrastructure for enterprise-scale Retool platform, including data access controls, admin tools, security monitoring, and organizational management systems. Requires 2-8 years full-stack experience with backend systems design and customer-focused problem-solving.

xAI

xAI

Palo Alto, CA
Member Of Technical Staff - Cloud Infrastructure
$180k+/yrOn-site5+ YOEDevOps / SRE

Designs, builds, and operates secure, scalable infrastructure including Kubernetes clusters and GPU hardware for large-scale AI workloads in classified US government environments. Requires 5+ years experience, Top Secret clearance, and expertise in IaC tools like Terraform and Ansible.

Zoo

Zoo

Los Angeles, CA

Senior Systems Software Engineer
$145k+/yrRemoteDevOps / SRE

Builds and maintains core backend systems, infrastructure, and automation for platform reliability and scalability. Owns troubleshooting across stack, integrations with external services, and observability. Requires strong Rust, systems engineering, and distributed systems experience.

Kraken

Kraken

United Kingdom
Senior AI Compute Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

This senior engineer will operate and optimize GPU and accelerator infrastructure for AI training, inference, evaluation, and experimentation. The role requires 5+ years of infrastructure experience, production GPU cluster operations, strong systems fundamentals, and expertise in serving, observability, reliability, and compute-cost optimization.

Octus

Octus

New York, NY

DevOps Engineer
$165k+/yrOn-site5+ YOEDevOps / SRE

Design, implement, and maintain cloud infrastructure and CI/CD pipelines. Collaborate with developers, SRE, and Security to ensure system reliability, scalability, and security.

Vesta

Vesta

United States

Software Engineer - Infrastructure
$200k+/yrRemoteDevOps / SRE

Builds and scales reliable cloud infrastructure, deployment systems, observability, and developer tooling to support mortgage market operations. Requires experience with strongly typed languages, PostgreSQL, Kubernetes, and major cloud providers.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Core Network Engineering
$230k+/yrOn-siteDevOps / SRE

Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.

Turquoise Health

Turquoise Health

Remote

Platform Operations Engineer
$153k+/yrRemote3+ YOEDevOps / SRE

Builds and scales platform infrastructure on AWS EKS with GitOps via ArgoCD, manages CI/CD with GitHub Actions, drives observability using Datadog/Sentry/CloudWatch, and ensures reliability through SLOs and incident response. Requires 3+ years SRE/DevOps experience and Kubernetes expertise.

Everlaw

Everlaw

Oakland, CA

Senior Software Engineer, Automation & Developer Tools
$131k+/yrHybrid5+ YOEDevOps / SRE

Builds and maintains automated testing infrastructure, CI/CD pipelines, and developer tools using Cypress and GitHub Actions/CircleCI. Mentors QA engineers and drives testing best practices in a legal SaaS platform. Requires 5+ years Cypress experience and expertise in major languages.

Tigerdata

Tigerdata

Spain
Senior Platform Engineer
No salary listedRemote3+ YOEDevOps / SRE

Builds and maintains Kubernetes-based infrastructure for managed TimescaleDB cloud services, develops Go microservices and operators, automates database operations, and ensures platform scalability and reliability. Requires 3+ years experience with Go, Kubernetes, and PostgreSQL.

OpenAI

OpenAI

San Francisco, CA

Networking Operating System Firmware Engineer
$266k+/yrHybridDevOps / SRE

Develops and maintains custom networking operating system firmware for AI supercomputers, integrating Linux kernel, switch ASICs, and control-plane services. Requires deep expertise in SONiC, SAI, routing protocols, and platform bring-up across hardware and software boundaries.

Stellar Cyber

Stellar Cyber

Spain
Senior SRE Engineer
No salary listedRemote5+ YOEDevOps / SRE

Senior SRE responsible for operating and improving highly available cloud platforms, distributed data systems, observability, incident response, and deployment automation. Requires 5+ years of SRE, DevOps, or platform engineering experience, advanced Kubernetes expertise, cloud proficiency, and strong Python and Bash skills.

Assembled

Assembled

New York, NY

Software Engineer - Platform
$135k+/yrOn-site5+ YOEDevOps / SRE

Build scalable infrastructure, integrations, and data platforms powering workforce management and AI agent products at enterprise scale. Requires 5+ years in backend/platform systems, with expertise in AWS, Kubernetes, Go/Python, and datastores like Postgres and Snowflake.

Zoox

Zoox

Foster City, CA

Senior Software Integration Engineer
$225k+/yrHybrid7+ YOEDevOps / SRE

Leads end-to-end integration of agentic platform tools with clients, backends, and cloud ops on Kubernetes. Requires 7+ years experience, strong Python, HTTP/auth expertise, and cross-functional collaboration for reliable tool contracts and releases.

NODA AI

NODA AI

Austin, TX

Senior Platform Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Platform Engineer architects and maintains infrastructure for autonomous vehicle orchestration platform, building CI/CD pipelines, managing AWS/on-premises/edge deployments, and ensuring DoD security compliance. Requires 5+ years in DevOps with Docker, AWS, and IaC expertise.

Crusoe

Crusoe

San Francisco, CA
Senior Production Engineer, Operational Excellence
$172k+/yrOn-site5+ YOEDevOps / SRE

Senior Production Engineer ensures reliability, scalability, and performance of GPU cloud infrastructure powering AI workloads. Drives observability, incident response, automation, and operational improvements in large-scale distributed systems.

Anthropic

Anthropic

San Francisco, CA
Staff+ Software Engineer, Developer Productivity
$405k+/yrHybridDevOps / SRE

Leads technical strategy and builds scalable developer infrastructure including build systems, CI/CD pipelines, and tooling for large monorepo environments. Requires 3+ years leading complex projects, proficiency in Python/Rust/Go, and experience with container orchestration.

Idme

Idme

Mountain View, CA

Staff Site Reliability Engineer
$218k+/yrOn-site10+ YOEDevOps / SRE

Leads infrastructure transformation from monoliths to scalable microservices at massive scale, architects observability/CI/CD systems, unifies complex stacks, and mentors engineers. Requires 10+ years coding internal tools, 5+ years cloud (GCP/AWS), Bachelor's in CS.

Phylo

Phylo

South San Francisco, CA
Member of Technical Staff - Infrastructure
CA$150k+/yrHybrid2+ YOEDevOps / SRE

Build and operate Kubernetes-based infrastructure for secure, reliable enterprise AI deployments across cloud and customer-managed environments. The role requires production Kubernetes experience, cloud infrastructure expertise, infrastructure as code, networking, security, and deployment automation.

Sesame

Sesame

San Francisco, CA
SWE - Backend Infrastructure Engineer
$175k+/yrOn-site3+ YOEDevOps / SRE

Builds and scales core infrastructure including ML training/serving, Kubernetes clusters, and low-latency voice/audio pipelines. Requires 3+ years in infrastructure/ML systems, hands-on reliability engineering, and Kubernetes expertise.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Engineer, Infrastructure, Training Systems
$350k+/yrOn-siteDevOps / SRE

Designs and optimizes distributed training systems scaling across thousands of GPUs for large AI models. Requires strong systems engineering, PyTorch/JAX expertise, and collaborative mindset to boost research productivity.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Engineer, Infrastructure, Numerics
$350k+/yrOn-siteDevOps / SRE

Designs and optimizes distributed training infrastructure for large-scale LLMs, focusing on low-precision numerics, kernel optimizations, and communication frameworks to enable stable, scalable trillion-parameter model training. Requires strong systems engineering, deep learning frameworks knowledge, and collaborative research mindset.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Engineer, Infrastructure, Kernels
$350k+/yrOn-siteDevOps / SRE

Designs and optimizes high-performance ML kernels (CUDA, CuTe, Triton) for large-scale LLM training, focusing on GPU efficiency, low-precision formats, and distributed compute. Collaborates with researchers to bridge algorithms and hardware.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Engineer, Infrastructure, Inference
$350k+/yrOn-siteDevOps / SRE

Designs, optimizes, and scales infrastructure for high-performance AI model inference, focusing on latency, throughput, efficiency, and reliability. Collaborates with researchers to enable production deployment of large-scale models using deep learning frameworks and distributed systems.

Standard Template Labs

Standard Template Labs

New York, NY

Senior Platform Software Engineer
$160k+/yrOn-site5+ YOEDevOps / SRE

Build and operate internal developer platforms, cloud infrastructure, CI/CD systems, and observability tooling that improve engineering productivity and production reliability. The role requires hands-on experience with AWS, Kubernetes, distributed systems, and infrastructure automation.

Nuro

Nuro

Mountain View, CA

Software Reliability Engineer
$146k+/yrOn-siteDevOps / SRE

Builds and operates resilient systems for autonomous vehicle fleet reliability, including pipelines for signal analysis, automated triage tools, internal workflows, and leading investigations. Requires production software experience and strong debugging skills in Python, Go, Bash, C++.

Twenty

Twenty

Arlington, VA

Forward Deployed Site Reliability Engineer (TS/SCI Required)
No salary listedOn-site5+ YOEDevOps / SRE

Forward Deployed Site Reliability Engineer responsible for owning reliability, observability, incident response, and deployments of a mission-critical platform in a restricted, air-gapped government AWS environment. Requires 5+ years SRE/production ops experience, strong Linux/Docker/Terraform skills, LGTM stack proficiency, and active TS/SCI clearance.

Ai2

Ai2

Seattle, WA

Senior Software Engineer, Platform
$126k+/yrOn-site8+ YOEDevOps / SRE

Builds foundational platform architecture for AI research agents, including SDKs, APIs, execution frameworks, and benchmarking infrastructure to enable researchers to develop intelligent systems over scholarly literature. Requires strong Python skills, 8+ years experience, cloud infrastructure, and AI integration expertise.

OpenAI

OpenAI

San Francisco, CA

Performance & Systems Engineer, Codex
$295k+/yrHybridDevOps / SRE

Optimizes performance across Codex AI system's stack including LLM inference, cloud orchestration, and agent behavior to reduce latency and costs. Collaborates with researchers and engineers on high-impact improvements in a high-ownership role.

Harvey

Harvey

San Francisco, CA

Staff Software Engineer, Developer Experience (DevEx)
$231k+/yrHybrid7+ YOEDevOps / SRE

Builds and scales developer platforms, CI/CD systems, testing infrastructure, and AI integrations to boost engineering velocity and reliability at an AI-native company. Requires 7+ years backend experience, leadership, and tools like Python, Kubernetes, and Terraform.

Railway

Railway

Remote

Senior Infra Engineer: Observability
No salary listedRemoteDevOps / SRE

Build high-scale observability pipelines and alerting engines handling 1M+ RPS for logs/metrics, develop Golang/Rust gRPC services and APIs, and manage immutable infrastructure with Terraform/Ansible in a distributed systems environment.

Lightning AI

Lightning AI

New York, NY
Infrastructure Engineer (GPU & Compute)
$180k+/yrRemote5+ YOEDevOps / SRE

Owns GPU diagnostics, validation workflows, and automation for bare-metal infrastructure supporting AI/ML workloads. Requires 5+ years in systems engineering with strong Linux, Python, and NVIDIA tools expertise.

Lightning AI

Lightning AI

New York, NY
Infrastructure Engineer (Storage)
$180k+/yrRemote5+ YOEDevOps / SRE

Operate and scale distributed storage systems like VAST and Ceph for AI/ML workloads, build Python automation tools, manage Linux bare-metal systems, and collaborate on infrastructure optimizations. Requires 5+ years in infrastructure engineering with storage expertise.

AKASA

AKASA

San Francisco, CA

Sr. Software Engineer, DevOps
$180k+/yrHybrid5+ YOEDevOps / SRE

Builds and scales reliable infrastructure for SaaS applications using Kubernetes, Terraform, and GitHub CI/CD. Focuses on observability with Grafana/Prometheus, automation to reduce toil, production troubleshooting, and cross-team collaboration. Requires 5+ years Python experience.

Palantir

Palantir

Warsaw, Poland
Edge Infrastructure Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Build, deploy, and operate physical and cloud infrastructure for Palantir’s air-gapped and edge environments. The role requires 5+ years managing large-scale systems, hands-on Linux, Kubernetes, networking, hardware, and distributed-systems experience, plus on-site availability in Warsaw.

Palantir

Palantir

Paris, France

Edge Infrastructure Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Operates and scales reliable edge and bare-metal infrastructure for Palantir’s production platform across isolated, on-premises, and cloud environments. The role requires at least five years managing large-scale systems, strong Linux, Kubernetes, networking, and server-hardware expertise, plus French proficiency and significant travel.

Databricks

Databricks

New York, NY
Staff Software Engineer - AI Research Infrastructure
$190k+/yrOn-site5+ YOEDevOps / SRE

Builds and operates research infrastructure for large-scale AI model training and inference across GPU fleets. Partners with scientists and engineers to create scheduling, orchestration, and dev tooling for efficient experimentation. Requires 5+ years in distributed systems and systems programming.

Brave

Brave

United States

Senior DevEx/CI (desktop/mobile) Engineer
No salary listedRemote15+ YOEDevOps / SRE

Builds and supports automation, CI/CD processes for desktop/mobile apps across 1000+ repositories in multiple languages and platforms to scale to 200M users. Requires 15+ years experience in build/release engineering, OS expertise, and tools like TeamCity, Jenkins, AWS.

Ema

Ema

San Francisco, CA

Software Engineer, DevOps
$135k+/yrHybrid3+ YOEDevOps / SRE

Designs and builds scalable infrastructure for AI products, focusing on cloud platforms, Kubernetes orchestration, CI/CD pipelines, and observability. Requires 3+ years in infrastructure engineering and bachelor's/master's in CS.

Anthropic

Anthropic

New York, NY
Incident Response Manager - Product & Engineering
$290k+/yrHybrid5+ YOEDevOps / SRE

Leads incident response operations for product and engineering, serving as on-call commander to coordinate cross-functional teams, manage communications, and improve processes during high-stakes incidents. Requires 5+ years in incident management with technical depth in infrastructure and cloud systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Productivity - Model Performance
$230k+/yrOn-siteDevOps / SRE

Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.