Skip to content
1,067 jobs

Job results

Nectarsocial

Nectarsocial

Palo Alto, CA

Senior Site Reliability Engineer
$200k+/yrHybrid5+ YOEDevOps / SRE

Nectar Social is seeking a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of their production systems. This role involves defining SLOs, leading incident response, improving infrastructure performance, and partnering with engineering teams to embed reliability into system design.

Ironclad

Ironclad

San Francisco, CA
Senior Staff Site Reliability Engineer
$245k+/yrHybrid8+ YOEDevOps / SRE

Ironclad is seeking a Senior Staff Site Reliability Engineer to provide technical leadership and strategic direction for the SRE team, champion engineering excellence, and drive architectural resilience for their cloud platform.

Replit

Replit

Foster City, CA

Staff Infrastructure Engineer
$220k+/yrHybrid8+ YOEDevOps / SRE

As a Staff Infrastructure Engineer, you will ensure the reliability, scalability, and performance of Replit's infrastructure. You will drive automation, optimize performance, elevate developer experience, and mentor the engineering team on best practices for resilient systems.

Speakeasy

Speakeasy

San Francisco, CA

Platform Engineer - AI Control Plane
No salary listedOn-siteDevOps / SRE

As a Platform Engineer (SRE) focusing on the AI Control Plane, you will identify architectural changes, foster a culture of reliability, design operational processes, participate in on-call rotations, build monitoring systems, and debug production issues.

Build

Build

New York, NY
Senior Software Engineer (Core Infrastructure)
$170k+/yrOn-site5+ YOEDevOps / SRE

Owns the foundational infrastructure that orchestrates autonomous agents and long-running workflows in production. The role combines systems architecture, reliability engineering, end-to-end infrastructure ownership, and direct responsibility for security and compliance.

Convex

Convex

San Francisco, CA

Senior Software Engineer, Infra/Systems
$240k+/yrHybrid6+ YOEDevOps / SRE

Convex is seeking a Senior Software Engineer to design, build, and maintain their global cloud infrastructure. This role involves working on core systems, improving performance and reliability, and owning architectural decisions.

Reltio

Reltio

Lisbon, Portugal

Staff Engineer, Release Management
No salary listedRemote8+ YOEDevOps / SRE

Leads release management and cloud infrastructure initiatives for a highly available SaaS platform, mentoring the DevOps/Release team and improving automation, deployment, observability, security, and reliability. Requires 8+ years of enterprise SaaS development or operations experience and 6+ years with highly available cloud applications.

TetraScience

TetraScience

Boston, MA

Technology Lead, DevOps Engineering
No salary listedRemote7+ YOEDevOps / SRE

Hands-on technical lead owning cloud infrastructure, CI/CD pipelines, and deployment automation for a multi-tenant SaaS platform. Architect and build production systems using AWS, Terraform, CloudFormation, and Python in a GxP-regulated environment.

Flourish

Flourish

New York, NY

CloudOps Engineer
$152k+/yrOn-site3+ YOEDevOps / SRE

Design, build, and automate secure AWS cloud-native infrastructure with Kubernetes and Terraform. Enable dev teams with self-service platforms, CI/CD pipelines, and SRE best practices.

Chess.com

Chess.com

United States

Site Reliability Engineer
No salary listedRemote5+ YOEDevOps / SRE

Design and operate multi-regional infrastructure for a high-traffic global gaming platform, owning on-call, monitoring, automation, and hybrid cloud migration to ensure reliability at massive scale.

Skydio

Skydio

San Mateo, CA
Senior Software Engineer, Infrastructure
$170k+/yrHybrid4+ YOEDevOps / SRE

Senior Infrastructure Engineer responsible for re-architecting Kubernetes infrastructure, improving continuous deployment, and making code changes across the stack to support drone platform needs.

Skydio

Skydio

San Mateo, CA
Software Engineer - Infrastructure
$140k+/yrHybrid2+ YOEDevOps / SRE

Infrastructure engineer responsible for maintaining and scaling Kubernetes fleets, improving CI/CD, and making product-level code changes in Python or Go to support autonomous drone platform needs.

Order.co

Order.co

United States

Senior Site Reliability Engineer
No salary listedRemoteDevOps / SRE

Senior SRE responsible for building and operating reliable, scalable infrastructure on AWS with Kubernetes and Terraform. Focus on observability, incident response, automation, and mentoring engineers on SRE best practices.

Pinterest

Pinterest

San Francisco, CA

Staff Software Engineer, Observability
$177k+/yrRemote7+ YOEDevOps / SRE

Staff Software Engineer building and scaling Pinterest's observability platform (metrics, logs, traces) for massive distributed systems. Requires 7+ years distributed systems experience, strong data engineering skills, and expertise with modern observability tools.

Drata

Drata

San Francisco, CA

Senior Platform Engineer, Interoperability
$151k+/yrHybrid5+ YOEDevOps / SRE

Lead development of scalable platform systems and infrastructure tools that enable internal and external developers to build faster, more reliable applications. Requires 5+ years of software engineering experience with 3+ years in Node.js.

Drata

Drata

San Francisco, CA

Staff Platform Engineer, Interoperability
$200k+/yrHybrid10+ YOEDevOps / SRE

Staff Platform Engineer building developer tooling, CI/CD automation, and scalable web applications using NodeJS, React, and AWS. Requires 10+ years experience and expertise in Temporal, Terraform, PostgreSQL, and Snowflake.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Full-Stack — Developer Experience
$185k+/yrOn-site5+ YOEDevOps / SRE

Build and operate scalable CI and Bazel-based build systems that accelerate engineering velocity and reliability for OpenAI's products and infrastructure.

Zoox

Zoox

Foster City, CA

Senior Manager, Network Engineering & Infrastructure
$272k+/yrHybrid10+ YOEDevOps / SRE

Lead and mentor a network engineering team responsible for designing, deploying, and operating multi-site enterprise network infrastructure across data centers, cloud, offices, and vehicle facilities. Requires 10+ years of network experience with 5+ years in senior leadership.

Abridge

Abridge

San Francisco, CA
Staff Platform Engineer
$228k+/yrHybrid10+ YOEDevOps / SRE

Lead platform architecture and operations for a multi-tenant, multi-cloud infrastructure serving a fast-growing AI healthcare company. Design and scale Kubernetes platforms, Terraform modules, CI/CD pipelines, and observability tooling while driving security, reliability, and developer velocity.

Earnin

Earnin

Mountain View, CA

Staff Site Reliability Engineer
$252k+/yrHybrid7+ YOEDevOps / SRE

Lead EarnIn's AI-first reliability engineering strategy. Define SLOs/SLIs, build AI agents for incident response and on-call automation, and partner with engineering teams to embed AI-assisted operations across production systems on AWS.

Render

Render

United States
Software Engineer, Network Infrastructure
$204k+/yrRemote6+ YOEDevOps / SRE

Design, build, and operate Render's core networking stack across data centers and clouds, focusing on Kubernetes and Linux internals, traffic routing, and hybrid connectivity at scale.

ModernFi

ModernFi

New York, NY

Senior Software Engineer - Platform & Infrastructure
$160k+/yrRemote5+ YOEDevOps / SRE

Founding Senior Platform Engineer building and owning AWS cloud infrastructure, reliability, observability, security/compliance (SOC 2, Vanta), and release tooling for a fintech platform serving banks and credit unions.

Julius

Julius

San Francisco, CA

Software Engineer - Infrastructure (Mid - Senior)
$150k+/yrOn-site3+ YOEDevOps / SRE

Build and scale secure multi-tenant container infrastructure and code sandboxes on AWS/GCP for an applied AI coding platform. Own reliability, observability, and performance for 500k+ containers per month.

Anthropic

Anthropic

San Francisco, CA
Performance Engineer, Inference Systems
$350k+/yrHybridDevOps / SRE

Performance engineer focused on cross-layer investigations of Anthropic's inference fleet for Claude, optimizing throughput, latency, reliability, and correctness while building observability and partnering with kernel and serving teams.

Twenty

Twenty

Arlington, VA
DevOps Engineer
No salary listedOn-siteDevOps / SRE

Design, build, and operate AWS infrastructure with Terraform, CI/CD pipelines, observability, and security for a national security technology platform.

Lightning AI

Lightning AI

New York, NY
Infrastructure Operations Engineer
$160k+/yrHybrid8+ YOEDevOps / SRE

Design, build, and maintain infrastructure platforms using Linux, AWS, Kubernetes, Terraform, and Ansible to support internal and customer-facing services. Participate in on-call rotations and collaborate across engineering and operations teams.

Metriport

Metriport

San Francisco, CA

Software Engineer, Internal Tools
$120k+/yrOn-site3+ YOEDevOps / SRE

Build and maintain internal tooling and automations to improve operational efficiency across engineering, operations, and business teams. Own end-to-end integrations, developer experience improvements, and security-compliant tool adoption.

April

April

Tel Aviv, Israel

SRE / DevOps Lead
No salary listedOn-site7+ YOEDevOps / SRE

Leads a small SRE team while owning the reliability, security, scalability, and delivery infrastructure of a GCP-based platform. The role combines hands-on cloud operations with incident management, compliance support, observability, and corporate IT leadership.

Reducto

Reducto

San Francisco, CA

Infrastructure Engineer
$150k+/yrOn-site5+ YOEDevOps / SRE

Founding Infrastructure Engineer to architect and scale resilient systems for AI/ML workloads, implement monitoring/observability, and automate infrastructure. Requires 5+ years production experience, Python, Kubernetes, and strong reliability focus.

Snowflake

Snowflake

Staff Software Engineer - Container Platform
$236k+/yrHybridDevOps / SRE

Staff-level engineer owning design and operations of Snowflake's large-scale Kubernetes container platform across AWS, Azure, and GCP. Focus on reliability, automation, and developer experience for internal engineering teams.

Fieldguide

Fieldguide

San Francisco, CA

Staff Software Engineer, App Platform
$210k+/yrRemote10+ YOEDevOps / SRE

Lead design and evolution of core platform services, APIs, and shared primitives that power every product surface and AI agent. Drive technical standards and architecture across SaaS, enterprise, and government environments while mentoring engineers.

OpenAI

OpenAI

San Francisco, CA

Tech Lead, Deployment & Operations — Custom Infrastructure
$342k+/yrHybrid8+ YOEDevOps / SRE

Lead deployment and operations for OpenAI’s custom silicon and systems into data center environments. Drive hardware bring-up, validation, production deployment, and fleet reliability at scale while leading a technical team.

Shield AI

Shield AI

Dallas, TX

Staff Engineer, XBAT DevOps (R4542) (Dallas, TX)
$158k+/yrOn-site7+ YOEDevOps / SRE

Own CI/CD and automation infrastructure for autonomy, embedded, and ground software, improving developer velocity and scaling simulation/HIL testing. Requires 7+ years experience with Azure DevOps, Docker, Kubernetes, and strong scripting skills.

Nooks

Nooks

San Francisco, CA

Senior Software Engineer, Core Infrastructure
$215k+/yrHybrid5+ YOEDevOps / SRE

Senior engineer on the Core Infrastructure team responsible for scaling data layers, observability, and developer tooling to support rapid multi-product growth at Nooks. Requires 5+ years experience scaling systems 10x+, strong distributed systems or infra background, and willingness to be in-office in San Francisco 3+ days/week.

Databricks

Databricks

New York, NY
Staff Software Engineer - AI Research Infrastructure
$199k+/yrOn-site5+ YOEDevOps / SRE

Build and operate the large-scale training and inference infrastructure that powers Databricks AI Research, enabling researchers to run experiments across thousands of GPUs. Partner with ML scientists and platform teams to deliver reliable, high-performance orchestration and tooling.

Together AI

Together AI

San Francisco, CA
AI Infrastructure Engineer
$190k+/yrOn-site5+ YOEDevOps / SRE

Builds and maintains AI infrastructure using Ansible, Terraform, and Kubernetes, ensuring scalability, reliability, and high availability. Handles on-call incident response, monitoring, debugging, and infrastructure growth planning. Requires 5+ years experience and CS bachelor's.

Shield AI

Shield AI

Dallas, TX

Staff Engineer, Field Quality (R4958)
$120k+/yrOn-site5+ YOEDevOps / SRE

Staff Engineer investigates field issues in autonomous hardware systems, drives root cause analysis, implements corrective actions, and improves reliability through cross-functional collaboration. Requires 5+ years in quality engineering for complex hardware in aerospace/defense/robotics.

Datadog

Datadog

Boston, MA
People Systems Developer
$145k+/yrHybrid3+ YOEDevOps / SRE

Builds and deploys AI-native workflows for HR systems like talent acquisition, onboarding, and performance management. Integrates LLMs and APIs into tools like Workday and Greenhouse; requires 3-5 years software engineering with AI/automation experience.

Wiz

Wiz

Tel Aviv, Israel

DevOps Engineer
No salary listedOn-site4+ YOEDevOps / SRE

The DevOps Engineer will design, scale, automate, and operate production-grade cloud infrastructure and distributed systems. The role requires 4+ years of DevOps experience, strong Docker and Kubernetes expertise, CI/CD knowledge, and Python, Bash, or Go skills.

Lovable

Lovable

Stockholm, Sweden

Staff / Principal Software Engineer, Platform
No salary listedOn-site10+ YOEDevOps / SRE

Build and operate the platform infrastructure and developer tooling powering Lovable’s AI product, including sandbox runtimes, schedulers, observability, networking, and reliability systems. The role requires 10+ years in platform, infrastructure, or developer experience engineering and onsite work in Stockholm.

Illumio

Illumio

Sunnyvale, CA

Sr. Site Reliability Engineer
$170k+/yrOn-site5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for ensuring reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure. Lead monitoring, incident response, on-call support, automation, and continuous improvement initiatives in a cybersecurity SaaS environment.

Odys Aviation

Odys Aviation

Mission Operations Specialist (Pilot)
No salary listedOn-siteDevOps / SRE

Serves as primary Remote Pilot in Command for eVTOL UAV flight tests, managing GCS operations, securing regulatory authorizations using FAA/EASA frameworks, and collaborating with engineering on test plans and telemetry analysis. Requires FAA Part 107 certification, UAV piloting experience, and bachelor's in aerospace or related field.

Zello

Zello

Austin, TX

Senior Site Reliability Engineer, Database Infrastructure
No salary listedHybrid7+ YOEDevOps / SRE

Owns the reliability, performance, and availability of production database infrastructure while contributing to observability, incident response, cloud modernization, and automation. Requires 7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, including substantial production database ownership.

ClickUp

ClickUp

Poland
Senior Database Reliability Engineer
No salary listedRemote7+ YOEDevOps / SRE

The Senior Database Reliability Engineer will operate and improve large-scale PostgreSQL environments in AWS, focusing on availability, performance, security, backups, recovery, and incident response. The role requires 7+ years of database administration or engineering experience, strong Linux and SQL skills, and production cloud database expertise.

Lovable

Lovable

Stockholm, Sweden

Software Engineer, Platform
No salary listedOn-site6+ YOEDevOps / SRE

Operates Lovable’s app runtime, deployment pipeline, and managed infrastructure services, with responsibility for reliability, vendor migrations, custom domains, and operational excellence. The role seeks an experienced platform, SRE, or infrastructure engineer with distributed-systems expertise and an uptime-focused mindset.

Supabase

Supabase

Remote

Software Engineer: IaC Platform Experience
No salary listedRemote5+ YOEDevOps / SRE

Owns and improves the Go-based Terraform provider for Supabase's developer platform, focusing on reliability, lifecycle management, schema evolution, and user migrations. Requires 5+ years experience with Go, deep Terraform expertise, and strong testing/CI/CD skills.

Supabase

Supabase

Remote

Edge Functions Engineer
No salary listedRemote5+ YOEDevOps / SRE

Develops and optimizes Supabase Edge Runtime, a Rust-based Deno host for global edge TypeScript functions. Evolves infrastructure for low-latency compute, integrates with Supabase stack, and improves developer tools. Requires 5+ years backend/systems experience with Rust, TypeScript, and scalable infra.

Scale AI

Scale AI

San Francisco, CA

Finance Systems & Automations Manager
$166k+/yrOn-site5+ YOEDevOps / SRE

Designs and builds integrations, automations, and AI agent workflows to streamline Finance, Accounting, People Ops, and Recruiting processes. Requires 5+ years in system integrations, iPaaS platforms like Workato, and experience with HR/Finance tools like NetSuite and Greenhouse.

Ai2

Ai2

Seattle, WA

Infrastructure Engineer
$100k+/yrOn-site5+ YOEDevOps / SRE

Builds and maintains automated IT infrastructure pipelines using Python, Terraform, and cloud providers to support company operations. Requires 5+ years experience with focus on automation, strong coding, and collaboration skills.

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer, Managed Orchestration (Managed Kubernetes)
$220k+/yrOn-site8+ YOEDevOps / SRE

Staff Software Engineer designs, builds, and scales managed Kubernetes and AI training clusters, focusing on reliability, performance, and orchestration using Go, Terraform, and GCP. Oversees architecture, CI/CD pipelines, and critical infrastructure projects requiring 8+ years experience.