Skip to content
1,067 jobs

Job results

Reltio

Reltio

Bengaluru, India

Senior Manager Engineering - SRE
No salary listedRemote12+ YOEDevOps / SRE

Leads a global 24×7 DevOps and Customer Enablement organization supporting highly available SaaS operations. The role requires 12+ years in DevOps, SRE, cloud operations, or production engineering, including substantial experience managing engineering teams and driving reliability, incident response, automation, and customer outcomes.

OnePay

OnePay

Bengaluru, India

Senior Platform Engineer
No salary listedRemote5+ YOEDevOps / SRE

Senior platform engineer responsible for shipping and operating scalable infrastructure and services supporting financial products. Requires at least five years building non-trivial products or services, with backend experience and strong collaboration, ownership, and communication skills.

Crusoe

Crusoe

San Francisco, CA

Principal Engineer, CAPE
$285k+/yrOn-site10+ YOEDevOps / SRE

Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.

Weave

Weave

Lehi, UT

Staff SIP Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Staff SIP Engineer responsible for growing, optimizing, and troubleshooting a distributed cloud-based VoIP platform built on Kamailio, FreeSWITCH, and Kubernetes. Requires deep expertise in SIP/RTP protocols, VoIP troubleshooting, and DevOps practices to ensure high reliability and call quality.

Sprinter Health

Sprinter Health

San Francisco, CA

AI Enablement Engineer
$180k+/yrHybrid7+ YOEDevOps / SRE

Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.

Crusoe

Crusoe

San Francisco, CA

Senior Performance Engineer
$170k+/yrOn-site5+ YOEDevOps / SRE

Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.

xAI

xAI

Palo Alto, CA

Software Engineer
$180k+/yrOn-site5+ YOEDevOps / SRE

Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.

Okta

Okta

Washington, DC

Staff TDI Site Reliability Engineer, Okta Federal
$174k+/yrHybrid7+ YOEDevOps / SRE

Staff SRE building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring in isolated environments to support national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI clearance.

Intercom

Intercom

Dublin, Ireland

Senior Engineer, Infrastructure Platform
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate scalable infrastructure platforms that improve reliability, developer velocity, security, and operational efficiency across R&D. The role requires senior-level experience with cloud-based distributed systems, automation, incident response, and AI-assisted engineering.

Lightning AI

Lightning AI

New York, NY
Senior Network Engineer
$170k+/yrOn-site10+ YOEDevOps / SRE

Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.

Fluidstack

Fluidstack

San Francisco, CA
Software Engineer, GPU Infrastructure
$175k+/yrOn-site5+ YOEDevOps / SRE

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.

Fluidstack

Fluidstack

San Francisco, CA
Software Engineer, Compute
$208k+/yrOn-site5+ YOEDevOps / SRE

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

Fluidstack

Fluidstack

San Francisco, CA
Site Reliability Engineer, Compute
$208k+/yrOn-site5+ YOEDevOps / SRE

Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.

Chess.com

Chess.com

United States

Senior Elasticsearch Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.

Encord

Encord

San Francisco, CA

DevOps Engineer
$150k+/yrOn-site4+ YOEDevOps / SRE

DevOps Engineer embedded in platform teams to build and operate scalable AI infrastructure on GCP/AWS. Own CI/CD, Kubernetes, observability, reliability (SLIs/SLOs), automation, and performance at petabyte scale. Requires 4-5 years production DevOps/SRE experience.

Crusoe

Crusoe

Tel Aviv, IL

Manager, Production Engineering
No salary listedOn-site8+ YOEDevOps / SRE

Leads the founding Tel Aviv Production Engineering team, combining people management with hands-on reliability engineering, incident response, automation, and firmware optimization. Requires 8+ years in infrastructure, SRE, or production engineering and 2+ years of direct engineering leadership.

Arena

Arena

San Francisco, CA

Infrastructure Engineer, TL
No salary listedHybrid4+ YOEDevOps / SRE

Infrastructure Engineer building low-latency, high-reliability APIs, streaming gateways, and observability for Arena's real-world AI model evaluation platform. Requires 4+ years backend/distributed systems experience with Go/Rust, LLM APIs, and cloud infra (K8s/Terraform).

Navan

Navan

Tel Aviv, Israel

Senior SRE AI Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Senior SRE AI Engineer responsible for operating reliable production infrastructure and AI-powered applications, with ownership of automation, observability, incident response, and provider integrations. Requires 5+ years of SRE or infrastructure engineering experience and hands-on cloud, platform, and AI operations expertise.

Fluidstack

Fluidstack

Austin, TX
Software Engineering, Commissioning Automation
$269k+/yrOn-site5+ YOEDevOps / SRE

Build and ship software that automates data center commissioning, including test orchestration, data capture, pass-fail analysis, and live reporting by integrating with BMS, EPMS, and test equipment. Requires experience building test automation or orchestration systems for hardware/infrastructure and replacing manual processes with trusted software.

CodeRabbit

CodeRabbit

San Francisco, CA

Senior/Staff Platform Engineer
$220k+/yrHybrid7+ YOEDevOps / SRE

Founding Platform Engineer owning compute, orchestration, and infrastructure for CodeRabbit's GenAI code review platform on multi-region GCP. Build from 0-to-1 with deep Kubernetes, distributed systems, and cloud infrastructure expertise.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Software Engineer, Developer Productivity, AI Tools
$350k+/yrOn-site5+ YOEDevOps / SRE

Build and standardize AI-powered coding tools, agents, and dev environments to accelerate internal software development while maintaining security and quality. Requires experience with productivity tooling for large codebases, container/CI platforms, and AI model APIs.

Illumio

Illumio

Sunnyvale, CA

Site Reliability Engineer II
$141k+/yrOn-site2+ YOEDevOps / SRE

Site Reliability Engineer II responsible for designing, deploying, and maintaining multi-cloud infrastructure (Azure primary, AWS/GCP) for Illumio's SaaS products. Focus on IaC, CI/CD pipelines, monitoring, incident response, automation, and improving reliability/scalability in collaboration with engineering and security teams. Requires 2+ years SRE/DevOps experience with Azure.

Okta

Okta

San Francisco, CA
Staff TDI Site Reliability Engineer, Okta Federal
$174k+/yrHybrid7+ YOEDevOps / SRE

Staff SRE on Okta's TDI team building and operating secure, air-gapped cloud infrastructure, CI/CD pipelines, and monitoring for national security missions. Requires 7+ years SRE/DevOps experience, deep AWS and automation skills, and active TS/SCI with polygraph clearance.

Onxmaps

Onxmaps

Bozeman, MT

Site Reliability Engineer III
$130k+/yrHybrid5+ YOEDevOps / SRE

Site Reliability Engineer responsible for deploying, monitoring, and maintaining highly available infrastructure on GCP using Terraform, Kubernetes, and various cloud services. Requires 5+ years experience (3+ in production), strong Kubernetes/IaC background, and on-call participation to ensure reliable systems for millions of users.

Illumio

Illumio

Sunnyvale, CA

Sr. Site Reliability Engineer
$170k+/yrOn-site5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for monitoring, incident response, and optimizing the reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure and SaaS services. Requires 5+ years SRE experience with strong cloud platform expertise.

DeepIntent

DeepIntent

Pune, India

Senior Software Engineer – Platform Operations
No salary listedHybrid5+ YOEDevOps / SRE

Senior Software Engineer responsible for production platform reliability, data pipeline troubleshooting, incident management, operational tooling, and automation. Requires 4+ years of engineering experience, strong Java and SQL skills, distributed data pipeline expertise, and cloud platform experience.

Fluidstack

Fluidstack

New York, NY
Compute Deployment Engineer
$150k+/yrHybrid5+ YOEDevOps / SRE

Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.

Fluidstack

Fluidstack

San Francisco, CA
Facilities Production Technical Lead
$188k+/yrOn-site7+ YOEDevOps / SRE

Lead technical direction for facilities production engineering, architecting telemetry pipelines from OT systems (BMS/EPMS/SCADA) into modern data stacks and setting controls integration standards across massive AI data center fleet.

Arena

Arena

San Francisco, CA

Site Reliability Engineer
No salary listedHybrid6+ YOEDevOps / SRE

Build and operate the scalable, low-latency infrastructure powering Arena's real-world AI model evaluation platform, including API gateways, observability, and enterprise features for frontier model routing and evaluation.

Together AI

Together AI

San Francisco, CA
Staff Software Engineer, GPU Infrastructure Lifecycle Management
$240k+/yrOn-site7+ YOEDevOps / SRE

Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.

Pinterest

Pinterest

San Francisco, CA

Staff Software Engineer, Capacity Engineering
$177k+/yrHybrid7+ YOEDevOps / SRE

Staff Software Engineer improving efficiency and performance of Pinterest's large-scale Kubernetes and distributed cloud infrastructure. Requires deep capacity/performance expertise, experience leading efficiency initiatives at scale, and strong AI collaboration skills.

OpenAI

OpenAI

San Francisco, CA
Data Center Compute Infrastructure
$230k+/yrOn-site5+ YOEDevOps / SRE

Build, scale, and operate OpenAI's global compute infrastructure for frontier AI models like GPT-5.6. Solve complex cross-disciplinary problems spanning distributed systems, hardware, ML infrastructure, power/cooling, manufacturing, supply chain, and data center development at unprecedented scale.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Infrastructure Engineer
No salary listedOn-site3+ YOEDevOps / SRE

Infrastructure Engineer responsible for hands-on installation, provisioning, maintenance, and troubleshooting of high-performance on-premise server hardware, Linux systems, and high-speed networking (100G/400G) in a data center environment. Requires 3+ years experience with Linux admin, x86 hardware, and network configuration.

Scale AI

Scale AI

Washington, DC

DevOps Engineer, Infrastructure & Security
$149k+/yrOn-site2+ YOEDevOps / SRE

DevOps Engineer building and enhancing CI/CD pipelines for Scale's lowside and highside products in classified environments. Integrate ML tasks into automated SDLC, incorporate security best practices, and collaborate across teams. Requires active TS/SCI clearance.

Databricks

Databricks

Mountain View, CA

Senior Software Engineer, AI Native Web Platform
$160k+/yrOn-site8+ YOEDevOps / SRE

Senior Software Engineer building and scaling the AI-native web platform layer at Databricks, including CI/CD, deployment automation, consent management, accessibility testing, and tech stack unification. Requires 8+ years experience in production web infrastructure.

OneSignal

OneSignal

United Kingdom
Senior or Staff Software Engineer, SRE/Platform Team
£100k+/yrRemote8+ YOEDevOps / SRE

This role builds and operates scalable platform infrastructure, automates operational and database workflows, and improves reliability, observability, and deployment practices. It requires at least eight years of platform experience, production systems expertise, Kubernetes, cloud, Linux, and SQL datastore experience.

Mercury

Mercury

San Francisco, CA
Senior Software Engineer
$201k+/yrRemote6+ YOEDevOps / SRE

Senior frontend platform engineer strengthening React, TypeScript, and Vite foundations, optimizing testing, CI/CD, automation, and observability to enable fast, high-quality frontend development across the organization. Requires 6-8+ years experience, leadership of complex projects, and deep frontend tooling expertise.

MongoDB

MongoDB

United States

Cloud Operations Engineer
$90k+/yrOn-site2+ YOEDevOps / SRE

Cloud Operations Engineer on 2nd shift weekends responsible for monitoring Atlas platform, diagnosing incidents, on-call rotations, automation, and ensuring uptime for MongoDB customers in FedRamp environments. Requires 2+ years DevOps/SRE experience, Linux expertise, cloud familiarity, and scripting skills.

Attain

Attain

Chicago, IL

Sr/Staff Site Reliability Engineer, Consumer Apps
No salary listedHybrid6+ YOEDevOps / SRE

Sr/Staff SRE building and automating infrastructure for a fintech platform using AI agents, Terraform, Kubernetes, and GCP services. Focus on eliminating toil, improving observability, and ensuring reliability at scale for consumer apps processing billions in transactions.

Forward Networks

Forward Networks

Santa Clara, CA

Site Reliability Engineer
$230k+/yrOn-site6+ YOEDevOps / SRE

Build the Site Reliability Engineering function from the ground up at Forward, defining SLOs, building observability infrastructure, leading incident response, and embedding reliability into the SDLC for their complex SaaS platform. Requires 6+ years SRE/DevOps experience, strong networking and Kubernetes skills, and a track record maturing SRE practices.

Chronograph

Chronograph

United States

Sr. Software Engineer, Platform Engineering
$175k+/yrRemote5+ YOEDevOps / SRE

Platform engineer responsible for designing, building, and operating cloud infrastructure on AWS and Kubernetes. Focus on developer tooling, automation with AI, performance tuning, security, observability, and on-call support for a SaaS platform serving private capital markets. Requires 5+ years in DevOps/SRE/Platform roles.

Snowflake

Snowflake

Senior Software Engineer, Snowpark Container Service
$200k+/yrHybrid7+ YOEDevOps / SRE

Senior Software Engineer building Snowflake's Snowpark Container Services, a managed Kubernetes-based platform for running containerized applications inside Snowflake's Data Cloud. Lead engineering efforts on highly scalable, reliable, multi-tenant container compute infrastructure.

Twilio

Twilio

New York, NY
Software Engineer L2
$117k+/yrRemote2+ YOEDevOps / SRE

Software Engineer L2 responsible for evolving and maintaining Twilio's Compute infrastructure, including VM orchestration, AWS Auto Scaling Groups, hardened AMIs, secure container images, and automation of operational tasks in a remote-first environment.

Aurelian

Aurelian

Seattle, WA

Staff Infrastructure Engineer
$200k+/yrOn-site6+ YOEDevOps / SRE

Staff Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents that support 911 emergency call centers. Requires 6+ years in infrastructure/platform/backend roles with experience in reliability and scale.

Aurelian

Aurelian

Seattle, WA

Senior Infrastructure Engineer
$160k+/yrOn-site4+ YOEDevOps / SRE

Senior Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents used in 911 emergency response centers. Requires 4+ years in infrastructure/platform/backend roles with experience in reliability and scale.

Forus

Forus

New York, NY

Software Engineer, Platform & Infrastructure
No salary listedOn-site5+ YOEDevOps / SRE

Software Engineer owning Forus' compute platform (EKS/Kubernetes), data layer (Postgres, OpenSearch, BigQuery), cloud cost optimization, reliability (SLOs, observability), and IaC primitives in a regulated healthcare environment. Requires production Kubernetes at scale, deep AWS/Terraform expertise, and database migration experience.

Forus

Forus

New York, NY

Software Engineer, Developer Productivity
No salary listedOn-siteDevOps / SRE

Software Engineer owning CI/CD, build systems, and infrastructure for AI coding agents to accelerate the entire engineering organization's velocity at an AI-powered healthcare network company. Requires strong infrastructure instincts, Python/Bash expertise, and experience driving developer productivity and adoption of agentic workflows.

General Intuition

General Intuition

New York, NY

Site Reliability / Infrastructure Engineer
$180k+/yrOn-siteDevOps / SRE

Own reliability and infrastructure for large-scale video and social systems, including on-call response, postmortems, database scaling, infrastructure as code, and CI/CD. The role requires deep GCP, Kubernetes, Terraform, Elasticsearch, and production database experience.

Runloop

Runloop

San Francisco, CA

Site Reliability Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Site Reliability Engineer responsible for the reliability, observability, performance, and security of a core AI agent platform including code sandboxes. Requires 5+ years software engineering experience (3+ in SRE/DevOps), strong CS fundamentals, expertise in containers, cloud IaC, monitoring, and distributed systems.

Stripe

Stripe

Seattle, WA

Staff Engineer, Deployment Platform
No salary listedHybrid10+ YOEDevOps / SRE

Staff Engineer owning end-to-end delivery of Stripe's deployment platform, including orchestrator evolution, Kubernetes fleet migration, anomaly detection, and reliability. Requires 10+ years experience leading large-scale infrastructure projects with deep expertise in distributed systems, containers, and operational excellence.