Skip to content
1,067 jobs

Job results

Imply

Imply

United States

Senior Software Engineer
$155k+/yrRemote6+ YOEDevOps / SRE

Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Trexquant

Trexquant

Stamford, CT

DevOps Engineer
No salary listedOn-site2+ YOEDevOps / SRE

The DevOps Engineer builds and scales automation, Kubernetes environments, CI/CD pipelines, and operational tooling for research and trading platforms. The role requires at least two years of Linux-centric DevOps, infrastructure, or SRE experience, with expertise in Kubernetes, pipeline engineering, and observability.

Snowflake

Snowflake

Senior Software Engineer - NatSec
$200k+/yrHybrid7+ YOEDevOps / SRE

Leads design, deployment, and operation of secure distributed cloud systems for public-sector and air-gapped environments. Requires active or obtainable TS/SCI clearance with polygraph, U.S. citizenship, and 7+ years of production experience.

Otter

Otter

Mountain View, CA

Production Engineer
$155k+/yrHybrid2+ YOEDevOps / SRE

Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.

Reddit

Reddit

United States

Staff Software Engineer, Observability
$217k+/yrRemote7+ YOEDevOps / SRE

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Lyft

Lyft

Toronto, Canada

Software Engineer, Observability
CA$108k+/yrHybrid3+ YOEDevOps / SRE

Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.

Phantom

Phantom

Remote

Staff DevOps Engineer
No salary listedRemote8+ YOEDevOps / SRE

Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.

Snowflake

Snowflake

Senior Software Engineer - Snowpark Container Service
$200k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.

Grafana Labs

Grafana Labs

United Kingdom
Software Engineer - Platform Metal
£72k+/yrRemoteDevOps / SRE

Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.

Nasuni

Nasuni

Hyderabad, India

Senior Software Engineer - Systems
No salary listedHybrid7+ YOEDevOps / SRE

Build and maintain Linux-based NAS appliance software, focusing on Python systems development, storage and filesystem performance, packaging, automation, and production debugging. The role requires 7–10 years of systems or platform engineering experience and deep Linux expertise.

Postman

Postman

Bengaluru, India

Staff Engineer – Observability Platform
No salary listedHybrid10+ YOEDevOps / SRE

Leads the technical vision and development of a scalable observability platform, improving reliability, performance, incident response, and engineering productivity across Postman. Requires 10+ years of software engineering experience with distributed systems, cloud-native architectures, and production operations.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

OpenAI

OpenAI

London, United Kingdom

Network Engineer
No salary listedHybridDevOps / SRE

The Network Engineer will design, operate, troubleshoot, and automate secure enterprise and cloud networks across offices, labs, and production services. The role combines network architecture and lifecycle planning with incident response, operational delivery, and automation using APIs, Infrastructure as Code, Git, testing, and CI/CD.

Vercel

Vercel

London, United Kingdom

Software Engineer, Compute
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate Vercel’s low-level compute infrastructure, including storage, state, clusters, and distributed workloads. The role requires 5+ years of software engineering experience, strong Go skills, and deep expertise in Linux, virtualization, schedulers, and reliable distributed systems.

ZoomInfo

ZoomInfo

Toronto, Canada

Software Engineer III
No salary listedOn-site5+ YOEDevOps / SRE

Build and operate delivery systems and infrastructure automation using Go and TypeScript/Node.js. The role combines production support, backend development, Kubernetes and cloud tooling, CI/CD, GitOps, and infrastructure-intelligence automation.

ZoomInfo

ZoomInfo

Toronto, Canada

Senior Software Engineer - CI/CD
No salary listedRemote5+ YOEDevOps / SRE

Builds and operates CI/CD platforms across GitHub Actions, Jenkins, Kubernetes, and cloud infrastructure. The role owns GitOps delivery, reusable developer tooling, observability, reliability, and migration initiatives across distributed engineering teams.

ZoomInfo

ZoomInfo

Toronto, Canada

Senior Software Engineer
No salary listedOn-site8+ YOEDevOps / SRE

Build and operate production-critical GitOps deployment platforms, shared tooling, APIs, and durable infrastructure workflows. The role requires 8+ years of software engineering experience, strong Go or TypeScript skills, and expertise with cloud, Kubernetes, and delivery systems.

6sense

6sense

Bengaluru, India

Staff Software Engineer - Infrastructure/DevOps
No salary listedOn-site10+ YOEDevOps / SRE

The role leads the design, automation, security, and reliability of multi-region cloud infrastructure and Kubernetes platforms. It requires extensive software and infrastructure engineering experience, strong AWS and infrastructure-as-code expertise, and proficiency in Python, Go, or Bash.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Airtable

Airtable

San Francisco, CA
Software Engineer, Infrastructure (2-8 YOE)
$148k+/yrHybrid2+ YOEDevOps / SRE

Backend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.

Twilio

Twilio

Ireland

DevOps Engineer
No salary listedRemoteDevOps / SRE

Build and lead the evolution of Twilio’s large-scale observability platform, including telemetry pipelines, query systems, developer tooling, and standards. The role requires expertise in observability systems, distributed systems, cloud infrastructure, and modern programming languages.

Twilio

Twilio

India

Senior Network Engineer
No salary listedRemote7+ YOEDevOps / SRE

Builds and operates Twilio’s global corporate network, VPN, zero-trust access, and cloud connectivity while monitoring performance and resolving incidents. The role requires substantial experience with Cisco, Palo Alto Networks, AWS networking, security protocols, and enterprise troubleshooting.

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Attentive

Attentive

United States

Staff Site Reliability Engineer
$180k+/yrRemote7+ YOEDevOps / SRE

Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA
Cluster Operations Software Engineer
No salary listedHybrid6+ YOEDevOps / SRE

Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.

xAI

xAI

Dublin, Ireland

Software Engineer - Network Software and Services
€80k+/yrOn-siteDevOps / SRE

Build scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.

Okta

Okta

Bellevue, WA

Senior Manager, Site Reliability Engineering - Infrastructure Platform
$176k+/yrHybrid6+ YOEDevOps / SRE

Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.

Fivetran

Fivetran

Bengaluru, India

Senior Software Engineer - Developer Productivity
No salary listedHybrid5+ YOEDevOps / SRE

Build internal developer-platform tooling, automation, observability, and AI-based workflows that improve engineering productivity across AWS, Azure, and Google Cloud. The role requires professional Python and/or Go experience, cloud experience, and strong production-service fundamentals.

Fivetran

Fivetran

Bengaluru, India
Staff Software Engineer - Infrastructure
No salary listedHybrid10+ YOEDevOps / SRE

Build and operate reliable, observable multi-cloud infrastructure for the dbt platform across AWS, Azure, and Google Cloud. The role emphasizes automation, Kubernetes administration, infrastructure as code, cloud cost optimization, developer experience, and operational reliability.

Fivetran

Fivetran

Bengaluru, India
Staff Software Engineer - Developer Productivity
No salary listedHybrid7+ YOEDevOps / SRE

Leads cross-team initiatives that improve developer productivity through internal platforms, self-service tooling, cloud automation, observability, and AI-assisted workflows. The role requires strong Python or Go experience, cloud-platform expertise, and a record of delivering ambiguous technical projects.

Fivetran

Fivetran

Dublin, Ireland
Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Owns the reliability, scalability, deployment automation, and incident response of production infrastructure for a large-scale data platform. The role requires expertise in managed Kubernetes, multi-cloud platforms, infrastructure as code, cloud networking, scripting, and Linux administration.

Fivetran

Fivetran

Dublin, Ireland
Staff Site Reliability Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Staff Site Reliability Engineer responsible for the performance, scalability, deployment robustness, vulnerability management, and incident response of a large-scale cloud data platform. The role requires deep Kubernetes and multi-cloud expertise, infrastructure automation, scripting, Linux administration, and cloud networking experience.

Fivetran

Fivetran

Bengaluru, India
Staff Site Reliability Engineer
No salary listedHybrid9+ YOEDevOps / SRE

Owns the reliability, scalability, deployment automation, and incident readiness of production infrastructure for a large-scale SaaS platform. Requires 9+ years of experience plus expertise in Kubernetes, cloud platforms, infrastructure tooling, scripting, Linux, networking, and databases.

Fivetran

Fivetran

Novi Sad, Serbia
Staff DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Own the reliability, scalability, deployment automation, and incident response of Fivetran’s production infrastructure. The role requires 7+ years of SaaS-scale experience plus deep expertise in Kubernetes, cloud platforms, infrastructure as code, Linux, networking, and programming.

Fivetran

Fivetran

Novi Sad, Serbia
Staff Site Reliability Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Owns the reliability, scalability, deployment automation, and incident response of production infrastructure for a SaaS data platform. The role requires 7+ years of experience, strong managed Kubernetes and cloud-platform expertise, and proficiency with infrastructure-as-code and scripting.

Temporal

Temporal

United States

Senior Software Engineer, Infrastructure Foundations
$176k+/yrRemote10+ YOEDevOps / SRE

Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer
$215k+/yrOn-site7+ YOEDevOps / SRE

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.

Fluidstack

Fluidstack

Remote

Principal Operations Engineer, Mechanical
$150k+/yrRemote10+ YOEDevOps / SRE

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

Anthropic

Anthropic

San Francisco, CA
Software Engineer, Infrastructure, Interpretability
$320k+/yrHybridDevOps / SRE

Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.

Alpaca

Alpaca

Americas
Senior DevOps Engineer
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates highly available GCP infrastructure and developer platforms for trading-critical systems. The role requires 5+ years of DevOps, platform, infrastructure, or SRE experience, with strong Terraform, Kubernetes, networking, CI/CD, observability, and incident-management skills.

Orkes

Orkes

EMEA

Site Reliability Engineer
$125k+/yrRemote5+ YOEDevOps / SRE

Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Datadog

Datadog

Paris, France

Senior Software Engineer - Incident Insights & Readiness
No salary listedHybrid5+ YOEDevOps / SRE

Build software, tooling, and operational frameworks that improve incident response, on-call practices, post-mortem learning, and reliability across Datadog. The role requires at least five years of software development experience, distributed-systems expertise, and strong cross-functional technical leadership.