Skip to content
1,067 jobs

Job results

Coinbase

Coinbase

Canada

Senior Software Engineer, Core Reliability
CA$191k+/yrRemote5+ YOEDevOps / SRE

Senior software engineer focused on improving production reliability, deployment safety, configuration and secrets management, and scalability across Coinbase’s service environment. The role requires 5+ years of experience with distributed systems, Ruby or Go, Terraform, cloud platforms, and observability tools.

Harvey

Harvey

San Francisco, CA
Senior Software Engineer, Production Engineering
$161k+/yrHybrid5+ YOEDevOps / SRE

Build and operate Harvey's core production infrastructure powering AI workloads, including Kubernetes, compute fleets, networking, and orchestration platforms. Drive reliability, scalability, security, and cost efficiency for rapidly growing LLM infrastructure while partnering across engineering teams.

Harvey

Harvey

New York, NY
Staff Software Engineer, Production Engineering
$231k+/yrHybrid10+ YOEDevOps / SRE

Staff Production Engineer building and operating Harvey's core compute, networking, Kubernetes, and workflow orchestration infrastructure to support rapidly growing AI workloads. Requires 10+ years experience with large-scale cloud infrastructure, Kubernetes, IaC, observability, and security.

Astronomer

Astronomer

New York, NY
Staff Software Engineer, Platform Infrastructure
$275k+/yrHybrid7+ YOEDevOps / SRE

Staff Software Engineer building foundational multi-cloud platform infrastructure for Astronomer's Astro DataOps platform. Requires deep distributed systems expertise, Kubernetes operator-level knowledge, strong Go proficiency, and experience driving technical strategy at scale.

Axle

Axle

Frederick, MD

Site Reliability Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Site Reliability Engineer modernizing a multi-cloud (AWS/Azure/GCP) environment into a scalable, observable Kubernetes-based platform using DevOps/SRE practices, AIOps, IaC, and AI-driven automation to support scientific and clinical research programs. Requires 6+ years SRE/DevOps experience with strong Linux, IaC, observability, and scripting skills.

Okta

Okta

San Francisco, CA
Senior Database Reliability Engineer (DBRE)
$143k+/yrHybrid4+ YOEDevOps / SRE

Designs, operates, and optimizes large-scale PostgreSQL and MySQL databases for mission-critical systems. Builds automation, monitoring, and high-availability infrastructure while leading incident response and collaborating with engineering teams. Requires 4+ years PostgreSQL experience.

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer - Observability
$147k+/yrHybrid5+ YOEDevOps / SRE

Senior SRE specializing in Splunk observability, building scalable platforms with infrastructure as code using Terraform and Go/Python/Ruby. Requires 5+ years Splunk experience and 3+ years SRE in high-availability systems.

Okta

Okta

Washington, DC

Principal Classified Systems Architect, Okta Federal
$224k+/yrOn-site12+ YOEDevOps / SRE

Leads architecture of air-gapped classified (SIPR/JWICS) developer platforms for DoD compliance, designs IaC for disconnected ops, integrates hardened tools like Big Bang/Iron Bank, and ensures secure, scalable Kubernetes infrastructure. Requires 12+ years experience with 5+ in classified DoD environments.

Okta

Okta

Washington, DC

Senior Manager, Site Reliability Engineering (Federal)
$207k+/yrHybridDevOps / SRE

Lead and mentor multiple SRE teams overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation tooling for Okta’s high-scale SaaS infrastructure on AWS.

Pinterest

Pinterest

San Francisco, CA

Sr. Site Reliability Engineer, tvScientific
$140k+/yrRemote4+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and scaling a cloud-native CTV advertising platform on AWS and Kubernetes. Requires deep Kubernetes and AWS expertise, GitOps with ArgoCD, IaC with Terraform, CI/CD, observability, and incident response experience.

Okta

Okta

Washington, DC

Staff Software Engineer - Federal
$161k+/yrRemote8+ YOEDevOps / SRE

Staff Software Engineer builds scalable security data platforms and automates infrastructure for Okta's Public Sector using Python, Terraform, and cloud technologies. Requires 8+ years experience in security engineering and data pipelines.

ConductorOne

ConductorOne

Portland, ME

Site Reliability Engineer
$180k+/yrHybrid5+ YOEDevOps / SRE

Own reliability and scalability of a horizontal identity platform, including core cloud and FedRAMP environments. Build observability, automate operations, lead incident response, and ensure new features are reliable from the start. Requires production SRE experience at scale with Kubernetes, IaC, and strong programming skills.

Datadog

Datadog

Atlanta, GA
Senior Software Engineer
$244k+/yrHybrid5+ YOEDevOps / SRE

Senior engineer building and improving Bazel-based build, test, and packaging tools for Datadog's large multi-language monorepo. Own projects end-to-end to boost developer productivity, performance, and CI efficiency at massive scale.

Lightning AI

Lightning AI

Singapore

Infrastructure Operations Engineer
SGD165k+/yrHybrid8+ YOEDevOps / SRE

Build and operate infrastructure platforms supporting internal and customer-facing workloads, with a focus on reliability, automation, cloud, Kubernetes, storage, and networking. The role requires extensive Linux and AWS experience plus hands-on infrastructure-as-code and systems engineering skills.

Crusoe

Crusoe

San Francisco, CA

Senior Production Engineer, Managed Cloud
$170k+/yrOn-site5+ YOEDevOps / SRE

Senior Production Engineer responsible for designing, operating, and optimizing reliable managed AI cloud services focused on scaling LLM workloads, defining SLIs/SLOs, building observability, and resolving issues in distributed systems for Crusoe's AI infrastructure.

Carbon

Carbon

Sunnyvale, CA

Staff Infrastructure Software Engineer
$210k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Software Engineer owning multi-cloud (AWS/GCP) platform, Kubernetes/Istio, Terraform IaC, networking to edge printers, and CI/CD pipelines at Carbon. Requires 7+ years production cloud infrastructure experience, expert Terraform, strong Kubernetes and networking skills.

Coinbase

Coinbase

United States
Senior IT Automation Engineer
No salary listedRemote5+ YOEDevOps / SRE

Owns the architecture and delivery of complex IT automation workflows and shared platform components. The role requires at least five years of IT, automation, or systems engineering experience, plus expertise with workflow orchestration, Git-based development, identity tools, and collaboration APIs.

Zocdoc

Zocdoc

New York, NY

Staff Platform Engineer, AI Enablement
$210k+/yrHybrid7+ YOEDevOps / SRE

Lead the design and build of internal developer platforms and AI guardrails at Zocdoc. Focus on enabling both engineers and non-technical teams with secure, scalable, easy-to-use tools, CI/CD, and AI workflows while driving adoption through empathy and measurable outcomes. Requires 7+ years platform experience and a Bachelor's degree.

Glean

Glean

Palo Alto, CA

Lead Site Reliability Engineer
$200k+/yrHybrid8+ YOEDevOps / SRE

Leads SRE team to ensure high availability, scalability, and reliability of cloud services through automation, incident management, and technical leadership. Requires 8+ years SRE experience, team management, and expertise in cloud platforms and containerization.

Onebrief

Onebrief

United States

Senior Site Reliability Engineer
$180k+/yrHybrid5+ YOEDevOps / SRE

Join Onebrief's Infrastructure & Security team as an SRE focused on improving application reliability directly in the TypeScript codebase for mission-critical military planning software. Requires active Secret clearance, 5+ years shipping application code, strong observability and incident response skills, and willingness to work onsite in Arlington, VA.

Snowflake

Snowflake

Principal Software Engineer
$236k+/yrHybrid12+ YOEDevOps / SRE

Principal Software Engineer building AI-first developer platforms and tooling for Snowflake's Snowsight UX. Requires 12+ years building tools for large codebases, experience with AI/LLMs, fluency in Go/Java/Python, and strong distributed systems fundamentals.

Starburst

Starburst

India

Senior Software Engineer
No salary listedRemote5+ YOEDevOps / SRE

Senior Software Engineer responsible for building and operating the reliability, scalability, efficiency, and observability infrastructure of Starburst Galaxy. The role requires cloud architecture and orchestration experience, Java and TypeScript development, and infrastructure-as-code expertise.

Pilot.com

Pilot.com

California
Sr. Software Engineer, Infrastructure
$133k+/yrRemote5+ YOEDevOps / SRE

Senior infrastructure engineer building scalable abstractions, platform tooling, and the full infrastructure stack (AWS to product platform) that accelerates Pilot's R&D and business growth. Requires 5+ years software engineering experience, production Python, Terraform/AWS, frontend familiarity, mentoring ability, and strong collaboration/communication skills.

Nominal

Nominal

New York, NY
Software Engineer, Developer Infrastructure
$130k+/yrOn-site4+ YOEDevOps / SRE

Build, optimize, and maintain large-scale build systems (Bazel priority) and developer infrastructure including CI/CD, observability, and release automation for a fast-growing hardware-software platform company. Requires 4+ years experience with build systems at scale and large monorepos.

Plenful

Plenful

San Francisco, CA

Senior DevOps Engineer
No salary listedOn-site7+ YOEDevOps / SRE

Senior DevOps Engineer building and scaling cloud infrastructure, CI/CD pipelines, observability, and developer tooling for a healthcare AI automation platform. Requires 7+ years experience, Terraform, AWS, serverless, containers, and Postgres.

Chainguard

Chainguard

United States
Software Engineer
$121k+/yrRemote5+ YOEDevOps / SRE

Senior Software Engineer building the platform that powers secure, reproducible builds and distribution of open-source libraries across Java, JavaScript, Python/AI/ML ecosystems at Chainguard. Lead services, automation, pipelines, and developer tooling with a focus on supply-chain security, CVE remediation, and AI-assisted patching.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

AI Inference Core - Senior SW Engineer for Platform & DevOps
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate the platform layer behind Cerebras engineering infrastructure, including CI/CD, Kubernetes, deployment automation, cloud and on-premises systems, developer environments, and observability. The role requires 5+ years of infrastructure or software engineering experience and strong debugging and systems fundamentals.

Sardine

Sardine

Germany
DevOps Engineer
€115k+/yrRemote6+ YOEDevOps / SRE

Owns and evolves cloud infrastructure, CI/CD, observability, developer tooling, and FinOps governance for a distributed engineering organization. Requires 6+ years in DevOps, SRE, or software engineering, plus strong Kubernetes, cloud, security, and automation experience.

Supabase

Supabase

Remote

Release Engineer
No salary listedRemote5+ YOEDevOps / SRE

Own reliability, SLOs, and observability for Supabase's deployment pipelines, control plane, and release systems as part of the Release Engineering / SRE team. Drive safe, observable deploys, disaster recovery, incident response, and toil reduction in a fully remote, async environment.

Wispr Flow

Wispr Flow

San Francisco, CA

Staff Platform Engineer, Infrastructure
$270k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer building scalable infrastructure to support low-latency voice AI product used hundreds of times daily by millions. Architect core platform components connecting product and ML, anticipate scaling bottlenecks, improve developer experience, and set technical direction for growing team.

OnePay

OnePay

United States

Platform Engineer, Builder Experience
$140k+/yrRemote2+ YOEDevOps / SRE

Platform Engineer building core services, developer tooling, and frameworks for high-scale distributed systems and agentic applications at a consumer fintech company. Requires 2-9 years experience with AWS, distributed systems, CI/CD, and building platforms that accelerate engineering velocity.

Chainguard

Chainguard

United States

Senior Software Engineer, Developer Platform
$157k+/yrRemote4+ YOEDevOps / SRE

Senior Software Engineer building Chainguard's internal Developer Platform "Factory", including monorepo CI/CD pipelines, Agentic AI platform for automated changes, and paved-road build infrastructure to reduce developer toil and accelerate secure artifact delivery.

Onebrief

Onebrief

United States

Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Infrastructure Engineer building and securing Kubernetes-based platforms, cloud-native deployments, and air-gapped appliances for military planning software. Requires 5+ years production infrastructure experience, deep Kubernetes and cloud expertise, security fundamentals, and full-stack engineering skills in languages like Go or Python.

Crusoe

Crusoe

Dublin, Ireland

Staff Network Production Engineer, Network Ops
No salary listedOn-site10+ YOEDevOps / SRE

Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks for GPU-based HPC infrastructure. The role requires 10+ years of production network operations experience, advanced routing knowledge, automation skills, and participation in 24/7 on-call support.

Harper

Harper

San Francisco, CA

Platform Engineer
$140k+/yrOn-site5+ YOEDevOps / SRE

Build and own core infrastructure, observability, and tooling for a hyper-growth AI insurance platform running 200+ services and thousands of daily agentic AI decisions. Focus on scale, reliability, developer velocity, and AI eval systems.

Plenful

Plenful

San Francisco, CA

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for defining SLOs, owning production health, leading incident response, building observability with Datadog/Grafana/OpenTelemetry, and optimizing performance/scalability on AWS for a healthcare AI platform. Requires 5+ years SRE experience, distributed systems expertise, and strong automation skills.

Plenful

Plenful

San Francisco, CA

Staff DevOps Engineer
No salary listedHybrid10+ YOEDevOps / SRE

Staff DevOps Engineer building and scaling cloud infrastructure, CI/CD pipelines, observability, and developer tooling for a healthcare AI automation platform. Requires 10+ years experience, strong IaC and AWS skills, and focus on reliability for backend/ML teams.

SnapLogic

SnapLogic

Thailand

Software Engineer
No salary listedOn-siteDevOps / SRE

Maintains and enhances a high-scale core platform through monitoring, production troubleshooting, defect resolution, and reliability optimization. Requires Python proficiency, production-systems experience, and a bachelor's degree in computer engineering or a related field.

Tulip

Tulip

Somerville, MA

Observability Tech Lead
No salary listedHybrid5+ YOEDevOps / SRE

Tech Lead for Observability at Tulip, mentoring on best practices, SLIs/SLOs, and reliability while designing, building, and maintaining core observability infrastructure, tooling, and AI-enhanced monitoring for distributed systems and production incidents.

Tulip

Tulip

Somerville, MA

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for observability, incident response, and building reliability tooling for Tulip's AI-native operations platform. Requires 5+ years with Prometheus, OpenTelemetry, and AI-driven observability tools, plus strong systems reasoning and mentoring skills.

Kong

Kong

Washington

Senior SRE, Managed Gateways
$118k+/yrRemote7+ YOEDevOps / SRE

Senior Site Reliability Engineer owning production reliability and enterprise customer implementations for Kong's fast-growing Managed Gateways SaaS product across AWS, GCP, and Azure. Requires deep Kubernetes, cloud-native, and Golang expertise plus customer-facing technical leadership.

Clickhouse

Clickhouse

United States

Senior Cloud Engineer
No salary listedRemote6+ YOEDevOps / SRE

Design and operate secure, scalable ClickHouse Cloud platforms across regulated cloud, hybrid, on-premises, and disconnected environments. The role requires 6+ years of distributed-systems experience and strong Kubernetes, infrastructure automation, cloud, database, and security expertise.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Software Engineer, Cluster Deployment
No salary listedOn-siteDevOps / SRE

Build and maintain automation tooling for large-scale AI compute cluster deployments, turning bare-metal infrastructure into repeatable, pushbutton workflows using Python, Ansible, Terraform, Kubernetes and observability tools. Ideal for new graduates or early-career engineers seeking hands-on production infrastructure experience.

Elicit

Elicit

Oakland, CA

Infrastructure Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.

Crusoe

Crusoe

San Francisco, CA
Senior Software Engineer, Developer Experience
$172k+/yrOn-site5+ YOEDevOps / SRE

Senior Software Engineer on the Developer Experience team building internal tools, libraries, CI/CD pipelines, and paved paths to accelerate engineering productivity and eliminate toil across the full SDLC at Crusoe. Requires strong Go, Kubernetes, DevOps/SRE background and experience creating developer infrastructure.

Instacart

Instacart

United States
Staff Software Engineer, Bazel & Go
$221k+/yrRemote10+ YOEDevOps / SRE

Staff Software Engineer leading technical direction for Instacart's Bazel build system (remote execution, caching, performance) and Go platform (frameworks, libraries, patterns). Hands-on role driving initiatives to improve build times, CI reliability, and developer velocity for 1000+ engineers. Requires 10+ years experience with deep Bazel and Go expertise.

Tulip

Tulip

Somerville, MA

Lead DevOps Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Lead DevOps Engineer owning multi-cloud SaaS infrastructure at scale for Tulip's AI-native frontline operations platform. Design resilient cloud architecture, CI/CD automation, observability, and mentor engineers while partnering with application teams. Requires 5-7+ years infrastructure experience and leadership.

Clickhouse

Clickhouse

EMEA

Senior Cloud Data Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

Build and operate scalable cloud-native infrastructure for ClickHouse’s serverless database platform, including Kubernetes-based management, metrics systems, and distributed data-plane capabilities. Requires 5+ years of software development experience and production expertise with cloud platforms and Go, C++, or Java.

Clickhouse

Clickhouse

EMEA

Senior Cloud Data Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

Build and operate ClickHouse’s cloud-native database infrastructure, including Kubernetes-based management, metrics systems, and highly available distributed services. The role requires 5+ years of software development experience, production expertise in Go, C++, or Java, and experience with public cloud and data infrastructure.

GitLab

GitLab

United States

Staff Engineer, People Technology
$126k+/yrRemote7+ YOEDevOps / SRE

Staff Engineer on the People Technology team building and maintaining scalable automations, agents, and integrations across Workday, Greenhouse, Slack, Workato, and GCP. Requires strong software engineering skills applied to HR systems with heavy use of AI/LLMs to eliminate manual work.