Skip to content
1,067 jobs

Job results

Headway

Headway

San Francisco, CA
Staff Infrastructure Engineer
$265k+/yrRemote8+ YOEDevOps / SRE

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.

Prove AI

Prove AI

Ireland

Senior Manager, Platform Engineering
€120k+/yrRemote5+ YOEDevOps / SRE

Leads Ireland-based Platform Developer Enablement and SRE teams, defining platform strategy, developer self-service, reliability objectives, and observability standards. Requires senior software, SRE, or platform engineering experience, management leadership, and expertise in cloud infrastructure, Kubernetes, Terraform, CI/CD, and distributed systems.

OpenAI

OpenAI

London, United Kingdom

Software Engineer, GPU Infrastructure - ChatGPT Engineering
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.

MongoDB

MongoDB

Gurugram, India

Senior Site Reliability Engineer
No salary listedHybrid6+ YOEDevOps / SRE

The Senior Site Reliability Engineer will operate and improve Kubernetes-based distributed infrastructure for AI application workloads, focusing on scalability, observability, reliability, and tenant isolation. The role requires 6+ years of distributed-systems experience, production Kubernetes expertise, cloud infrastructure knowledge, and strong programming skills.

MongoDB

MongoDB

Bengaluru, India

Staff Site Reliability Engineer
No salary listedHybrid10+ YOEDevOps / SRE

Provides technical leadership for the reliability architecture and operational foundations of a multi-region, multi-cloud platform for AI applications. The role requires deep Kubernetes and distributed-systems expertise, infrastructure programming, cloud knowledge, and experience setting SRE standards and mentoring engineers.

MongoDB

MongoDB

Gurugram, India

Senior Platform Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate a self-service internal development platform that helps engineering teams deploy and run production services reliably. The role requires strong backend programming, production Kubernetes operations, cloud infrastructure, observability, networking, and distributed-systems experience.

MongoDB

MongoDB

Dublin, Ireland
Staff Engineer
No salary listedHybrid10+ YOEDevOps / SRE

Staff Engineer responsible for architecting and operating MongoDB’s large-scale observability collection and ingestion infrastructure. The role requires 10+ years of experience with distributed or highly concurrent systems, expert programming skills, and strong database, performance, and systems fundamentals.

MongoDB

MongoDB

Dublin, Ireland

Lead, Cloud Operations Engineering
No salary listedHybrid5+ YOEDevOps / SRE

Leads Cloud Operations Engineering activities, combining technical leadership, incident response, production troubleshooting, automation, and team coaching. The role requires expertise in Linux, networking, cloud infrastructure, monitoring, distributed systems, and at least two programming languages.

MongoDB

MongoDB

Cork, Ireland
Site Reliability Engineer , Storage Layer Services
No salary listedRemote6+ YOEDevOps / SRE

This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.

Cloudflare

Cloudflare

Atlanta, GA
Principal Systems Engineer, DevTools
$200k+/yrHybrid7+ YOEDevOps / SRE

Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.

Kraken

Kraken

LATAM

Site Reliability Engineer - Telemetry
No salary listedRemote3+ YOEDevOps / SRE

Operates and scales shared telemetry infrastructure spanning metrics, logs, traces, alerting, dashboards, and profiling. The role requires at least three years of production engineering experience, distributed-systems troubleshooting, Infrastructure as Code, container orchestration, incident response, and on-call participation.

Shield AI

Shield AI

United States

Sr. Staff Platform/Data Reliability Engineer, Databricks
$180k+/yrRemote12+ YOEDevOps / SRE

Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.

The Voleon Group

The Voleon Group

Berkeley, CA
Senior Software Engineer, Developer Experience
$225k+/yrHybrid5+ YOEDevOps / SRE

Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

Databricks

Databricks

Costa Rica

Senior IT Site Reliability Software Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Builds and operates resilient, observable cloud infrastructure and automation for internal IT services. The role requires 5+ years of production software engineering experience, strong Python skills, infrastructure-as-code expertise, and hands-on cloud and container experience.

Databricks

Databricks

Bengaluru, India

Senior Software Engineer - Observability
No salary listedOn-site5+ YOEDevOps / SRE

Develops large-scale observability infrastructure for Databricks products and platform systems, including logging, metrics, tracing, dashboards, and alerting. The role requires 5+ years of production software engineering experience and familiarity with distributed systems and observability practices.

Fusion Health

Fusion Health

Woodbridge, NJ

DevOps Engineer
$120k+/yrHybrid5+ YOEDevOps / SRE

Owns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.

Tatari

Tatari

Poland
Senior SRE
No salary listedRemote5+ YOEDevOps / SRE

The Senior SRE will design, automate, and operate high-throughput AWS and Kubernetes infrastructure, improving reliability, observability, CI/CD, and cost efficiency. The role requires 4–6 years of production SRE, DevOps, or systems engineering experience and strong Terraform, Linux, scripting, and incident-response skills.

Clear Street

Clear Street

New York, NY

Senior Software Engineer - Platform Engineer
$150k+/yrHybrid5+ YOEDevOps / SRE

Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.

GitLab

GitLab

Canada
Site Reliability Engineer, Intermediate to Senior Staff
$126k+/yrRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.

Grafana Labs

Grafana Labs

United Kingdom
Staff Software Engineer - Databases SRE
£104k+/yrRemote8+ YOEDevOps / SRE

Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.

OnePay

OnePay

United States

Platform Engineer
$170k+/yrRemote7+ YOEDevOps / SRE

Builds and operates core platform services, Kafka-based event streaming, and developer frameworks for high-scale distributed systems in fintech. Requires 7+ years experience with AWS, Kubernetes, and cloud-native infrastructure.

Mattermost

Mattermost

United States

Lead Site Reliability Engineer
$145k+/yrRemote7+ YOEDevOps / SRE

Leads the architecture, reliability, observability, and operational excellence of secure cloud and hybrid infrastructure for a mission-critical collaboration platform. Requires 5+ years in SRE, DevOps, or cloud infrastructure, with expertise in Kubernetes, Terraform, AWS, and regulated environments.

Zoox

Zoox

Foster City, CA
Staff Software Engineer, HPC
$230k+/yrHybrid7+ YOEDevOps / SRE

Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.

Solace

Solace

United States

Senior Platform Engineer
No salary listedRemote5+ YOEDevOps / SRE

Build and operate scalable cloud infrastructure and application platforms, enabling frequent deployments, resilient systems, observability, and self-healing capabilities. The role requires strong troubleshooting, Linux and cloud experience, networking knowledge, and expertise in one or more platform engineering focus areas.

Kraken

Kraken

LATAM

Infrastructure Engineer - Core Infrastructure
No salary listedRemote3+ YOEDevOps / SRE

Operates and scales Kraken’s core infrastructure platforms, with a focus on OpenStack, Ceph, Linux, distributed systems, and automation. The role requires 3+ years of infrastructure or software engineering experience and supports reliable compute and storage services across cloud and on-premises environments.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Continuous Integration
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate highly reliable continuous integration infrastructure, intelligent test-selection systems, and incident-response automation at scale. The role requires 10+ years of experience with large-scale CI/CD systems, container orchestration, and developer productivity tooling.

Okta

Okta

Bengaluru, India

Staff Site Reliability Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Leads reliability engineering for large-scale, customer-facing cloud services, including incident response, observability, automation, and platform improvements. Requires deep Kubernetes, cloud infrastructure, infrastructure-as-code, software engineering, and technical leadership experience.

xAI

xAI

Memphis, TN
Operations Engineer
No salary listedOn-site1+ YOEDevOps / SRE

Improves facility operations through process standardization, maintenance and construction workflow optimization, operational dashboards, data analysis, and automation. Requires an engineering bachelor's degree, at least one year of operations or process-improvement experience, and Excel, SQL, or Python skills.

Zoox

Zoox

Foster City, CA
Release Engineer
$140k+/yrHybrid3+ YOEDevOps / SRE

Own end-to-end releases for vehicle software components while coordinating cross-functional readiness and building automation for branch management, CI/CD, visibility, and deployment. The role requires 3–5 years of release or build engineering experience, Git and scripting expertise, and a bachelor's or master's degree or equivalent experience.

Lyft

Lyft

Toronto, Canada

Software Engineer, Async Platform
CA$108k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable asynchronous platform infrastructure, developing automation and tooling while improving reliability, performance, and operational efficiency. The role requires 5+ years of software development, automation, and systems engineering experience, plus cloud and distributed-systems expertise.

Webflow

Webflow

Argentina

Senior Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

Owns and evolves Webflow’s highly available, multi-cloud infrastructure, including Kubernetes, networking, infrastructure as code, observability, and AI-powered automation. The role requires 5+ years operating customer-facing cloud infrastructure and deep AWS experience.

Webflow

Webflow

Argentina

Staff DevOps Engineer, Delivery Loop
No salary listedRemote7+ YOEDevOps / SRE

Leads Webflow’s deployment strategy and GitOps platform, improving CI/CD reliability, progressive delivery, and developer productivity. The role requires 7+ years in DevOps, SRE, or infrastructure engineering and deep experience with Kubernetes, AWS, Docker, and infrastructure as code.

Tatari

Tatari

New York, NY
Senior Software Engineer, Internal Tools & Automation
$160k+/yrHybrid5+ YOEDevOps / SRE

Leads developer experience and AI tooling across the engineering organization, building internal agents, ephemeral environments, and faster CI/CD workflows. Requires 5+ years in cloud infrastructure, platform, or developer tooling, plus hands-on AI assistant and LLM workflow experience.

Cloudflare

Cloudflare

Atlanta, GA
Network Deployment Engineer
$126k+/yrHybrid2+ YOEDevOps / SRE

Deploys and expands global datacenter and physical network infrastructure, coordinating contractors, vendors, installations, and operational processes. Requires at least two years of datacenter or Linux systems administration experience plus networking, configuration management, scripting, and project coordination skills.

Okta

Okta

Bengaluru, India

Staff Site Reliability Engineer, Security - GCP
No salary listedOn-site8+ YOEDevOps / SRE

Hardens and operates large-scale GCP and AWS infrastructure, automating security remediation, incident response, IAM, and reliability improvements. The role requires deep cloud security and DevSecOps experience, strong infrastructure-as-code skills, and expertise in Kubernetes, Linux, and security automation.

xAI

xAI

Memphis, TN
Manager, Data Center Operations
No salary listedOn-site5+ YOEDevOps / SRE

Manages data center technicians and critical infrastructure supporting AI compute systems, including power, cooling, networking, hardware deployments, incidents, vendors, and capacity expansion. Requires 5+ years in data center operations and 3+ years managing technical teams.

xAI

xAI

Memphis, TN
Supervisor, Data Center Operations
No salary listedOn-site5+ YOEDevOps / SRE

Supervise data center technicians while overseeing server and network infrastructure installation, maintenance, troubleshooting, and operational improvement. The role requires 5+ years of relevant hardware and repair experience, technical leadership, Linux proficiency, and scripting experience.

Kustomer

Kustomer

New York, NY

Software Engineer, Cost Optimization
$130k+/yrHybrid4+ YOEDevOps / SRE

Own end-to-end cost optimization across AWS infrastructure and AI/LLM usage, building tooling, dashboards, and guardrails while balancing cost, performance, and reliability. Requires 4+ years of AWS infrastructure experience, cost-management expertise, and proficiency with Terraform and a programming language.

Notion

Notion

Hyderabad, India

Software Engineer, Infrastructure
No salary listedHybrid4+ YOEDevOps / SRE

The engineer will evolve and operate Notion’s async task runner and configuration management platform, supporting reliability and scalability for more than 100 million users. The role requires at least four years of software development experience and knowledge of distributed systems, production operations, and infrastructure tradeoffs.

Cloudflare

Cloudflare

Singapore

Network Deployment Engineer - APJC
No salary listedHybrid5+ YOEDevOps / SRE

Deploys and expands Cloudflare’s global data center and network infrastructure, coordinating vendors and contractors while automating provisioning and operational workflows. The role requires at least five years of relevant infrastructure experience, strong networking and Linux skills, and proficiency with Python or Bash automation.

Cognition

Cognition

San Francisco, CA
Site Reliability Engineer
No salary listedOn-siteDevOps / SRE

Owns production reliability (SLOs, monitoring, incident response) and platform engineering (CI/CD, infrastructure as code) for AI developer tools used by hundreds of thousands. Requires deep production systems experience, strong coding skills, and cloud proficiency.

Cognition

Cognition

San Francisco, CA

Research Engineer, Infrastructure
No salary listedOn-siteDevOps / SRE

Builds and owns distributed training infrastructure, experiment orchestration, data pipelines, and performance optimizations for large-scale AI research on GPU clusters. Requires deep systems expertise, Python/C++/PyTorch proficiency, and ML understanding to accelerate frontier research.

Cognition

Cognition

San Francisco, CA
Software Engineer, Infrastructure
$260k+/yrOn-siteDevOps / SRE

Build and operate the compute, orchestration, networking, developer platform, and reliability systems underlying AI agents and developer tools. The role requires large-scale infrastructure experience, Kubernetes and cloud expertise, Python proficiency, and a strong security and observability mindset.

Stripe

Stripe

Bengaluru, India

Staff Engineer, Core Infrastructure
No salary listedOn-site12+ YOEDevOps / SRE

Leads the design and validation of resilient regional infrastructure for globally scaled payment systems, including launches, failovers, migrations, and CI/CD readiness gates. Requires 12+ years of software or infrastructure engineering experience, distributed-systems expertise, cloud experience, and strong technical leadership.

PostHog

PostHog

United States

Site Reliability Engineer
No salary listedRemote5+ YOEDevOps / SRE

Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.

PostHog

PostHog

San Francisco, CA

SRE - Infra
No salary listedRemoteDevOps / SRE

Owns and automates production infrastructure on multi-region AWS with EKS clusters, focusing on scaling, reliability, and self-healing systems. Requires deep Kubernetes, Terraform, and Linux expertise for large-scale stateful workloads.

OPSWAT

OPSWAT

Timișoara, Romania

DevOps Engineer
No salary listedOn-site2+ YOEDevOps / SRE

The DevOps Engineer designs, automates, and supports secure, scalable AWS infrastructure, Kubernetes environments, CI/CD pipelines, and monitoring systems. The role requires 2–4 years of DevOps experience plus expertise in infrastructure as code, testing, troubleshooting, and production support.

Applied Intuition

Applied Intuition

Munich, Germany

Software Engineer - Virtualization
No salary listedOn-site2+ YOEDevOps / SRE

Build and improve CI/CD, deployment, build, and validation infrastructure for embedded automotive software across SIL/HIL and cloud environments. The role requires 2+ years of experience, a bachelor's degree, and familiarity with automotive toolchains and embedded systems.

StackAI

StackAI

San Francisco, CA
Senior DevOps Engineer
$130k+/yrHybrid5+ YOEDevOps / SRE

The Senior DevOps Engineer will design and operate secure, scalable cloud infrastructure, lead Kubernetes operations, automate environments, and improve CI/CD and observability. The role requires 5+ years of DevOps or infrastructure experience and expertise with Kubernetes, Terraform, Docker, and cloud platforms.