Skip to content
1,061 jobs

Job results

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Socure

Socure

Bengaluru, India

Software Engineer-II SRE
No salary listedOn-site2+ YOEDevOps / SRE

Build and operate reliable, scalable production systems across AWS, Kubernetes, infrastructure automation, CI/CD, and observability. The role requires 2–4 years of SRE, DevOps, platform, or cloud infrastructure experience and strong automation skills.

DuploCloud

DuploCloud

United States

DevOps Engineer
$80k+/yrRemote2+ YOEDevOps / SRE

Customer-facing DevOps Engineer helping organizations implement secure, compliant cloud infrastructure through the DuploCloud platform. Requires 2–3 years of cloud or DevOps experience, containerization expertise, public cloud knowledge, and strong customer communication skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff, Systems Infrastructure
$200k+/yrOn-siteDevOps / SRE

Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Shield AI

Shield AI

Melbourne, Australia

Staff C++ Build and Release Engineer
No salary listedOn-site7+ YOEDevOps / SRE

Own reproducible C++ build and release workflows, containerized CI/CD, artifact promotion, and edge deployment for autonomy and vision products. The role requires strong production C++, build-system, container, Linux, and constrained-environment delivery experience.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Senior Software Engineer - Cloud Infrastructure
$190k+/yrOn-site5+ YOEDevOps / SRE

Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

GitLab

GitLab

United Kingdom

Site Reliability Engineer, Infrastructure Platforms
No salary listedRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

Mercury

Mercury

San Francisco, CA
Senior Software Engineer - SRE
$190k+/yrRemote5+ YOEDevOps / SRE

Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.

Okta

Okta

Bengaluru, India

Senior Site Reliability Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

Bloomerang

Bloomerang

United States

Senior Software Engineer, Site Reliability
$115k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.

Clear Street

Clear Street

London, United Kingdom

Senior Production Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Distributed Software Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate distributed infrastructure software that automates, schedules, observes, and repairs large-scale AI compute clusters. The role requires 5+ years of infrastructure or distributed-systems experience, strong Go and Python skills, and deep Kubernetes expertise.

Okta

Okta

Bengaluru, India

SRE Operations Engineer
No salary listedHybrid1+ YOEDevOps / SRE

Supports the reliability and day-to-day operation of Okta’s Customer Identity Cloud by monitoring platform health, handling service requests, executing runbooks, and troubleshooting production issues. Requires cloud operations experience, infrastructure knowledge, and familiarity with Kubernetes and monitoring tools.

Kong

Kong

Bengaluru, India

Site Reliability Engineer 2, Managed Gateways
No salary listedOn-site2+ YOEDevOps / SRE

Owns reliability, scalability, and performance for managed gateway services by automating cloud operations, monitoring production systems, and resolving incidents. Requires at least two years of production SRE experience plus proficiency in Golang or Python, Kubernetes, and major cloud platforms.

Fluidstack

Fluidstack

United States

Principal Operations Engineer, Network
$258k+/yrRemote7+ YOEDevOps / SRE

This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.

Fab2

Fab2

Austin, TX
Infrastructure Software Engineering Intern
$114k+/yrOn-siteDevOps / SRE

Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Writer

Writer

New York, NY
Infrastructure Engineer
$140k+/yrHybrid5+ YOEDevOps / SRE

Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.

Writer

Writer

London, United Kingdom

Infrastructure Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.

Coinbase

Coinbase

United States

Senior Network Engineer
$186k+/yrRemote8+ YOEDevOps / SRE

Own and optimize ultra-low-latency network infrastructure for institutional trading across cloud, on-premises, and colocated environments. The role requires 8+ years of network or infrastructure experience, deep routing and multicast expertise, production incident leadership, and automation skills.

Duolingo

Duolingo

Pittsburgh, PA
Senior Site Reliability Engineer
$183k+/yrOn-site5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale distributed systems, infrastructure, reliability, and incident response. Requires 5+ years of SRE or DevOps experience plus programming and container orchestration expertise.

Descript

Descript

San Francisco, CA

Software Engineer, Infrastructure
$220k+/yrRemote8+ YOEDevOps / SRE

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

Snowflake

Snowflake

Menlo Park, CA

Principal Software Engineer - Performance Engineering
$264k+/yrOn-site12+ YOEDevOps / SRE

Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.

Anthropic

Anthropic

San Francisco, CA
Staff+ Site Reliability Engineer, Safeguards ML Infra
$320k+/yrHybrid8+ YOEDevOps / SRE

Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.

Commure

Commure

Mountain View, CA
Senior Software Engineer, Infrastructure
$170k+/yrHybrid6+ YOEDevOps / SRE

Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

MongoDB

MongoDB

Palo Alto, CA
Senior Network Engineer
$118k+/yrHybrid6+ YOEDevOps / SRE

Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

PrizePicks

PrizePicks

United States

Senior Site Reliability Engineer
$120k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.

Fortanix

Fortanix

Santa Clara, CA

Senior/Staff Infrastructure & Platform Engineer
$155k+/yrOn-site7+ YOEDevOps / SRE

Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Cohere

Cohere

United States
Software Engineer, GPU Infrastructure
No salary listedHybrid7+ YOEDevOps / SRE

Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.

Ontic

Ontic

Austin, TX

Associate DevOps Engineer
$100k+/yrHybridDevOps / SRE

Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.

Scale AI

Scale AI

San Francisco, CA

Staff Network Engineer, App Platform
No salary listedOn-site7+ YOEDevOps / SRE

Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.

Together AI

Together AI

San Francisco, CA
Senior Network Engineer
No salary listedOn-site8+ YOEDevOps / SRE

Designs, operates, and troubleshoots large-scale multi-vendor networks supporting high-performance AI compute infrastructure. Requires 8+ years of production networking experience, deep routing expertise, automation skills, and hands-on RoCE or InfiniBand experience.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Motive

Motive

Buffalo, NY
Staff Platform Engineer
$164k+/yrOn-site7+ YOEDevOps / SRE

Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.

Together AI

Together AI

San Francisco, CA
AI Infrastructure Systems Engineer
No salary listedHybrid3+ YOEDevOps / SRE

Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.

Domino

Domino

Argentina

Technical Cloud Operations Lead
No salary listedRemoteDevOps / SRE

Leads technical cloud operations by standardizing production changes, building service procedures, and improving automation, documentation, and self-service. The role requires production cloud or SaaS operations experience, strong operational judgment, and familiarity with cloud-native systems such as AWS and Kubernetes.

GitLab

GitLab

Bengaluru, India

Senior Release Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.

Acryldata

Acryldata

Bengaluru, India

DevOps
No salary listedRemote5+ YOEDevOps / SRE

Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.