Latest DevOps / SRE jobs
Job results
Build and harden Linux networking and connectivity systems for autonomous vehicles, spanning high-throughput internal networks, time synchronization, cellular failover, embedded hardware integration, and fleet observability. The role requires strong systems programming in Go and Python plus deep Linux networking and hardware-boundary debugging experience.
Leads and mentors a DevOps team while remaining hands-on in building secure, multi-region Kubernetes platforms, cloud infrastructure, GitOps delivery, and reliability practices. The role also supports AI platform infrastructure and requires strong expertise in Kubernetes, Terraform, security, observability, and team leadership.
Build and operate large-scale, geographically distributed infrastructure supporting massive file metadata, data volumes, analytics, and concurrent connections. The role requires 5+ years of software development experience and expertise in backend systems, programming, operating systems, and distributed infrastructure.
Leads reliability engineering for large-scale, customer-facing cloud services, driving automation, observability, incident response, and operational excellence. The role requires deep Kubernetes, cloud infrastructure, infrastructure-as-code, software engineering, and distributed systems expertise, along with cross-team technical leadership and mentoring.
Operates and improves reliability for a trading-critical brokerage platform across cloud infrastructure, Kubernetes, observability, messaging, and PostgreSQL. Requires 4+ years of production operations experience, strong PostgreSQL fundamentals, incident response expertise, and proficiency in Go or Python.
Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud, with deep Kubernetes, automation, observability, and SRE expertise. The role requires 8+ years in SRE, DevOps, or infrastructure engineering and experience leading large-scale migrations and cross-team reliability initiatives.
Leads platform engineering for autonomous robots across embedded Linux, networking, middleware, CI/CD, containerized deployments, and cloud-connected services. The role requires hands-on C, Python, Bash, networking, and Linux systems expertise while mentoring a small engineering team.
Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.
Build and operate resilient, large-scale platform services across AWS and Azure, managing Kubernetes clusters, infrastructure automation, deployment orchestration, and observability. The role requires 4+ years of software engineering experience with Kubernetes, Terraform, Go, and distributed systems.
Leads a storage engineering team while architecting and operating highly available Linux-based storage, datacenter, and data-protection infrastructure. The role requires deep Ceph experience, PB-scale archiving and backup expertise, hands-on troubleshooting, and team leadership.
Own platform initiatives that improve cloud scalability, reliability, automation, and developer productivity. The role requires 5+ years of cloud infrastructure experience plus hands-on expertise with infrastructure as code, Kubernetes, CI/CD, programming or scripting, and AI-assisted engineering.
Fluidstack is seeking a Principal Operations Engineer, Electrical to be the senior technical authority for electrical infrastructure across their hyperscale AI data center portfolio. This role involves leading site assessments, driving technical readiness, reviewing designs, and feeding operational learnings back into the design and manufacturing organization.
Build and operate Discord’s greenfield Enterprise Platform, turning identity, device management, infrastructure, and application delivery into reusable self-service software. The role requires production software engineering experience, IAM and Terraform expertise, and end-to-end ownership of complex platform projects.
Build the foundational Enterprise Platform engineering discipline, turning identity, device management, infrastructure, and application delivery into self-service, API-driven software. The role requires 6+ years of production engineering experience, deep IAM expertise, and strong Terraform and CI/CD skills.
Rotates across credit, equity, forex, and derivatives desks to monitor execution algorithms, manage portfolio risk, interface with brokers, and build trading automation under senior mentorship in a machine learning-driven hedge fund.
Build and operate automation, monitoring, validation, and remediation systems for a large-scale GPU fleet supporting AI training and inference. The role requires 3+ years of distributed systems or infrastructure software experience and strong Python, Go, or Rust skills.
Builds tools, libraries, and deployment frameworks that improve developer experience, build performance, and software delivery across automotive and cloud platforms. The role requires experience with programming, CI and deployment technologies, Linux or ARM systems, and configuration management.
Staff software engineer responsible for improving reliability across Anthropic’s AI serving systems, from SDK and network layers through infrastructure and accelerators. The role focuses on SLOs, observability, high availability, incident response, and resilience across distributed systems.
Software Engineer building automated customer deployment systems, IaC templates, and complex cloud networking setups (AWS/GCP) for Glean's Work AI platform. Requires 4+ years experience in cloud infrastructure, Terraform, and strong networking fundamentals.
Build reliable, productized systems for customer cloud setup and deployment across AWS and GCP. The role requires 4+ years of software engineering experience, backend fundamentals, infrastructure-as-code expertise, and strong networking knowledge.
Leads the reliability, observability, incident management, QA, and release engineering functions for a cloud-native healthcare SaaS platform. Requires 12+ years of SRE, infrastructure, or platform engineering experience and significant engineering leadership experience.
The Senior Site Reliability Engineer will build and scale cloud infrastructure, observability, networking, and platform services that support Carta’s applications. The role requires strong experience with cloud platforms, infrastructure as code, Kubernetes, monitoring, Python, API services, and reliability practices.
Improves the reliability, observability, and deployment safety of Block’s critical platforms while leading high-severity incident response and on-call operations. Requires strong production reliability experience, incident management skills, and 5+ years of software development experience.
Build and operate secure, scalable AWS infrastructure and developer enablement systems for internal engineering teams. The role requires 5+ years in SRE, DevOps, or systems engineering, with strong expertise in AWS governance, Terraform, Python automation, CI/CD, Kubernetes, and observability.
Lead and scale the Data Center Infrastructure Management (DCIM) program, owning the Nlyte platform and architecting workflow and change management policies. This role drives operational excellence and ensures the performance, availability, and security of Cloudflare's global network.
Staff-level engineer who will establish company-wide strategy and platform capabilities for software quality, production reliability, observability, and safe delivery. The role requires cross-team technical leadership, distributed-systems experience, cloud expertise, and strong programming skills.
Defines the technical direction and builds the secure, reliable, cloud-agnostic platform enabling enterprise customers to deploy Lovable across major cloud providers. The role requires Staff-level infrastructure expertise, strong programming skills, distributed-systems experience, and architectural ownership.
Build and operate highly available Kubernetes platforms on AWS, with responsibility for platform creation, scaling, service mesh, automation, security, and incident response. The role requires extensive infrastructure experience and expertise in Kubernetes, Helm, Karpenter, Istio, CI/CD, and cloud-native systems.
Owns the systems that build, package, update, secure, and observe software running on customer-hosted edge appliances. The role requires 5+ years of infrastructure engineering experience, deep Linux expertise, release and rollback ownership, and proficiency with systems languages, packaging, IaC, containers, and CI/CD.
Build and optimize developer productivity infrastructure spanning local development, CI, testing, code review, AI-assisted workflows, and engineering metrics. The role requires 7+ years of software engineering experience and proficiency in Go, TypeScript, Python, Rust, and shell.
Senior Production Engineer responsible for optimizing Crusoe's virtualization, hypervisor, and Linux kernel stack to deliver high-performance AI and HPC compute infrastructure. Requires 5+ years experience with kernel internals, KVM/QEMU, low-level debugging, and performance tuning for GPUs and DPUs.
Senior Production Engineer focused on building automation, self-healing tools, and reliability for Crusoe's SDN infrastructure that powers AI and HPC workloads. Requires 5+ years experience automating network provisioning with deep expertise in SDN platforms, Linux networking, Kubernetes CNIs, and protocols like BGP.
Build and optimize distributed, fault-tolerant cloud storage systems (block, file, object) for Crusoe's AI/HPC infrastructure. Ensure high availability, performance, and reliability through automation, incident response, and collaboration with hardware/kernel teams. Requires 5+ years in storage engineering/SRE with deep Linux and IaC expertise.
Staff Site Reliability Engineer on Okta's Federal SRE team building and operating highly reliable, scalable, secure cloud services for emerging products. Requires deep Kubernetes, IaC, and reliability engineering expertise plus active TS/SCI clearance and FedRAMP/IL6 experience.
Provide front-line operational support and incident response for critical FCM financial processes at NinjaTrader. Design monitors, maintain runbooks and DAGs, implement observability with OpenTelemetry/Kafka, and collaborate with engineering to improve resiliency of distributed systems.
Leads reliability engineering for a large-scale platform, designing multi-region failover, observability, disaster recovery, and resilient distributed systems. Requires 8+ years of engineering experience, expert Java or Scala proficiency, AWS operations expertise, and strong incident-management leadership.
Builds, secures, operates, and automates data-center and network infrastructure for a global cryptocurrency exchange. The role requires network engineering experience, hands-on physical infrastructure and firewall expertise, cloud networking knowledge, and an automation-focused mindset.
The Staff Software Reliability Engineer will design, build, optimize, and operate scalable streaming and distributed data-platform infrastructure supporting analytics and machine learning. The role requires 5+ years of industry experience, strong software engineering skills, and expertise with distributed data technologies and reliability practices.
The Senior Site Reliability Engineer will improve the reliability, resilience, monitoring, and incident response of Auth0’s large-scale production systems. The role requires 3+ years in SRE or cloud operations, experience with Go, shell scripting, and Terraform, and participation in rotational 24/7 on-call coverage.
Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud. The role requires deep Kubernetes expertise, SRE practices, infrastructure as code, automation, observability, and at least eight years of relevant experience.
Leads reliability strategy, architecture, and operational excellence for large-scale cloud products, while building automation and internal platforms. The role requires deep Kubernetes, cloud infrastructure, distributed systems, observability, and reliability engineering expertise, plus strong cross-organizational technical leadership.
Build and evolve the cloud infrastructure, CI/CD platforms, and GitOps standards powering Okta’s AI platform and broader workloads. The role requires 6+ years of hands-on software, cloud, and large-scale infrastructure experience, with deep expertise in AWS, Kubernetes, Argo tooling, Terraform, and Helm.
Senior individual contributor responsible for designing, securing, automating, and optimizing AWS and Azure infrastructure in a hybrid environment. The role requires 6+ years of infrastructure experience, strong IAM and cloud security expertise, infrastructure-as-code skills, and operational leadership.
Staff Software Engineer on the Developer Experience team at Grow Therapy, owning high-impact platform and AI-native tooling to accelerate 120+ engineers. Sets org-wide standards for build systems, CI/CD, monorepo performance, and agentic AI workflows while driving monolith decomposition and measuring adoption.
Build and own the infrastructure platform supporting Topsort’s real-time auction engine, APIs, and developer tooling. The role requires production Kubernetes, cloud, infrastructure-as-code, CI/CD, distributed-systems, observability, and security experience.
Operates and improves large-scale cloud infrastructure, focusing on reliability, observability, incident response, automation, and disaster recovery. The role requires production cloud experience, scripting or programming skills, containers, infrastructure as code, CI/CD, and modern monitoring tools.
Maintains VMware, Jenkins, Linux, and automation infrastructure that powers software build, test, and release workflows. The role requires 2–4 years of infrastructure experience, strong systems administration skills, and hands-on expertise with VMware, Jenkins, and infrastructure-as-code tools.
The DevOps Engineer will design and operate highly available, multi-tenant cloud infrastructure, automation, CI/CD, observability, security, and data or ML systems. The role requires 7+ years of infrastructure experience, cloud expertise, and the ability to work in a fast-paced global startup environment.
Build and maintain GenAI tooling, automation, and integrations that accelerate software engineering workflows for robotics and autonomy teams. Requires hands-on LLM tool development, DevOps automation, secure cloud/container experience, and ability to obtain SECRET clearance.
Build and operate Kubernetes-based compute orchestration infrastructure used across Coinbase, while developing developer tooling, automation, and AI-enabled workflows. The role requires 5+ years of software engineering experience, including substantial experience operating Kubernetes or comparable systems in production.