Latest DevOps / SRE jobs
Job results
Builds AI-native developer tooling, agentic workflows, and CI/CD automation that improve engineering velocity and software delivery. The role requires 4+ years of software engineering experience, infrastructure or internal tooling expertise, and production experience with LLM-powered systems.
Design, operate, and automate the global network and reliability layer for a high-performance NVIDIA DGX SuperPOD supporting ML workloads. Own architecture, observability, incident response, and security for mission-critical infrastructure.
Senior Infrastructure Engineer building and operating AWS cloud infrastructure for healthcare data platform. Requires Python, Terraform, CI/CD expertise, and big data tools experience.
Build and operate foundational data infrastructure including Airflow, Flink, DynamoDB, and RDS using Terraform and Kubernetes. Requires 2-4 years of infrastructure/platform experience and strong Python skills.
Hands-on DevOps role owning AWS infrastructure, building developer tooling, and driving technical roadmap at an early-stage YC startup. Requires 6+ years infra/DevOps experience and strong AWS/K8s/Terraform skills.
Design and own the OpenUSD-based asset pipeline for a high-fidelity sensor simulation platform. Build automated DCC-to-engine pipelines, custom schemas, material conversion, and validation systems at library scale.
Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.
Staff Infrastructure Engineer responsible for re-architecting Kubernetes infrastructure, improving continuous delivery, and making code changes across the stack to support drone platform needs.
Owns weekly releases and phased production rollouts for a globally distributed remote execution platform. The role focuses on Terraform-managed infrastructure, CI health, incident coordination, operational reliability, and continuous improvement of release processes.
Senior Network Engineer responsible for operating and architecting Cloudflare's global edge and backbone network, building AI-powered operational tooling, and mentoring engineers. Requires deep expertise in BGP, MPLS, network automation, and production LLM agent development.
Operates and improves large-scale, multi-region blockchain RPC infrastructure across Kubernetes, cloud, and bare-metal environments. The role requires production operations, infrastructure-as-code, GitOps, observability, incident response, and on-call experience.
Senior engineer responsible for architecting and maintaining scalable AWS cloud infrastructure, leading modernization initiatives, and ensuring PCI/SOC2 compliance. Requires 5+ years experience with Terraform, Kubernetes, observability, and production cloud systems.
Hands-on Infrastructure Tech Lead building and scaling AWS cloud infrastructure from scratch for an AI-driven enterprise analytics platform. Owns architecture, IaC, security/compliance (SOC 2), and operational excellence.
Senior Cloud Engineer owning AWS/GCP infrastructure, Kubernetes/GitOps platforms, and CI/CD systems. Designs and operates scalable, secure cloud infrastructure while mentoring engineers and enabling AI/ML tooling.
Design end-to-end datacenter network architectures for AI training and inference workloads. Own topology selection, fabric design, physical infrastructure integration, and produce deployable HLDs/LLDs across multiple GPU platforms and customer requirements.
Owns the reliability, availability, and operational health of an agentic AI platform across customer environments. The role focuses on cloud infrastructure, deployment automation, observability, incident response, and production stability, requiring 3–5 years of DevOps, infrastructure, or deployment engineering experience.
Build observability tools and platforms (metrics, logging, tracing, alerting) using Go, OpenTelemetry, and Kubernetes. Requires 5+ years experience building high-quality software that other engineers use.
Deploy and support Nominal's self-hosted platform in customer environments including air-gapped and regulated sites. Own Linux, Kubernetes, and bare-metal infrastructure reliability while partnering directly with customer IT and security teams.
Network Reliability Engineer responsible for operating and engineering Cloudflare's core data center network, building automation tools, and leveraging LLMs for deployment and troubleshooting. Requires 3+ years of network/SRE experience, strong Go/Python skills, and Linux networking expertise.
Senior IC role owning reliability, performance, and scalability of PostgreSQL (Aurora), OpenSearch, Redis, and CDC pipeline to Snowflake. Sets standards for ORM usage, migration safety, and observability at scale.
Own end-to-end health, repair automation, and qualification of a hyperscale GPU/TPU compute fleet. Build metrics pipelines, firmware tooling, and self-healing repair workflows across Kubernetes and bare metal.
Senior SRE owning reliability, monitoring, and automation for Coinbase's AI infrastructure on AWS and Kubernetes. Requires 5+ years cloud automation experience and strong incident response skills.
Staff SRE on the IT Operations team owning reliability, automation, and observability for Coinbase's AI infrastructure on AWS and Kubernetes. Requires 8+ years of cloud infrastructure experience and strong incident response leadership.
Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.
Staff-level engineer building developer tools, infrastructure, and automation to accelerate Crusoe engineering productivity. Requires Go, Kubernetes, CI/CD, and strong DevOps/SRE experience.
Senior engineer building and operating large-scale HPC infrastructure for AI model training. Owns job scheduling, automation, and performance optimization across GPU clusters.
Leads the design and production adoption of self-service infrastructure platforms, multi-region cloud foundations, and reliable delivery workflows. The role requires 8+ years of hands-on software or platform engineering experience, strong Go or comparable programming skills, and deep expertise in a core infrastructure domain.
Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.
Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.
Principal Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 8+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.
Seasoned Platform Engineer designing and maintaining scalable distributed systems and infrastructure on AWS. Builds foundational patterns, IaC, CI/CD pipelines, and observability for Python services using Kubernetes and serverless. Requires 6+ years of platform engineering experience.
Design, deploy, and secure highly available ClickHouse Cloud platforms across regulated cloud, hybrid, on-premises, and disconnected environments. Requires 6+ years of distributed-systems experience plus expertise in Kubernetes, infrastructure automation, databases, cloud platforms, and Go or Python.
Lead and develop a distributed SRE team, setting vision and operating model for reliability practices. Own core infrastructure services (Kubernetes, CI/CD, observability) and drive service ownership frameworks across engineering teams.
Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.
As a Software Engineer, Developer Experience, you will build tools and systems to improve the developer workflow, focusing on CI/CD, build toolchains, and supply chain security. You will work to make shipping to production fast, safe, and reliable for the engineering team.
The Director of Platform & Reliability Engineering will lead an engineering organization responsible for secure, scalable, and highly reliable products. This role involves setting the vision for internal platforms, cloud infrastructure, developer enablement, and production operations.
Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.
As a Staff Data Platform Engineer, you will own the reliability, stability, and operational health of the data platform infrastructure. This role focuses on systems and infrastructure, ensuring proper deployment, monitoring, maintenance, and promotion across environments.
The Member of Technical Staff, DevOps will own progressive delivery, GitOps, and on-demand environment tooling to improve deployment safety and speed for engineering teams. This role requires a platform-as-a-product mindset and experience with infrastructure as code and CI/CD pipelines.
Vapi is seeking a Site Reliability Engineer to drive 99.99% call completion for their Voice AI platform. This role involves running incident command, owning SLOs and error budgets, building reliability culture, and shipping code for platform services in Go or TypeScript.
As a Senior Developer Productivity Engineer, you will own the build, test, and deployment processes for a 50+ person engineering team. You will improve monorepo productivity, drive excellence in testing, and support multi-cloud/multi-region infrastructure to enable fast and safe shipping.
As a Senior/Staff Distributed Systems Engineer, you will design and evolve core control, data, and observability systems for LiveKit's platform, focusing on latency, availability, and operational simplicity. You will implement resilient architectures and build tools to enhance reliability and developer velocity.
This role is for a Software Engineer on the Cloud Infrastructure team, focusing on designing, building, and operating foundational cloud primitives and deployment models. The engineer will own the roadmap and technical strategy for agent-driven cloud infrastructure management, ensuring secure and scalable solutions for various customer environments.
As a Senior Site Reliability Engineer, you will own the end-to-end reliability and scalability of AWS infrastructure and Kubernetes platforms. This role involves designing, operating, and continuously improving production systems with a strong focus on automation and observability.
Senior Infrastructure Engineer responsible for building and operating platform primitives including Kubernetes, CI/CD, observability, and developer tooling at a high-growth AI and data platform company.
As a Senior Infrastructure Engineer, you will be responsible for architecting and maintaining scalable, reliable cloud infrastructure, leading incident management, and improving operational processes. This role requires strong proficiency in AWS, infrastructure-as-code, and experience with monitoring and observability tools.
As a Senior Infrastructure Engineer, you will design, implement, and maintain cloud infrastructure on GCP, focusing on CI/CD pipelines, Kubernetes, and Terraform. This role requires a strong background in DevOps/SRE and a passion for building foundational systems in a fast-paced environment.
Seeking a Senior DevOps/Site Reliability Engineer to build, operate, and scale reliable cloud-native infrastructure and distributed data platforms. This role requires expertise in Kubernetes, cloud infrastructure, observability, automation, CI/CD, and incident management.
As a Senior Production Engineer, you will ensure the reliability and scalability of Crusoe’s AI-optimized cloud platform, focusing on designing and operating managed AI services for LLM workloads. You will build automation and reliability tooling, define SLIs/SLOs, and optimize large-scale training and inference clusters.
Flint is seeking an Infrastructure Engineer to own the systems powering their AI-generated pages at scale. This 0-to-1 role involves building production-grade cloud architecture, CI/CD, deployments, observability, and security, with a focus on managing parallel background agents.