Latest DevOps / SRE jobs
Job results
Senior SRE scales and maintains high-performance research compute clusters for machine learning in finance, ensuring uptime, reliability, and observability across on-prem and cloud infrastructure. Requires 5+ years SRE/DevOps experience with HPC frameworks, IaC, and cloud services.
Designs and builds developer tools, SDKs, APIs, and CI/CD pipelines to boost engineering productivity. Requires 8+ years experience in platform tooling, proficiency in Java/Python/TypeScript/Go, and cloud platforms like AWS/GCP.
Builds, automates, and maintains scalable, secure Linux infrastructure using modern tooling. Requires 7+ years Linux sysadmin experience, deep OS/network knowledge, Python/Ruby scripting, config management (Ansible/Puppet/Chef), and monitoring (Prometheus/Grafana).
Leads design and delivery of enterprise API integrations using Workato for People Tech ecosystem, automating workflows across ERP, CRM, HCM systems. Requires 7+ years experience with Workato recipes, iPaaS patterns, API security, and stakeholder collaboration.
Staff engineer owns architecture of Inference Cloud Platform, building distributed systems for high-QPS AI workloads with focus on availability, latency, reliability, and global scale. Requires 8+ years experience in large-scale cloud systems and backend languages like Go, C++, Python.
Principal Engineer leads Inference Cloud Platform, defining architecture for multi-region, high-QPS AI inference systems. Focuses on reliability, performance optimization, production code, and cross-team technical strategy. Requires 10+ years in distributed systems.
Leads automation and platform engineering for ultra-reliable AI inference infrastructure, architecting self-service GitOps pipelines, observability, and tooling to eliminate toil across datacenters. Requires 8+ years SRE experience with large-scale clusters and tools like Argo CD and Prometheus.
Owns deployment and operational infrastructure for Multigres distributed Postgres platform. Builds and maintains Go-based Kubernetes operator, architects cloud deployments on EKS/other K8s, manages storage/networking, and develops tooling for reliable scaling.
Builds and owns security infrastructure including runtime security, IAM, secrets management, CI/CD hardening, and compliance for cloud/container environments. Embeds with engineering teams to enable secure-by-default development using AWS, Terraform, and policy-as-code tools. Requires 8+ years DevSecOps experience.
Leads architecture and development of scalable managed Kubernetes and AI orchestration systems, providing technical direction for cloud infrastructure reliability and performance. Requires 10+ years in software engineering with deep expertise in Go, Kubernetes, and large-scale systems.
Performance Engineer optimizes infrastructure and application performance for ChatGPT and OpenAI API, focusing on latency, throughput, and efficiency at scale. Requires 7+ years in high-scale systems with expertise in profiling, tracing, and cross-layer optimizations.
Owns and maintains C++ build systems for autonomous aircraft software, improves developer velocity by optimizing CI/CD pipelines, integrates testing with simulations, and implements monitoring to resolve issues quickly. Requires 7+ years experience with deep expertise in build tools and DevOps practices.
Designs, deploys, and improves scalable blockchain infrastructure using Kubernetes, Terraform, and cloud tools. Drives AI enablement, builds observability with Prometheus/Grafana, manages multi-cloud networks, and leads incident response. Requires 5+ years in SRE/infrastructure with strong automation focus.
Designs, manages, and secures AWS cloud infrastructure across multi-account environments using Terraform for autonomous vehicle platforms. Requires 3+ years AWS experience, strong IaC skills, and security focus; GovCloud/FedRAMP preferred.
Builds infrastructure and simulation engines for training autonomous AI agents in complex environments using reinforcement learning. Requires strong engineering skills in distributed systems, containerization, networking, and data systems.
Infrastructure Engineer on the Foundation team builds and maintains highly available systems and developer tooling to ensure platform stability and productivity for processing mortgage transactions. Requires deep curiosity, full ownership from design to maintenance, and ability to solve hard problems under pressure.
Designs, architects, and scales production GPU/HPC clusters globally. Debugs hardware/software issues, automates operations, and mentors juniors. Requires 5+ years experience and hybrid SF presence.
Build and scale backend platform systems including real-time EHR integrations, data lakes from TB to PB scale, and core infrastructure for 100x growth. Requires 5+ years in scalable backends, Kubernetes, AWS, PostgreSQL, and DevOps practices.
Owns build and deployment automation, platform health, scaling, and self-service tools in a multi-account AWS environment. Requires 5+ years DevOps experience, deep AWS knowledge, IaC tools like Terraform/Pulumi, Docker, Python/Bash scripting.
Owns end-to-end production-critical infrastructure for analytics platform, building performant backend systems in Rust or C++ and operating distributed services at scale on Kubernetes. Requires strong systems experience in performance optimization, debugging, and on-call reliability.
Designs, builds, and scales customer-facing managed services with a focus on edge agents running on customer infrastructure. Provides technical oversight for high-reliability systems using eBPF, Kubernetes, and low-level Linux metrics; leads cross-team collaboration and mentors engineers.
Owns end-to-end system reliability, incident response, observability, and proactive stability improvements in a serverless AWS environment. Requires 2+ years software engineering with production-facing experience, strong debugging, and hands-on AWS/Go/TypeScript skills.
Senior SRE improves platform reliability using AI-driven automation, leads incident response and oncall for critical services, and implements safe deployment practices. Requires 5+ years experience in backend/platform engineering, observability, and high-availability systems.
Designs, builds, and operates reliable cloud infrastructure for real-time voice AI systems. Owns Kubernetes clusters, CI/CD pipelines, observability, and security using AWS and IaC tools. Requires 5+ years DevOps experience with strong Python and async programming skills.
Senior SRE ensures reliability of MongoDB's multi-tenant cloud storage layer by defining SLOs, building resilient infrastructure, optimizing performance, and participating in on-call. Requires 6+ years experience with distributed systems, Python/Go, Kubernetes, and cloud platforms.
Build and own desktop automation execution platform integrating AI agents with Windows sessions via remote protocols like VNC/RDP. Requires deep systems knowledge in OS APIs, accessibility frameworks, and automation infrastructure for reliable enterprise workflows.
Staff SRE on the Fabric team builds and maintains secure multi-cloud networking infrastructure for service communication, leveraging deep networking expertise to ensure resilience and scalability. Requires 10+ years experience in distributed systems and networking fundamentals.
Staff-level software engineer builds and scales platform infrastructure across teams, including dev tools, service infra, multicloud, auth, connectivity, API distributability, and ML adaptation systems. Requires 8+ years full-stack experience with Staff leadership, focusing on robust, scalable solutions in fast-paced AI environment.
Own and secure multi-region cloud infrastructure, delivery pipelines, and network boundaries while driving vulnerability remediation, incident response, and compliance readiness. The role requires 6+ years of DevOps or security experience and deep AWS security expertise.
Lead design, fit-out, and commissioning of data center sites focused on power, cooling, and IT infrastructure for high-density GPU workloads. Build and manage a 20-person operations team, oversee multi-site portfolio, and establish processes from scratch.
Build scalable infrastructure and data pipelines for AI/ML applications in legal tech. Collaborate with ML and dev teams to optimize performance and enhance developer productivity; requires cloud expertise and bachelor's/master's in CS.
Owns performance simulation platform infrastructure for validating autonomy code changes on robot hardware, including benchmarking orchestration, data pipelines, observability, and fleet management. Requires 5+ years experience in Python/C++, Linux systems, data engineering, with technical leadership.
Builds and maintains performance simulation platform with bench-top rigs, cloud orchestration, and data pipelines to validate autonomy code changes for real-time performance on robot hardware. Requires 3+ years experience in Python/C++, Linux systems, data engineering, with technical leadership.
Builds and operates core infrastructure including Kubernetes clusters, geo-deployments, edge security, and cost optimization to support rapid growth and team productivity. Requires deep AWS/K8s experience and strong engineering fundamentals.
Leads infrastructure team managing cloud, networking, storage, and compute for high-scale developer tool. Sets technical direction, codes, hires, and optimizes costs/regional deployments with deep Kubernetes/AWS expertise.
Leads technical architecture for data center operations, overseeing global ticket queues, fleet supportability, power topology, resilience planning, and hardware failure escalations for AI infrastructure. Requires 10+ years in data center ops or HPC with deep NVIDIA GPU expertise.
Infrastructure Engineer owns platform reliability, security, and performance in a HIPAA-compliant environment handling healthcare data. Designs infrastructure with Terraform, improves CI/CD workflows, and ensures resilient systems using PostgreSQL and modern cloud tools.
Leads R&D for next-generation data center architectures, evaluating emerging technologies across power, cooling, compute, and facilities. Defines reference designs, guides pilot-to-hyperscale progression, and influences cross-functional teams. Requires 15+ years experience and systems thinking across engineering domains.
Designs end-to-end infrastructure architecture for AI/ML workloads, including multi-cloud strategies, GPU orchestration, storage for massive datasets, and capacity planning to support production inference and research training at scale. Requires 7+ years in systems architecture with deep expertise in Kubernetes, AWS, and GPU infrastructure.
Builds and operates core distributed systems, infrastructure, and service architecture for Phylo's agentic AI platform, ensuring reliable scaled execution across cloud and enterprise environments. Requires 3+ years in backend/infrastructure engineering with Kubernetes and cloud expertise.
Senior engineer building self-service internal platform capabilities at Docker, focusing on multi-region networking, continuous deployment, EKS foundations, and AI-assisted operations to enable faster, safer provisioning for engineering teams.
Builds and operates supercomputing infrastructure for AI research including 1000+ GPU Kubernetes clusters, distributed data pipelines processing petabytes, and fault-tolerant training systems. Requires strong distributed systems intuition and experience with Python, PyTorch, and large-scale infrastructure.
Designs, builds, and operates GPU supercomputing environments for large-scale AI training and inference. Automates cluster management, extends orchestration systems, and optimizes performance metrics in collaboration with researchers.
Builds and scales core infrastructure for AI model training, data systems, and developer tools in a high-impact team. Requires backend proficiency (Python/Rust), experience with large-scale clusters like Kubernetes, and end-to-end project ownership.
Build and lead reliability engineering practices for ClickHouse Core, improving production performance, observability, incident response, and database operations. The role requires at least five years of reliability, QA, or customer-facing engineering experience, plus production SQL database and cloud expertise.
The Database Reliability Engineer will improve the reliability, scalability, performance, and incident response practices for ClickHouse Core. The role requires at least five years of reliability, QA, or customer-facing engineering experience, production database operations expertise, and strong scripting and debugging skills.
Build and lead reliability practices for ClickHouse Core, improving database performance, observability, incident response, and scalability. The role requires at least five years of reliability, QA, or customer-facing engineering experience plus production SQL database operations and cloud expertise.
The Database Reliability Engineer will improve ClickHouse Core’s reliability, scalability, performance, and production operations through observability, incident response, chaos initiatives, and database debugging. The role requires at least five years of reliability, QA, or customer-facing engineering experience and production SQL database expertise.
Build and lead reliability practices for ClickHouse Core, improving production performance, observability, incident response, and scalability. The role requires at least five years in reliability, QA, or customer-facing engineering plus production database operations experience.
Senior Platform Engineer owns cloud services, Kubernetes clusters, and developer tooling on AWS to enable engineering teams. Requires 5+ years DevOps/SRE experience, strong Kubernetes/Terraform/AWS skills, and on-call participation.