Latest DevOps / SRE jobs
Job results
Builds and maintains AI platform infrastructure for agentic systems, including connectors, execution environments, governance, and self-service tools to enable safe, scalable AI use across engineering and business teams. Requires 8+ years experience with LLM agents, GCP, and cloud-native tech.
The Staff SRE Engineer will drive reliability, scalability, observability, and operational efficiency across highly available cloud and distributed systems. The role requires 5+ years of SRE, DevOps, or platform experience, advanced Kubernetes expertise, strong automation skills, and leadership in incident management.
Site Reliability Engineer owns reliability of multi-cloud Kubernetes infrastructure for AI/ML platform, builds observability tooling as code, automates mitigations, leads incident response, and defines SLOs/SLIs. Requires extensive Kubernetes and observability experience.
Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.
Build and maintain scalable infrastructure, CI/CD workflows, and cloud or datacenter automation for Cerebras’s AI software stack. The role requires a current university student or new graduate with software development experience and proficiency in Python, shell scripting, containers, Jenkins, and cloud platforms.
Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.
Builds and maintains bare metal provisioning, orchestration engine, and internal tools for Railway's infrastructure platform. Optimizes fleet efficiency and develops resilient services using Golang/Rust, Ansible, and Terraform for distributed systems.
Develops and optimizes Kubernetes-based infrastructure for high-performance AI inference services, including deployment, scaling, debugging, and integration with ML workflows. Requires Master's in CS and 1+ year experience with Docker, Kubernetes, Python, and related tools.
Develops resilient, high-availability software for AI inference on AWS, including deployment workflows, container orchestration with Docker/Kubernetes, monitoring, and debugging. Requires Master's in CS and 18 months experience with AWS services, IaC tools, and Python.
Builds governance infrastructure for enterprise-scale Retool platform, including data access controls, admin tools, security monitoring, and organizational management systems. Requires 2-8 years full-stack experience with backend systems design and customer-focused problem-solving.
Designs, builds, and operates secure, scalable infrastructure including Kubernetes clusters and GPU hardware for large-scale AI workloads in classified US government environments. Requires 5+ years experience, Top Secret clearance, and expertise in IaC tools like Terraform and Ansible.
Builds and maintains core backend systems, infrastructure, and automation for platform reliability and scalability. Owns troubleshooting across stack, integrations with external services, and observability. Requires strong Rust, systems engineering, and distributed systems experience.
This senior engineer will operate and optimize GPU and accelerator infrastructure for AI training, inference, evaluation, and experimentation. The role requires 5+ years of infrastructure experience, production GPU cluster operations, strong systems fundamentals, and expertise in serving, observability, reliability, and compute-cost optimization.
Design, implement, and maintain cloud infrastructure and CI/CD pipelines. Collaborate with developers, SRE, and Security to ensure system reliability, scalability, and security.
Builds and scales reliable cloud infrastructure, deployment systems, observability, and developer tooling to support mortgage market operations. Requires experience with strongly typed languages, PostgreSQL, Kubernetes, and major cloud providers.
Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.
Builds and scales platform infrastructure on AWS EKS with GitOps via ArgoCD, manages CI/CD with GitHub Actions, drives observability using Datadog/Sentry/CloudWatch, and ensures reliability through SLOs and incident response. Requires 3+ years SRE/DevOps experience and Kubernetes expertise.
Builds and maintains automated testing infrastructure, CI/CD pipelines, and developer tools using Cypress and GitHub Actions/CircleCI. Mentors QA engineers and drives testing best practices in a legal SaaS platform. Requires 5+ years Cypress experience and expertise in major languages.
Builds and maintains Kubernetes-based infrastructure for managed TimescaleDB cloud services, develops Go microservices and operators, automates database operations, and ensures platform scalability and reliability. Requires 3+ years experience with Go, Kubernetes, and PostgreSQL.
Develops and maintains custom networking operating system firmware for AI supercomputers, integrating Linux kernel, switch ASICs, and control-plane services. Requires deep expertise in SONiC, SAI, routing protocols, and platform bring-up across hardware and software boundaries.
Senior SRE responsible for operating and improving highly available cloud platforms, distributed data systems, observability, incident response, and deployment automation. Requires 5+ years of SRE, DevOps, or platform engineering experience, advanced Kubernetes expertise, cloud proficiency, and strong Python and Bash skills.
Build scalable infrastructure, integrations, and data platforms powering workforce management and AI agent products at enterprise scale. Requires 5+ years in backend/platform systems, with expertise in AWS, Kubernetes, Go/Python, and datastores like Postgres and Snowflake.
Leads end-to-end integration of agentic platform tools with clients, backends, and cloud ops on Kubernetes. Requires 7+ years experience, strong Python, HTTP/auth expertise, and cross-functional collaboration for reliable tool contracts and releases.
Senior Platform Engineer architects and maintains infrastructure for autonomous vehicle orchestration platform, building CI/CD pipelines, managing AWS/on-premises/edge deployments, and ensuring DoD security compliance. Requires 5+ years in DevOps with Docker, AWS, and IaC expertise.
Senior Production Engineer ensures reliability, scalability, and performance of GPU cloud infrastructure powering AI workloads. Drives observability, incident response, automation, and operational improvements in large-scale distributed systems.
Leads technical strategy and builds scalable developer infrastructure including build systems, CI/CD pipelines, and tooling for large monorepo environments. Requires 3+ years leading complex projects, proficiency in Python/Rust/Go, and experience with container orchestration.
Leads infrastructure transformation from monoliths to scalable microservices at massive scale, architects observability/CI/CD systems, unifies complex stacks, and mentors engineers. Requires 10+ years coding internal tools, 5+ years cloud (GCP/AWS), Bachelor's in CS.
Build and operate Kubernetes-based infrastructure for secure, reliable enterprise AI deployments across cloud and customer-managed environments. The role requires production Kubernetes experience, cloud infrastructure expertise, infrastructure as code, networking, security, and deployment automation.
Builds and scales core infrastructure including ML training/serving, Kubernetes clusters, and low-latency voice/audio pipelines. Requires 3+ years in infrastructure/ML systems, hands-on reliability engineering, and Kubernetes expertise.
Designs and optimizes distributed training systems scaling across thousands of GPUs for large AI models. Requires strong systems engineering, PyTorch/JAX expertise, and collaborative mindset to boost research productivity.
Designs and optimizes distributed training infrastructure for large-scale LLMs, focusing on low-precision numerics, kernel optimizations, and communication frameworks to enable stable, scalable trillion-parameter model training. Requires strong systems engineering, deep learning frameworks knowledge, and collaborative research mindset.
Designs and optimizes high-performance ML kernels (CUDA, CuTe, Triton) for large-scale LLM training, focusing on GPU efficiency, low-precision formats, and distributed compute. Collaborates with researchers to bridge algorithms and hardware.
Designs, optimizes, and scales infrastructure for high-performance AI model inference, focusing on latency, throughput, efficiency, and reliability. Collaborates with researchers to enable production deployment of large-scale models using deep learning frameworks and distributed systems.
Build and operate internal developer platforms, cloud infrastructure, CI/CD systems, and observability tooling that improve engineering productivity and production reliability. The role requires hands-on experience with AWS, Kubernetes, distributed systems, and infrastructure automation.
Builds and operates resilient systems for autonomous vehicle fleet reliability, including pipelines for signal analysis, automated triage tools, internal workflows, and leading investigations. Requires production software experience and strong debugging skills in Python, Go, Bash, C++.
Forward Deployed Site Reliability Engineer responsible for owning reliability, observability, incident response, and deployments of a mission-critical platform in a restricted, air-gapped government AWS environment. Requires 5+ years SRE/production ops experience, strong Linux/Docker/Terraform skills, LGTM stack proficiency, and active TS/SCI clearance.
Builds foundational platform architecture for AI research agents, including SDKs, APIs, execution frameworks, and benchmarking infrastructure to enable researchers to develop intelligent systems over scholarly literature. Requires strong Python skills, 8+ years experience, cloud infrastructure, and AI integration expertise.
Optimizes performance across Codex AI system's stack including LLM inference, cloud orchestration, and agent behavior to reduce latency and costs. Collaborates with researchers and engineers on high-impact improvements in a high-ownership role.
Builds and scales developer platforms, CI/CD systems, testing infrastructure, and AI integrations to boost engineering velocity and reliability at an AI-native company. Requires 7+ years backend experience, leadership, and tools like Python, Kubernetes, and Terraform.
Build high-scale observability pipelines and alerting engines handling 1M+ RPS for logs/metrics, develop Golang/Rust gRPC services and APIs, and manage immutable infrastructure with Terraform/Ansible in a distributed systems environment.
Owns GPU diagnostics, validation workflows, and automation for bare-metal infrastructure supporting AI/ML workloads. Requires 5+ years in systems engineering with strong Linux, Python, and NVIDIA tools expertise.
Operate and scale distributed storage systems like VAST and Ceph for AI/ML workloads, build Python automation tools, manage Linux bare-metal systems, and collaborate on infrastructure optimizations. Requires 5+ years in infrastructure engineering with storage expertise.
Builds and scales reliable infrastructure for SaaS applications using Kubernetes, Terraform, and GitHub CI/CD. Focuses on observability with Grafana/Prometheus, automation to reduce toil, production troubleshooting, and cross-team collaboration. Requires 5+ years Python experience.
Build, deploy, and operate physical and cloud infrastructure for Palantir’s air-gapped and edge environments. The role requires 5+ years managing large-scale systems, hands-on Linux, Kubernetes, networking, hardware, and distributed-systems experience, plus on-site availability in Warsaw.
Operates and scales reliable edge and bare-metal infrastructure for Palantir’s production platform across isolated, on-premises, and cloud environments. The role requires at least five years managing large-scale systems, strong Linux, Kubernetes, networking, and server-hardware expertise, plus French proficiency and significant travel.
Builds and operates research infrastructure for large-scale AI model training and inference across GPU fleets. Partners with scientists and engineers to create scheduling, orchestration, and dev tooling for efficient experimentation. Requires 5+ years in distributed systems and systems programming.
Builds and supports automation, CI/CD processes for desktop/mobile apps across 1000+ repositories in multiple languages and platforms to scale to 200M users. Requires 15+ years experience in build/release engineering, OS expertise, and tools like TeamCity, Jenkins, AWS.
Designs and builds scalable infrastructure for AI products, focusing on cloud platforms, Kubernetes orchestration, CI/CD pipelines, and observability. Requires 3+ years in infrastructure engineering and bachelor's/master's in CS.
Leads incident response operations for product and engineering, serving as on-call commander to coordinate cross-functional teams, manage communications, and improve processes during high-stakes incidents. Requires 5+ years in incident management with technical depth in infrastructure and cloud systems.
Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.