Latest DevOps / SRE jobs
Job results
Designs and builds scalable AI infrastructure for deploying enterprise AI agents, automates LLMOps/MLOps workflows, optimizes GPU workloads, and ensures production reliability using Kubernetes and cloud platforms. Requires 5+ years in MLOps/infrastructure and deep expertise in Python and distributed systems.
Builds and scales backend systems, cloud infrastructure, and deployment pipelines for an AI platform serving property managers. Requires 3+ years experience with Node.js, IaC tools like Terraform, and DevOps practices.
Operates, deploys, and improves reliable Palantir infrastructure and services for government environments. The role emphasizes production troubleshooting, automation, scalable systems, cross-functional collaboration, and security clearance eligibility.
Builds performant, scalable infrastructure for Palantir's Foundry platform, including distributed systems, search/indexing, and container orchestration. Requires strong coding skills in backend languages like Java/Rust/Go and familiarity with storage/data systems; engineering degree preferred.
Builds and maintains scalable cloud infrastructure powering an AI-driven R&D platform for biotech, focusing on reliability, security, AI workloads, and developer tooling. Requires 5+ years in platform engineering with expertise in cloud IaC, CI/CD, and observability.
Builds and maintains cloud infrastructure abstractions for scalable, reliable product platforms like ChatGPT. Requires 5+ years in core infrastructure, Kubernetes at scale, and cloud abstractions; onsite in San Francisco with on-call duties.
Designs and evolves production Ceph storage clusters, builds APIs and orchestration services for block/object storage using Go and gRPC. Requires experience with distributed systems, filesystems like ZFS/BTRFS, and building scalable infrastructure.
Designs and operates secure multi-tenant platforms using Kubernetes and AWS, managing provisioning, monitoring, CI/CD with Terraform and GitHub Actions, and ensuring security/compliance. Requires 4+ years DevOps experience with strong fundamentals in listed technologies.
Designs, builds, and maintains global cloud infrastructure for Convex's backend platform, focusing on performance, reliability, and scalability. Requires 5+ years backend experience with systems at scale, strong architecture skills, and on-call ownership in a startup environment.
Designs and builds internal tools, AI agents, and automations for GTM, Operations, and Finance teams using full-stack skills. Partners cross-functionally to optimize sales lifecycle and drive business impact with Python, SQL, and AI expertise.
Design, build, and operate a multi-tenant caching platform powering OpenAI's inference, identity, and products. Requires 5+ years in distributed systems with deep Redis/Memcached expertise and Kubernetes experience.
Builds, deploys, and optimizes large-scale Kubernetes and Slurm clusters for AI training and inference. Requires 3+ years in ML infrastructure, expert Kubernetes/Slurm admin, Python/C++, and PyTorch experience.
Platform Engineer building high-scale distributed document indexing and search systems for an AI platform serving top financial institutions. Requires 5+ years experience in backend/distributed systems with Python/Java/Go and cloud infrastructure.
Staff Engineer leads architecture rewrites, builds scalable cloud infrastructure with IaC, optimizes data systems, and establishes DevOps practices for AI-powered property management platform. Requires 5+ years backend experience with Node.js and cloud expertise.
Builds and maintains secure, scalable infrastructure for AI-powered agents, deploys ML workloads, manages cloud environments, IaC, CI/CD, and ensures observability and security on AWS/Azure/GCP/Kubernetes.
Build and improve large-scale infrastructure platforms supporting Palantir products and developer workflows. The role is suited to a new graduate with an engineering background, programming experience, and familiarity with cloud, storage, frontend, and infrastructure technologies.
Staff SRE manages and scales Kubernetes clusters on AWS EKS, automates infrastructure with IaC tools, optimizes performance, maintains blockchain nodes and databases, and improves system reliability using monitoring tools. Requires 5+ years SRE experience with strong Kubernetes and AWS proficiency.
Build and scale managed Kubernetes and AI training clusters, developing operators, controllers, and infrastructure using Go, Terraform, and GCP. Design reliable, high-performance systems competing with GKE/EKS, with 5+ years experience required.
Forward Deployed Site Reliability Engineer responsible for building, operating, and maintaining scalable infrastructure in air-gapped on-prem environments for US Government customers. Requires 4+ years Linux admin experience, hardware/networking knowledge, scripting skills, 50% travel availability, and active Top Secret clearance.
Senior Platform Engineer designs, builds, and scales cloud infrastructure (80% focus) including Kubernetes and GCP, implements CI/CD pipelines, security practices, and developer tools to boost engineering velocity and maintain compliance at hyperscale.
Senior SRE focuses on ensuring reliability, availability, and performance of distributed database systems in cloud-native environments. Requires 4+ years experience with Kubernetes, Docker, cloud platforms (AWS/GCP/Azure), IaC tools, and scripting in Python/Go/Java.
DevOps engineer building and maintaining decentralized Web3 infrastructure, including IPFS clusters, Arweave redundancy, cross-chain bridges, blockchain nodes, and Graph Protocol subgraphs. Requires proficiency in IPFS, Arweave, The Graph, Docker, and Kubernetes.
Builds and owns scalable platform infrastructure on AWS and Kubernetes, including deployment pipelines, automation, and observability tools. Provides technical leadership to enable efficient engineering workflows and production reliability.
Leads architecture, development, and scaling of high-throughput API-based transaction processing platforms handling massive volumes. Mentors engineers and requires 8+ years experience in distributed systems, cloud platforms like AWS, and languages like Go/Python/C++.
Builds and scales distributed systems infrastructure for high-performance CI workloads, focusing on performance/reliability in filesystem, network, and storage. Requires 2+ years experience with C/C++/Go/Rust/Zig and onsite work in NYC.
Designs and engineers distributed systems for compute, scheduling, and orchestration of ML and ETL pipelines handling large-scale video data. Requires 3+ years building data infrastructure, proficiency in Go/Python, and experience with petabyte-scale pipelines and CI/CD.
Builds scalable distributed systems including queueing, state stores, and execution layers for developer tools platform. Requires experience with Go, distributed systems at scale, and strong engineering judgment. Works with US PST overlap.
Builds and maintains scalable cloud infrastructure using Terraform and Kubernetes, owns CI/CD pipelines and observability, supports multi-cloud deployments, and assists customers with self-hosting. Requires 5+ years in DevOps/SRE, deep AWS experience, and programming skills.
The Senior DevOps Engineer will build and operate scalable, secure multi-region cloud infrastructure, networking, monitoring, and delivery pipelines across cloud and private environments. The role requires 5+ years of DevOps or cloud engineering experience, strong Linux and AWS expertise, and participation in production on-call support.
Senior/Staff SRE maintains high-availability production infrastructure on AWS EKS with Terraform, manages sharded MongoDB clusters, participates in on-call rotation, and engages with customers to enhance reliability for 1B+ daily API calls.
Builds and maintains foundational systems, tools, and processes to boost developer productivity and engineering velocity at OpenAI. Requires 5+ years engineering experience, including infrastructure tooling, with core tech like Kubernetes, Python, and Terraform. Onsite in SF HQ.
Builds scalable data and ML infrastructure supporting multi-cloud and deployment models. Partners with founders and engineers on core platforms, tooling, and research features for reliable production systems.
Builds and maintains infrastructure components for ML inference platform using Python and Go. Implements Kubernetes deployments, monitoring systems, and resource management for efficient model serving, requiring Kubernetes knowledge and ML basics.
Build and operate low-latency infrastructure and distributed backend systems powering a high-throughput search stack. The role requires strong cloud, Linux, automation, observability, and systems-language experience across AWS, Kubernetes, Rust, and Go.
Designs, implements, and operates infrastructure systems for model training and deployment on a massive GPU fleet. Requires experience with hyperscale compute, Kubernetes, public clouds like Azure, and strong programming skills.
This engineer automates deployment and operation of production databases, manages migrations, improves reliability, and participates in incident response and on-call rotations. The role requires programming experience, cloud and orchestration expertise, database knowledge, and a strong Linux foundation.
Builds and operates global datacenters bridging physical infrastructure to software, focusing on performance, resilience, and reliability. Requires experience with Golang, gRPC, Postgres, OS primitives, oncall rotations, and delivering 0→N projects remotely.
Builds and scales massive Kubernetes clusters for OpenAI's frontier supercomputers, automates bare-metal provisioning, and ensures reliability across data centers for AI model training. Requires expertise in distributed systems, Kubernetes operations, and infrastructure automation.
Builds and optimizes core platform services including search, indexing, infrastructure, auth, billing, and developer tooling. Collaborates with AI teams and requires 3+ years full stack experience with strong systems programming skills.
Designs, builds, and maintains high-performance distributed systems for a serverless AI platform. Requires 5+ years experience in production code, large-scale systems, cloud, OS foundations, and performance optimization; onsite in NYC.
Develops APIs and Kubernetes controllers to enable containerization across Palantir platforms like Foundry, Gotham, Apollo for diverse infrastructures. Requires 3+ years software experience, Go/Java proficiency, system design, and bachelor's in CS.
Senior Site Reliability Engineer building agentic AI platforms, MCP servers, and automation tools to achieve zero KTLO. Drives SecDevOps culture with focus on security, reliability, observability while partnering on product roadmap and incident response. Requires 6+ years SRE/DevOps experience plus deep expertise in AWS, EKS, Terraform, Kubernetes.
Builds low-level infrastructure software from first principles to power a cloud platform, focusing on OS primitives, distributed systems, and scalable gRPC services in Golang/Rust for high developer leverage.
Builds and operates scalable infrastructure systems including Kubernetes clusters, distributed databases, and cloud services to support AI music platform at consumer scale. Requires 5+ years experience in infrastructure engineering with strong ownership and scaling expertise.
Builds high-performance data processing systems for real-time analysis of massive unstructured data at enterprise scale, focusing on storage, indexing, query optimization, and Braintrust's btql language. Requires expert systems programming in C++ or Rust, concurrency, databases, and OS knowledge.
Senior Software Engineer builds and maintains APIs for Kubernetes-based network infrastructure, managing traffic flows across zero-trust clusters. Requires 3+ years experience, Kubernetes expertise, Go/Java skills, and networking knowledge.
Builds and runs end-to-end ML inference infrastructure for generative AI image models across mobile and web platforms. Requires expertise in distributed systems, Kubernetes, Kafka, Redis, GPUs, and multi-cloud environments.
Site Reliability Operations Analyst streamlines workflows, stabilizes projects, and acts as first responder for Palantir deployments to free engineers for technical work. Requires 3+ years project management, travel willingness, and strong judgment under pressure.
Builds and secures scalable infrastructure, CI/CD pipelines, monitoring, and incident-response processes. The role requires DevOps automation experience, cloud expertise, Terraform and Ansible proficiency, and a bachelor's degree or equivalent experience.
The Senior Site Reliability Engineer will optimize Kubernetes and GPU infrastructure for cost, throughput, and reliability while enabling backend teams through tooling and instrumentation. The role requires 4–5 years of systems experience, backend development depth, Kubernetes expertise, GCP and Terraform fluency, and hybrid on-premises infrastructure experience.