Latest DevOps / SRE jobs
Job results
Staff/Senior Software Engineer building data platforms, simulation systems, or technical infrastructure for autonomous driving technology. Requires 5+ years experience, strong Python/C++/Go skills, and expertise in distributed systems or related areas.
Forward Deployed Reliability Engineer ensures stability of Palantir's mission-critical workflows by handling on-call incidents, automating solutions, and driving product improvements. Requires proficiency in Python, Java, SQL, and a technical background.
Builds and scales platform infrastructure enabling AI agents to integrate with tools like GitHub and Salesforce. Evolves APIs, manages code runtimes, optimizes performance, and collaborates with product teams on backend distributed systems.
Entry-level SiteOps engineer supporting deployment, validation, monitoring, and first-line troubleshooting of AI clusters in data center environments. Requires a relevant engineering degree or equivalent experience, familiarity with server hardware, networking, and Linux, and readiness to work hands-on in data centers.
Architects and manages multi-tenant infrastructure handling billions of requests monthly, designs build/deploy systems, drives observability, and ensures high availability. Requires 8+ years experience with TypeScript/Python, Terraform, Cloudflare, and AWS.
Staff engineer owns core platform and systems architecture, leads technical initiatives, makes high-leverage decisions for speed/reliability/scale, and partners with product/leadership. Requires 9+ years experience in production systems at staff/principal level.
Designs, builds, and operates scalable cloud infrastructure on AWS with Kubernetes and serverless tech to support AI healthcare products. Requires 4+ years experience in cloud-native platforms, IaC, automation, and reliability practices for regulated environments.
Senior Systems Engineer owns end-to-end feature domains across vehicle, autonomy, and operations for mobility-as-a-service. Leads cross-functional teams through full product lifecycle, creates requirements and MBSE models, requiring 8+ years experience and bachelor's in engineering/CS/physics.
Software engineer focused on reliability and uptime of OpenAI's GPU/HPC compute fleet through automation, monitoring tools, and performance optimization. Requires proficiency in Python/Go, Linux, networking, and data analysis skills.
Founding Infra Engineer owns Kubernetes-based platform, CI/CD pipelines, observability, security, and PostgreSQL optimization for a healthcare financial OS. Requires 5+ years Infra/SRE experience; NYC-based hybrid role.
Designs and maintains reliable infrastructure for petabyte-scale video processing across multi-cloud environments. Owns incident response, observability (Prometheus, OpenTelemetry), security, and CI/CD tooling with 3+ years experience in scalable systems.
Lead AI reliability engineering for Postman's API and agentic systems, building monitoring, observability, and automation for high availability. Requires strong SRE/DevOps background in large-scale AI infrastructure and cloud platforms.
Build and operate scalable, highly available cloud-native database infrastructure for ClickHouse’s serverless platform. The role requires 5+ years of distributed-systems experience, production programming in Go, C++, or Java, cloud infrastructure expertise, and operational on-call experience.
Builds and maintains infrastructure for software development, ML models, and AI operations including CI/CD pipelines, cloud platforms, and automation tools. Requires BS/MS in CS, experience with containerization, IaC, and programming in Python/Go/Java.
Integrates autonomy software stack for AI robotics platforms, including multi-agent systems, sensor processing, and hardware deployment across simulation, HIL, and flight environments. Requires 7+ years experience, Python/C++, CI/CD expertise, and strong systems integration skills.
Leads hands-on integration and testing of autonomy systems on small unmanned aircraft, debugging issues across hardware, software, and flight environments. Requires 4-7 years experience in robotics/aerospace testing, Python/C++ proficiency, and bachelor's in engineering or related field.
Leads integration of electrical, thermal, mechanical, and networking systems in modular data centers for high-density AI compute. Ensures compatibility with power sources, optimizes thermal management, and supports transition to liquid cooling with 6+ years systems engineering experience.
Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.
Build and extend scalable infrastructure for Databricks' data and AI platform, including multi-cloud systems and Kubernetes at massive scale. Requires 5+ years experience in Java/Scala/Go/C++/Python, distributed systems, and cloud technologies.
Lead Infrastructure Engineer shapes Atticus's infrastructure roadmap, builds security/reliability foundations, and empowers product teams with self-service tools. Requires 5+ years infra/SRE experience, GCP/Terraform expertise, and broad technical knowledge across networking, observability, and CI/CD.
Owns digital infrastructure for AI research, managing compute access, auto-scaling, resource visibility, and reproducibility using Kubernetes and observability tools. Requires systems intuition, operational rigor, and pragmatism for experimental workloads.
Owns and evolves Vercel's CI/CD infrastructure, designing scalable microservices, ensuring reliability for millions of daily builds, and leading end-to-end projects. Requires 6+ years experience with JavaScript/TypeScript, AWS, Terraform, and distributed systems.
Designs and operates scalable cloud infrastructure on AWS, focusing on Kubernetes orchestration, reliability practices, and observability for AI healthcare products. Requires 8+ years experience with IaC, containerization, and cross-team leadership.
Designs, builds, and operates scalable multi-cloud infrastructure powering ML training, inference, and data curation pipelines. Collaborates with teams on AWS-focused systems using Kubernetes and IaC tools like Terraform.
Owns and evolves Kubernetes-based infrastructure for secure, compliant AI deployments in financial services, including observability with Datadog, IaC with Terraform, and incident response. Requires 8+ years experience with Docker, K8s, AWS, and Python at scale.
Designs and operates large-scale infrastructure for secure, scalable AI agent runtimes, untrusted code execution, and multi-cloud deployments. Requires strong expertise in distributed systems, containers, Kubernetes, and security.
Build and operate scalable infrastructure powering FlexAI’s AI and PaaS platform. The role focuses on Kubernetes, infrastructure as code, CI/CD, observability, incident response, and reliability practices, requiring 4+ years of DevOps, SRE, or infrastructure engineering experience.
Develops core features for cloud-based build systems and compilers, optimizing scalability and performance. Requires expertise in build tools like Bazel/CMake, Linux/cloud infrastructure, and languages like Java/C++/Rust.
Product Reliability Engineers own end-to-end service health for Palantir's critical platforms, combining on-call incident response with forward-looking work on observability, resilience, code improvements, and infrastructure migrations. Requires strong backend coding (Java), troubleshooting skills, and ownership in dynamic environments; US security clearance eligibility needed.
Owns and builds core parts of the Forge platform end-to-end, including IDE, compiler, runtime, and infra for AI-powered English programming language. Requires staff+ engineer experience building products 0-1 with deep technical craft.
Designs and implements cloud database features for Postgres platform, ensures operational excellence, performance, and observability for large-scale Postgres/Timescale instances on Kubernetes. Requires deep Postgres expertise, Golang, and experience managing stateful workloads at scale.
Designs and builds distributed failure detection, tracing, and observability systems for large-scale AI training jobs. Requires deep expertise in performance, distributed systems, hardware, networking, and low-level software engineering.
Designs, builds, and scales core infrastructure for Harvey's AI platform, focusing on multi-cloud systems (Azure, GCP), Kubernetes, and observability. Requires 4+ years in infrastructure engineering with strong IaC and distributed systems expertise.
Staff Software Engineer on the Platform team leads technical roadmap, scales distributed systems for high QPS and throughput, prototypes AI-powered primitives, and owns core infrastructure to enable Rippling's product ecosystem. Requires 8+ years experience with deep platform expertise.
Builds and scales core infrastructure using MicroVMs, Kubernetes, AWS, and GCP to support AI agent development sandboxes. Requires 2+ years experience in infrastructure engineering and systems programming.
DevOps Engineer builds and maintains full-stack systems for billing, authentication, onboarding, and integrations on Flux's AI hardware platform. Requires 5+ years SRE/DevOps experience, TypeScript, observability tools like Datadog/Sentry, and IaC with Pulumi across GCP/AWS/Firebase.
Define and implement reliability systems for a growing AI cloud infrastructure platform, including architectural improvements, operational processes, monitoring, and incident response. Requires 5+ years production coding and 2+ years on-call experience with strong cloud skills.
Build and operate reliable infrastructure and testing services for autonomous-vehicle development. The role requires 5+ years supporting production services and SRE responsibilities, plus proficiency in Python or Golang and experience with automation, observability, CI/CD, and resilient infrastructure.
Owns infrastructure, deployment pipelines, monitoring, and security using AWS and IaC tools. Requires 3+ years DevOps experience, strong AWS and Terraform proficiency, and onsite work 4-5 days/week in NYC or SF.
Technical advisor partnering with sales to demo CodeRabbit's AI code review platform, develop custom solutions, and support pre-sales PoVs. Requires 2+ years customer-facing experience, cloud knowledge (AWS/GCP/Azure), and command line comfort.
Operates and scales reliable infrastructure for serving Cohere’s language models, building Kubernetes automation, observability, resilience, and customized production deployments. Requires 5+ years running large-scale production infrastructure and experience with distributed systems, cloud platforms, Linux, GPU workloads, and high-performance server development.
Builds and maintains shared libraries, SDKs, and AWS infrastructure (ECS, RDS, MSK) to enable faster, safer development. Architects golden paths, optimizes observability, and mentors engineers with deep Go backend and cloud expertise.
Senior Infrastructure Engineer builds and maintains scalable cloud infrastructure, CI/CD pipelines, and monitoring systems using AWS, Kubernetes, and Terraform. Collaborates with engineering teams to enhance reliability, efficiency, and security; requires 4+ years experience.
Build and maintain highly available AI cloud infrastructure virtualizing ML hardware like GB200 GPUs and BlueField DPUs, enabling self-serve Kubernetes/Slurm clusters for internal and external customers. Requires 5+ years experience with distributed systems, backend development (Golang preferred), and cloud providers.
Owns systems and tooling to optimize developer workflows, CI/CD pipelines, and local environments for faster software delivery. Requires 5+ years in DevOps/SRE, proficiency in Python/Go/JS, and CI/CD expertise.
Builds and scales core cloud infrastructure for high availability, performance, and developer productivity. Collaborates across teams to evolve backend architecture, automate workflows, and maintain production systems using Kubernetes, Docker, Postgres, and Node.
Builds developer infrastructure tools including CI/CD pipelines, build systems, and self-serve frameworks to enable fast, reliable engineering workflows. Requires 5+ years experience with Python/TypeScript, automation focus, and dev tooling expertise.
Infrastructure Engineer owns system stability, observability, and debugging at scale. Requires 3+ years experience with Go, Kubernetes, and tools like Datadog/Prometheus for production incident response.
Senior Software Engineer on the Platform team builds and optimizes core APIs and pipelines for document parsing using LLMs, integrates cutting-edge models, reduces latency/costs, and collaborates with ML engineers. Requires 5+ years experience with 2+ years in production LLMs and Python expertise.
Builds and maintains blockchain infrastructure including validators, Kubernetes deployments, and observability systems to enable fast engineering velocity. Requires expertise in Rust/Go/Python, Terraform, Prometheus/Grafana, Linux, and Ethereum ecosystem.