Latest DevOps / SRE jobs
Job results
Builds and maintains cloud infrastructure systems on Google Cloud to support engineering teams in deploying, scaling, and observing production workloads. Requires 3+ years experience in infrastructure engineering.
Builds and maintains scalable infrastructure for real-time analytics and ML workloads, focusing on reliability, automation, CI/CD, monitoring, and incident response. Requires 8+ years SRE/DevOps experience with Kubernetes, Terraform, Linux, and observability tools.
Technical visionary architecting Docker's foundational platform for accounts, billing, data, governance, and infrastructure. Drives cross-company strategy enabling enterprise growth, requiring 12+ years experience in large-scale distributed systems.
Designs and operates CI/CD pipelines and release infrastructure for multi-component systems including bootloaders, firmware, and OTA updates. Requires strong automation skills in Python/Bash, Linux expertise, and experience with build systems for embedded/consumer products.
Designs and maintains data infrastructure and pipelines to power Palantir's internal decision-making on product usage, stability, costs, and revenue. Partners cross-functionally to analyze data and deliver executive insights using Python and Palantir's platform. Requires 3+ years experience and quantitative degree.
Build and operate resilient systems for Vanta's FedRAMP and enterprise environments, define reliability frameworks, and partner with teams to ensure scalable, compliant infrastructure using AWS and modern tooling.
Hands-on network engineer deploying and validating large-scale AI datacenter fabrics, configuring switches, troubleshooting physical/optical layers, and coordinating cross-functional teams. Requires 3-7 years datacenter experience and 70-80% travel to onsite locations.
Builds scalable infrastructure platforms for data collection systems, including orchestration, dev platforms, and multi-tenant data lakes supporting exabyte-scale robotics data. Requires Python fluency and experience with distributed systems; AI/robotics background preferred.
Builds, maintains, and scales multi-cloud GPU infrastructure for AI training/inference, focusing on reliability, performance tuning, automation, and security in a fast-paced startup. Requires 8+ years SRE experience with deep Linux, cloud, and high-performance networking expertise.
Owns production infrastructure for clinical AI platform, ensuring 99.9%+ stability. Designs/scales Kubernetes-based systems, optimizes TypeScript/Python/ML CI/CD pipelines, and manages Terraform IaC in high-velocity environment.
Site Reliability Engineer enhances system observability, reliability, and availability at a prediction markets platform. Builds automation, optimizes cloud infrastructure (Kubernetes, Docker, Terraform), debugs issues, and participates in on-call rotations. Requires 4+ years software engineering experience.
Senior infrastructure engineer focused on building reliable, scalable systems for event ingestion, APIs, and billing. Requires 5+ years software engineering experience with 4+ years in infrastructure, strong debugging skills, and mentoring ability.
Designs and maintains cloud infrastructure, Kubernetes clusters for GPU/ML workloads, implements GitOps with ArgoCD and Terraform IaC. Requires 5+ years DevOps experience, Kubernetes expertise, AWS/GCP proficiency, and Python.
Senior SRE ensures reliability, scalability, and performance of legal AI platform by managing global infrastructure, leading incident response, automating operations, and optimizing costs. Requires 5+ years SRE experience, IaC expertise, cloud proficiency, and strong programming skills.
Staff SRE ensures reliability, scalability, and performance of legal AI platform across global regions. Leads incident management, automates operations, optimizes infrastructure costs, and mentors teams. Requires 10+ years SRE experience, IaC, cloud platforms, and observability tools.
Designs, builds, and scales infrastructure for a prediction market exchange including AWS, Kubernetes, high-performance APIs, and clearing systems. Requires 3+ years experience with strong fundamentals in cloud, containers, and DevOps tooling.
Develops foundational developer-experience and infrastructure systems for the Codeship team, improving CI/CD, development environments, engineering workflows, and shipping velocity. Requires 5+ years of backend or infrastructure software engineering experience, with Python or Go and scalable systems expertise.
Senior infrastructure engineer responsible for reliability, automation, observability, and operations of ClickHouse’s Postgres integration across multi-cloud environments. The role requires 7+ years of infrastructure experience, strong PostgreSQL and AWS expertise, and proficiency with Terraform, Kubernetes, and Go.
Own reliability, automation, observability, and operations for ClickHouse’s Postgres integration across multi-cloud environments. The role requires 7+ years of infrastructure or SRE experience, strong Postgres and AWS expertise, and proficiency with Terraform, Kubernetes, and Go.
Designs and builds performant, stable client infrastructure for Cursor's desktop app across macOS, Windows, and Linux. Requires deep experience in client systems, performance, reliability, and high-performance desktop applications.
Lead product vision and roadmap for Render's infrastructure platform supporting millions of developers. Requires 8+ years in product management focused on developer tools, infrastructure, or data products, with strong AI and developer experience interest.
Build and operate a scalable, secure platform supporting Ashby’s growing product and engineering organization. The role combines infrastructure engineering, reliability, developer tooling, security, database optimization, and hands-on software development with substantial end-to-end ownership.
Build and operate Ashby’s scalable platform, improving reliability, security, deployment workflows, and developer experience. The role requires strong software engineering skills, infrastructure automation experience, operational judgment, and comfort owning projects end-to-end in a distributed environment.
Staff Platform Engineer builds and scales infrastructure, optimizes compilers and databases, implements deployment tools like canary deploys and feature flags, and ensures reliability with SLOs/SLIs on AWS/Kubernetes. Requires strong coding skills in TypeScript/Node.js and handling diverse infra challenges end-to-end.
Designs and deploys high-performance network architectures for compute infrastructure and distributed AI simulation environments. Optimizes for latency, throughput, fault tolerance; requires 5+ years in high-speed networking protocols like Ethernet, InfiniBand, RDMA.
Scales infrastructure, builds automation and internal tooling, and enhances observability on GCP/GKE for a remote-first SaaS platform. Requires IaC/GitOps expertise, observability practices, and familiarity with message queues, Prometheus, and Golang.
Build and operate high-throughput infrastructure for AI engineering workloads, including sandbox runtimes, schedulers, networking, and reliability systems. The role requires deep production infrastructure experience, systems programming expertise, and strong knowledge of orchestration and isolation.
Own and scale Lovable’s developer platform, improving engineering velocity, observability, reliability, and application frameworks. The role requires 7+ years of platform experience, strong programming skills, and expertise with Kubernetes, Docker, and modern infrastructure practices.
Designs, builds, and maintains cloud-native platform services using Kubernetes for multi-tenant deployments. Develops microservices in Ruby/Java/Go, manages CI/CD pipelines with GitOps, and ensures production reliability across AWS/Azure. Requires 3-5 years experience and strong Kubernetes expertise.
Leads architecture and scaling of GCP serverless infrastructure powering a high-traffic social events app. Drives reliability, developer velocity via CI/CD and AI tooling, mentors engineers. Requires 7+ years backend/infra experience with Node.js/TypeScript.
Builds and maintains scalable infrastructure on GCP, including serverless systems, CI/CD pipelines, and observability tools for a high-growth social events app. Requires 3+ years in infrastructure/backend engineering with serverless experience and strong ownership mindset.
New grad Forward Deployed Infrastructure Engineer building, operating, and maintaining scalable, reliable infrastructure and services for Palantir's government platforms. Requires strong coding skills, comfort with production systems, automation focus, and US security clearance eligibility.
Leads reliability strategy and hands-on implementation for a large-scale, low-latency AI inference service. The role requires 7+ years in backend, infrastructure, or reliability engineering, strong backend programming skills, and deep expertise in distributed-system reliability.
Designs and operates high-performance inference and training infrastructure for ML models, focusing on GPU scheduling, distributed systems, and cloud optimization. Requires 2+ years experience with cloud platforms like AWS/GCP/Azure.
Senior SRE measures software performance, defines SLOs/SLAs, optimizes infrastructure with Temporal/Kubernetes/AWS, handles on-call, and improves developer experience/scalability for growing B2B workflows. Requires 5+ years SRE/DevOps experience.
Designs, builds, and maintains scalable infrastructure for real-time telemetry platform supporting mission-critical systems. Requires 8+ years in distributed systems, cloud environments (AWS/GCP/Azure), Kubernetes, Docker, and DevOps tools.
Builds and maintains scalable infrastructure for real-time telemetry platform using cloud, containers, and DevOps tools. Requires 3+ years in distributed systems, hands-on with Kubernetes, Docker, AWS/GCP/Azure.
Builds and maintains privacy-focused telecommunications infrastructure, including monitoring, high-availability systems, and FedRamp compliance. Requires 4+ years SRE experience, AWS expertise, and fluency in Golang/Rust/Java/Python.
Builds and maintains scalable, reliable infrastructure including testing tools, automation, and resource management platforms for AI systems. Collaborates cross-functionally to ensure high availability, performance, and fault tolerance in a fast-paced environment.
Leads physical and logical deployment of global network infrastructure for AI data centers, including rack/stack, cabling, automation with Python/Ansible, testing, and partner coordination. Requires 8+ years experience with Arista, Juniper, Mellanox, BGP/EVPN, and physical layer expertise.
Operates and automates production infrastructure for a high-scale AI inference service. The role requires production Kubernetes experience, Python or Go proficiency, observability expertise, and a focus on reliability, automation, and reducing operational toil.
Builds and scales cloud infrastructure for Render's developer platform, focusing on container orchestration, networking, storage, and AI workloads. Requires 5+ years experience with Kubernetes, IaC tools like Terraform/Pulumi/Ansible, and production systems at scale.
Site Reliability Engineer builds and maintains scalable infrastructure for ML model deployment, automates CI/CD pipelines, and ensures reliability using tools like Kubernetes and Terraform. Collaborates cross-functionally, owns projects end-to-end, and mentors juniors; bachelor's in CS or related field required.
Leads infrastructure initiatives including AWS management, dev velocity improvements, AI/LLM observability, cost optimization, and compliance. Requires 5+ years infra experience, high agency, and strong communication to shape and grow the team.
Builds backend infrastructure and core platform for AI agent cloud, including VM hypervisors, LLM sandboxes, networking, and orchestration. Requires 5+ years in distributed systems and Linux administration for onsite role in San Francisco.
Builds scalable cloud infrastructure on AWS/GCP, architects self-hostable platforms, and improves developer tooling for enterprise AI agent deployment. Requires 5+ years backend/platform experience with IaC tools like Terraform.
Build and operate large-scale research infrastructure for generative AI training, including GPU clusters, distributed systems, telemetry, and reliability tooling. The role requires deep cloud infrastructure expertise, Kubernetes, infrastructure as code, and experience operating large-scale training platforms.
Owns infrastructure for AI document platform in life sciences, building CI/CD pipelines, managing cloud systems (AWS, Kubernetes), monitoring, and light security to ensure reliability and scale.
Builds and scales core infrastructure for a high-growth AI tax platform, focusing on reliable APIs, data pipelines, observability, and fault-tolerant distributed systems. Requires 7+ years experience with Node.js, PostgreSQL, Redis, AWS, Kubernetes, and observability tools.
Builds and operates large-scale infrastructure including GPU clusters, Kubernetes orchestration, AWS batch jobs, and observability tooling to power AI search systems. Requires experience with massive-scale systems and focus on reliability and optimization.