Latest remote DevOps / SRE jobs
Job results
Operate and scale distributed storage systems like VAST and Ceph for AI/ML workloads, build Python automation tools, manage Linux bare-metal systems, and collaborate on infrastructure optimizations. Requires 5+ years in infrastructure engineering with storage expertise.
Builds and supports automation, CI/CD processes for desktop/mobile apps across 1000+ repositories in multiple languages and platforms to scale to 200M users. Requires 15+ years experience in build/release engineering, OS expertise, and tools like TeamCity, Jenkins, AWS.
Senior engineer owning reliability, automation, and evolution of core cloud infrastructure systems. Builds tooling, integrations, and observability; troubleshoots across backend, infra, and frontend. Requires Rust proficiency, database expertise, and distributed systems experience.
Owns backend infrastructure including Postgres, job queues, OpenSearch, Redis, and data pipelines at a fast-growing SaaS startup. Scales systems handling millions of events daily, focusing on reliability, tradeoffs, and production excellence. Requires deep scaling expertise in production systems.
Leads architecture and strategy for deploying optimized NLP models in high-throughput, low-latency production environments using Kubernetes and cloud platforms. Mentors engineers and designs custom customer deployments with 8+ years infrastructure experience.
Builds and maintains scalable AWS infrastructure, CI/CD pipelines with GitHub, and secure AI Ops platform for clinical AI/ML products. Requires 5+ years in DevOps/SRE, Kubernetes proficiency, Terraform expertise, and healthcare compliance experience.
Owns and scales platform infrastructure including edge/cloud services on Cloudflare, GCP, Vercel and data layers like Spanner, ClickHouse, Postgres to serve millions of LLM requests daily. Requires 5+ years in production infrastructure with cloud platforms, databases, and full-stack TypeScript expertise.
Owns and improves GitHub Actions CI pipelines, triages flaky tests across Cypress/Jest suites, manages database test infrastructure, and supports releases for a distributed engineering team. Requires CI/CD experience, proactive mindset, and strong debugging skills.
Senior infrastructure engineer owning GPU and AWS cloud infrastructure, partnering with teams to ensure reliability and build developer tools in a fast-paced AI environment. Requires deep AWS/GPU expertise and senior-level ownership.
Infrastructure Engineer builds and maintains internal engineering services, improves observability, CI/CD pipelines, and cloud infrastructure using tools like Kubernetes and AWS. Requires experience with distributed systems, infrastructure as code, and operating managed services in a remote environment.
Designs and builds distributed systems for scheduling, workflows, and storage abstractions powering research and trading in hybrid environments. Requires 5+ years experience in scalable services, modern languages like Python/Go, and Linux with strong system design skills.
Senior SRE scales and maintains high-performance research compute clusters for machine learning in finance, ensuring uptime, reliability, and observability across on-prem and cloud infrastructure. Requires 5+ years SRE/DevOps experience with HPC frameworks, IaC, and cloud services.
Designs and builds developer tools, SDKs, APIs, and CI/CD pipelines to boost engineering productivity. Requires 8+ years experience in platform tooling, proficiency in Java/Python/TypeScript/Go, and cloud platforms like AWS/GCP.
Builds, automates, and maintains scalable, secure Linux infrastructure using modern tooling. Requires 7+ years Linux sysadmin experience, deep OS/network knowledge, Python/Ruby scripting, config management (Ansible/Puppet/Chef), and monitoring (Prometheus/Grafana).
Leads automation and platform engineering for ultra-reliable AI inference infrastructure, architecting self-service GitOps pipelines, observability, and tooling to eliminate toil across datacenters. Requires 8+ years SRE experience with large-scale clusters and tools like Argo CD and Prometheus.
Owns deployment and operational infrastructure for Multigres distributed Postgres platform. Builds and maintains Go-based Kubernetes operator, architects cloud deployments on EKS/other K8s, manages storage/networking, and develops tooling for reliable scaling.
Performance Engineer optimizes infrastructure and application performance for ChatGPT and OpenAI API, focusing on latency, throughput, and efficiency at scale. Requires 7+ years in high-scale systems with expertise in profiling, tracing, and cross-layer optimizations.
Owns build and deployment automation, platform health, scaling, and self-service tools in a multi-account AWS environment. Requires 5+ years DevOps experience, deep AWS knowledge, IaC tools like Terraform/Pulumi, Docker, Python/Bash scripting.
Senior SRE ensures reliability of MongoDB's multi-tenant cloud storage layer by defining SLOs, building resilient infrastructure, optimizing performance, and participating in on-call. Requires 6+ years experience with distributed systems, Python/Go, Kubernetes, and cloud platforms.
Build and own desktop automation execution platform integrating AI agents with Windows sessions via remote protocols like VNC/RDP. Requires deep systems knowledge in OS APIs, accessibility frameworks, and automation infrastructure for reliable enterprise workflows.
Infrastructure Engineer owns platform reliability, security, and performance in a HIPAA-compliant environment handling healthcare data. Designs infrastructure with Terraform, improves CI/CD workflows, and ensures resilient systems using PostgreSQL and modern cloud tools.
Designs end-to-end infrastructure architecture for AI/ML workloads, including multi-cloud strategies, GPU orchestration, storage for massive datasets, and capacity planning to support production inference and research training at scale. Requires 7+ years in systems architecture with deep expertise in Kubernetes, AWS, and GPU infrastructure.
Senior engineer building self-service internal platform capabilities at Docker, focusing on multi-region networking, continuous deployment, EKS foundations, and AI-assisted operations to enable faster, safer provisioning for engineering teams.
Build and lead reliability engineering practices for ClickHouse Core, improving production performance, observability, incident response, and database operations. The role requires at least five years of reliability, QA, or customer-facing engineering experience, plus production SQL database and cloud expertise.
The Database Reliability Engineer will improve the reliability, scalability, performance, and incident response practices for ClickHouse Core. The role requires at least five years of reliability, QA, or customer-facing engineering experience, production database operations expertise, and strong scripting and debugging skills.
Build and lead reliability practices for ClickHouse Core, improving database performance, observability, incident response, and scalability. The role requires at least five years of reliability, QA, or customer-facing engineering experience plus production SQL database operations and cloud expertise.
The Database Reliability Engineer will improve ClickHouse Core’s reliability, scalability, performance, and production operations through observability, incident response, chaos initiatives, and database debugging. The role requires at least five years of reliability, QA, or customer-facing engineering experience and production SQL database expertise.
Build and lead reliability practices for ClickHouse Core, improving production performance, observability, incident response, and scalability. The role requires at least five years in reliability, QA, or customer-facing engineering plus production database operations experience.
Senior Platform Engineer owns cloud services, Kubernetes clusters, and developer tooling on AWS to enable engineering teams. Requires 5+ years DevOps/SRE experience, strong Kubernetes/Terraform/AWS skills, and on-call participation.
Builds and maintains platform shared services including Redpanda event streaming and auth systems using Go. Owns infrastructure, deployments, and Kubernetes ops at scale in a startup environment. Requires 5+ years experience with 1+ in Go and advanced K8s skills.
Builds and evolves internal developer platforms using Kubernetes, Terraform, and GitOps to enhance reliability, scalability, and DevEx. Requires 5+ years in platform engineering, strong AWS and cloud-native expertise, with on-call responsibilities.
Senior Infrastructure Engineer designs, builds, and operates cloud infrastructure, developer tooling, observability, and reliability systems at scale, primarily on GCP. Requires high-velocity dev experience, IaC, database scaling, workflow orchestration, and production Python coding.
Senior Infrastructure Engineer designs and owns internal platforms including TypeScript Pulumi, Kubernetes environments, and monitoring to enable reliable software shipping. Requires 5+ years experience with AWS, Kubernetes, IaC, and distributed systems.
Senior Platform Engineer owns backend services, APIs, infrastructure, CI/CD pipelines, and platform tooling to enable scalable product delivery in fintech/crypto. Requires 7+ years backend experience, Golang proficiency, Kubernetes, and cloud infrastructure expertise.
Senior SRE on the Fleet Management team develops and maintains scalable Kubernetes runtime environments, provides internal support to engineering teams, and participates in 24/7 on-call with a focus on automation and blameless post-mortems. Requires 6+ years experience with distributed systems, containerization, and Go/Python proficiency.
Senior SRE engineer builds tooling and automation to enhance production system reliability, monitoring microservices, Kubernetes, and ML platforms. Requires 6+ years in software/SRE/DevOps, proficiency in Python/Go, IaC, and observability tools.
Designs and implements large-scale public cloud infrastructure, builds complex distributed systems and microservices. Requires 10+ years experience, expert skills in performance tuning, concurrency, multiple cloud providers like AWS/GCP/Azure, and graduate degree or equivalent.
Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.
Develops production trading systems and data pipelines for machine learning in finance, designing real-time distributed systems, integrating markets, owning observability, and leading cross-team projects. Requires 5+ years in scalable distributed systems and cloud expertise.
Build and lead developer experience tools including CI/CD pipelines, test frameworks, and AI-powered dev tools to enable Vanta engineers to ship scalable products quickly. Requires technical leadership in DevEx/platform teams and expertise in scaling developer workflows.
Senior DevOps Engineer responsible for designing and operating secure, scalable AWS infrastructure, CI/CD pipelines, container platforms, and AI-assisted observability. The role requires 7+ years of DevOps, SRE, or cloud engineering experience and strong expertise in infrastructure as code, AWS, containers, scripting, and compliance.
Build and maintain infrastructure for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning for AI workloads. Requires 3+ years managing large-scale bare-metal/cloud fleets, strong Python and deep Linux expertise.
The Senior Site Reliability Engineer will improve the reliability, availability, scalability, and performance of ClickHouse Cloud by designing distributed systems, managing observability and incident response, and driving automation and chaos initiatives. The role requires 8+ years of SRE experience and hands-on Go or Python expertise.
Builds and scales Rust-based services for edge-to-cloud distributed systems integrated with Kubernetes and major clouds (AWS, Azure, GCP). Requires 6+ years experience, preferably from FAANG/cloud providers, with strong distributed systems expertise.
Builds and operates Upbound Spaces, a multi-control plane management platform using Go and Kubernetes. Troubleshoots production issues, develops features, contributes to open-source Crossplane, and ensures scalability and reliability in cloud environments.
Builds and operates hybrid AI/ML infrastructure using Kubernetes, AWS, Terraform, and Slurm for GPU workloads. Requires 5+ years SRE/DevOps experience, expert Kubernetes, and bare metal management.
Builds and scales distributed platform systems for AI-powered mortgage origination, focusing on performance engineering across AWS, Kubernetes, Kafka, databases, and observability tools. Senior-level role requiring full-stack ownership in early-stage startup.
Design, build, and operate highly available distributed cloud infrastructure for a multi-cloud SaaS platform. The role requires 5+ years of software development experience, cloud and infrastructure-as-code expertise, and strong knowledge of networking and security.
Senior software engineer responsible for designing, automating, and operating secure, highly available distributed cloud infrastructure across multi-cloud environments. Requires 5+ years of experience with scalable systems, cloud platforms, infrastructure as code, and production debugging.
Design and operate highly available, distributed cloud infrastructure for a multi-cloud ClickHouse Cloud platform. The role requires 5+ years of software development experience, strong systems and cloud expertise, and proficiency with infrastructure automation and security.