Latest remote DevOps / SRE jobs
Job results
Senior software engineer designing, deploying, and operating scalable cloud infrastructure across distributed, multi-cloud environments. Requires 5+ years of experience with distributed systems, cloud platforms, infrastructure as code, Kubernetes, networking, and cloud security.
Design, build, and operate highly available distributed cloud infrastructure for a multi-cloud ClickHouse Cloud platform. The role requires 5+ years of software development experience, cloud infrastructure expertise, and strong knowledge of networking and security.
Senior Infrastructure Engineer owns critical infrastructure decisions, builds scalable platforms using AWS and Terraform, ensures security/compliance, and mentors teams. Requires 5+ years AWS experience and expertise in monitoring tools like Datadog.
The Senior Cloud Performance Engineer benchmarks and optimizes distributed database and cloud infrastructure performance while developing chaos engineering tools and initiatives. The role requires 6+ years of experience with scalable distributed systems, programming in Go, C/C++, or Java, Kubernetes, and a major public cloud provider.
Leads performance benchmarking, optimization, capacity planning, and chaos engineering for large-scale distributed cloud database systems. Requires 6+ years of software development experience, strong cloud infrastructure expertise, and proficiency in systems programming and production debugging.
Build tools and practices for benchmarking, optimizing, and testing the resilience of large-scale distributed cloud database systems. The role requires 6+ years of software development experience, strong distributed-systems expertise, cloud infrastructure knowledge, and production debugging skills.
Senior cloud performance engineer responsible for benchmarking and optimizing distributed database and cloud infrastructure systems, while building chaos-engineering tools to improve resilience and scalability. Requires 6+ years of software development experience and expertise in distributed systems and public-cloud infrastructure.
Leads platform operations team supporting developers on GitLab and Kubernetes-based DevOps platform. Resolves deployment issues, manages on-call support, trains team members, and ensures SLOs in microservices environment. Requires BS in STEM, Linux skills, and scripting experience.
Senior Platform Engineer owns and evolves infrastructure for reliability, performance, and cost optimization at scale. Partners with engineers on debugging, observability (Prometheus, Grafana), deployment pipelines (Kubernetes, Terraform), and on-call incident response. Requires 5+ years experience including DevOps/SRE.
Build and operate scalable, highly available cloud-native database infrastructure for ClickHouse’s serverless platform. The role requires 5+ years of distributed-systems experience, production programming in Go, C++, or Java, cloud infrastructure expertise, and operational on-call experience.
Lead Infrastructure Engineer shapes Atticus's infrastructure roadmap, builds security/reliability foundations, and empowers product teams with self-service tools. Requires 5+ years infra/SRE experience, GCP/Terraform expertise, and broad technical knowledge across networking, observability, and CI/CD.
Designs and operates scalable cloud infrastructure on AWS, focusing on Kubernetes orchestration, reliability practices, and observability for AI healthcare products. Requires 8+ years experience with IaC, containerization, and cross-team leadership.
Develops core features for cloud-based build systems and compilers, optimizing scalability and performance. Requires expertise in build tools like Bazel/CMake, Linux/cloud infrastructure, and languages like Java/C++/Rust.
Designs and implements cloud database features for Postgres platform, ensures operational excellence, performance, and observability for large-scale Postgres/Timescale instances on Kubernetes. Requires deep Postgres expertise, Golang, and experience managing stateful workloads at scale.
Operates and scales reliable infrastructure for serving Cohere’s language models, building Kubernetes automation, observability, resilience, and customized production deployments. Requires 5+ years running large-scale production infrastructure and experience with distributed systems, cloud platforms, Linux, GPU workloads, and high-performance server development.
Senior Infrastructure Engineer builds and maintains scalable cloud infrastructure, CI/CD pipelines, and monitoring systems using AWS, Kubernetes, and Terraform. Collaborates with engineering teams to enhance reliability, efficiency, and security; requires 4+ years experience.
Build and maintain highly available AI cloud infrastructure virtualizing ML hardware like GB200 GPUs and BlueField DPUs, enabling self-serve Kubernetes/Slurm clusters for internal and external customers. Requires 5+ years experience with distributed systems, backend development (Golang preferred), and cloud providers.
Builds and maintains blockchain infrastructure including validators, Kubernetes deployments, and observability systems to enable fast engineering velocity. Requires expertise in Rust/Go/Python, Terraform, Prometheus/Grafana, Linux, and Ethereum ecosystem.
Technical visionary architecting Docker's foundational platform for accounts, billing, data, governance, and infrastructure. Drives cross-company strategy enabling enterprise growth, requiring 12+ years experience in large-scale distributed systems.
Build and operate resilient systems for Vanta's FedRAMP and enterprise environments, define reliability frameworks, and partner with teams to ensure scalable, compliant infrastructure using AWS and modern tooling.
Builds, maintains, and scales multi-cloud GPU infrastructure for AI training/inference, focusing on reliability, performance tuning, automation, and security in a fast-paced startup. Requires 8+ years SRE experience with deep Linux, cloud, and high-performance networking expertise.
Designs and maintains cloud infrastructure, Kubernetes clusters for GPU/ML workloads, implements GitOps with ArgoCD and Terraform IaC. Requires 5+ years DevOps experience, Kubernetes expertise, AWS/GCP proficiency, and Python.
Senior infrastructure engineer responsible for reliability, automation, observability, and operations of ClickHouse’s Postgres integration across multi-cloud environments. The role requires 7+ years of infrastructure experience, strong PostgreSQL and AWS expertise, and proficiency with Terraform, Kubernetes, and Go.
Own reliability, automation, observability, and operations for ClickHouse’s Postgres integration across multi-cloud environments. The role requires 7+ years of infrastructure or SRE experience, strong Postgres and AWS expertise, and proficiency with Terraform, Kubernetes, and Go.
Lead product vision and roadmap for Render's infrastructure platform supporting millions of developers. Requires 8+ years in product management focused on developer tools, infrastructure, or data products, with strong AI and developer experience interest.
Build and operate a scalable, secure platform supporting Ashby’s growing product and engineering organization. The role combines infrastructure engineering, reliability, developer tooling, security, database optimization, and hands-on software development with substantial end-to-end ownership.
Build and operate Ashby’s scalable platform, improving reliability, security, deployment workflows, and developer experience. The role requires strong software engineering skills, infrastructure automation experience, operational judgment, and comfort owning projects end-to-end in a distributed environment.
Staff Platform Engineer builds and scales infrastructure, optimizes compilers and databases, implements deployment tools like canary deploys and feature flags, and ensures reliability with SLOs/SLIs on AWS/Kubernetes. Requires strong coding skills in TypeScript/Node.js and handling diverse infra challenges end-to-end.
Scales infrastructure, builds automation and internal tooling, and enhances observability on GCP/GKE for a remote-first SaaS platform. Requires IaC/GitOps expertise, observability practices, and familiarity with message queues, Prometheus, and Golang.
Builds and maintains privacy-focused telecommunications infrastructure, including monitoring, high-availability systems, and FedRamp compliance. Requires 4+ years SRE experience, AWS expertise, and fluency in Golang/Rust/Java/Python.
Builds and scales cloud infrastructure for Render's developer platform, focusing on container orchestration, networking, storage, and AI workloads. Requires 5+ years experience with Kubernetes, IaC tools like Terraform/Pulumi/Ansible, and production systems at scale.
Builds and scales backend systems, cloud infrastructure, and deployment pipelines for an AI platform serving property managers. Requires 3+ years experience with Node.js, IaC tools like Terraform, and DevOps practices.
Designs and evolves production Ceph storage clusters, builds APIs and orchestration services for block/object storage using Go and gRPC. Requires experience with distributed systems, filesystems like ZFS/BTRFS, and building scalable infrastructure.
Designs and builds internal tools, AI agents, and automations for GTM, Operations, and Finance teams using full-stack skills. Partners cross-functionally to optimize sales lifecycle and drive business impact with Python, SQL, and AI expertise.
Staff SRE manages and scales Kubernetes clusters on AWS EKS, automates infrastructure with IaC tools, optimizes performance, maintains blockchain nodes and databases, and improves system reliability using monitoring tools. Requires 5+ years SRE experience with strong Kubernetes and AWS proficiency.
DevOps engineer building and maintaining decentralized Web3 infrastructure, including IPFS clusters, Arweave redundancy, cross-chain bridges, blockchain nodes, and Graph Protocol subgraphs. Requires proficiency in IPFS, Arweave, The Graph, Docker, and Kubernetes.
Leads architecture, development, and scaling of high-throughput API-based transaction processing platforms handling massive volumes. Mentors engineers and requires 8+ years experience in distributed systems, cloud platforms like AWS, and languages like Go/Python/C++.
Builds scalable distributed systems including queueing, state stores, and execution layers for developer tools platform. Requires experience with Go, distributed systems at scale, and strong engineering judgment. Works with US PST overlap.
Builds and maintains scalable cloud infrastructure using Terraform and Kubernetes, owns CI/CD pipelines and observability, supports multi-cloud deployments, and assists customers with self-hosting. Requires 5+ years in DevOps/SRE, deep AWS experience, and programming skills.
Builds and operates global datacenters bridging physical infrastructure to software, focusing on performance, resilience, and reliability. Requires experience with Golang, gRPC, Postgres, OS primitives, oncall rotations, and delivering 0→N projects remotely.
Builds low-level infrastructure software from first principles to power a cloud platform, focusing on OS primitives, distributed systems, and scalable gRPC services in Golang/Rust for high developer leverage.
Builds and secures scalable infrastructure, CI/CD pipelines, monitoring, and incident-response processes. The role requires DevOps automation experience, cloud expertise, Terraform and Ansible proficiency, and a bachelor's degree or equivalent experience.
Develops and operates scalable, secure blockchain infrastructure, managing 1000-node networks and optimizing high-throughput systems. Requires 5-10 years experience in low-level systems programming with kernel/data structures expertise and startup background.
Builds and maintains internal platform infrastructure including Kubernetes clusters, stateful services like databases and Bitcoin/LN nodes, and observability tools. Advises dev teams on integrations; requires strong Linux, networking, cloud, and systems programming expertise.
Designs and builds infrastructure tools for Lightning Network including automated channel management, liquidity optimization algorithms, monitoring systems, and network health metrics. Requires systems programming expertise in Go/C/C++, Bitcoin knowledge, and secure scalable systems experience.
Senior SRE/DevOps Engineer owns and operates AWS infrastructure and Kubernetes-based application stacks for Metabase Cloud, debugs issues, builds automation tooling, and improves deployments. Requires 5+ years experience with strong Kubernetes, AWS, Terraform, and modern languages like Python/Go.