Latest remote DevOps / SRE jobs
Job results
Build and operate resilient, large-scale platform services across AWS and Azure, managing Kubernetes clusters, infrastructure automation, deployment orchestration, and observability. The role requires 4+ years of software engineering experience with Kubernetes, Terraform, Go, and distributed systems.
Leads a storage engineering team while architecting and operating highly available Linux-based storage, datacenter, and data-protection infrastructure. The role requires deep Ceph experience, PB-scale archiving and backup expertise, hands-on troubleshooting, and team leadership.
Own platform initiatives that improve cloud scalability, reliability, automation, and developer productivity. The role requires 5+ years of cloud infrastructure experience plus hands-on expertise with infrastructure as code, Kubernetes, CI/CD, programming or scripting, and AI-assisted engineering.
Fluidstack is seeking a Principal Operations Engineer, Electrical to be the senior technical authority for electrical infrastructure across their hyperscale AI data center portfolio. This role involves leading site assessments, driving technical readiness, reviewing designs, and feeding operational learnings back into the design and manufacturing organization.
Leads the reliability, observability, incident management, QA, and release engineering functions for a cloud-native healthcare SaaS platform. Requires 12+ years of SRE, infrastructure, or platform engineering experience and significant engineering leadership experience.
Improves the reliability, observability, and deployment safety of Block’s critical platforms while leading high-severity incident response and on-call operations. Requires strong production reliability experience, incident management skills, and 5+ years of software development experience.
Owns the systems that build, package, update, secure, and observe software running on customer-hosted edge appliances. The role requires 5+ years of infrastructure engineering experience, deep Linux expertise, release and rollback ownership, and proficiency with systems languages, packaging, IaC, containers, and CI/CD.
Build and optimize developer productivity infrastructure spanning local development, CI, testing, code review, AI-assisted workflows, and engineering metrics. The role requires 7+ years of software engineering experience and proficiency in Go, TypeScript, Python, Rust, and shell.
Build and operate Kubernetes-based compute orchestration infrastructure used across Coinbase, while developing developer tooling, automation, and AI-enabled workflows. The role requires 5+ years of software engineering experience, including substantial experience operating Kubernetes or comparable systems in production.
Senior software engineer focused on improving production reliability, deployment safety, configuration and secrets management, and scalability across Coinbase’s service environment. The role requires 5+ years of experience with distributed systems, Ruby or Go, Terraform, cloud platforms, and observability tools.
Senior Site Reliability Engineer responsible for operating and scaling a cloud-native CTV advertising platform on AWS and Kubernetes. Requires deep Kubernetes and AWS expertise, GitOps with ArgoCD, IaC with Terraform, CI/CD, observability, and incident response experience.
Staff Software Engineer builds scalable security data platforms and automates infrastructure for Okta's Public Sector using Python, Terraform, and cloud technologies. Requires 8+ years experience in security engineering and data pipelines.
Owns the architecture and delivery of complex IT automation workflows and shared platform components. The role requires at least five years of IT, automation, or systems engineering experience, plus expertise with workflow orchestration, Git-based development, identity tools, and collaboration APIs.
Senior Software Engineer responsible for building and operating the reliability, scalability, efficiency, and observability infrastructure of Starburst Galaxy. The role requires cloud architecture and orchestration experience, Java and TypeScript development, and infrastructure-as-code expertise.
Senior infrastructure engineer building scalable abstractions, platform tooling, and the full infrastructure stack (AWS to product platform) that accelerates Pilot's R&D and business growth. Requires 5+ years software engineering experience, production Python, Terraform/AWS, frontend familiarity, mentoring ability, and strong collaboration/communication skills.
Senior Software Engineer building the platform that powers secure, reproducible builds and distribution of open-source libraries across Java, JavaScript, Python/AI/ML ecosystems at Chainguard. Lead services, automation, pipelines, and developer tooling with a focus on supply-chain security, CVE remediation, and AI-assisted patching.
Owns and evolves cloud infrastructure, CI/CD, observability, developer tooling, and FinOps governance for a distributed engineering organization. Requires 6+ years in DevOps, SRE, or software engineering, plus strong Kubernetes, cloud, security, and automation experience.
Own reliability, SLOs, and observability for Supabase's deployment pipelines, control plane, and release systems as part of the Release Engineering / SRE team. Drive safe, observable deploys, disaster recovery, incident response, and toil reduction in a fully remote, async environment.
Platform Engineer building core services, developer tooling, and frameworks for high-scale distributed systems and agentic applications at a consumer fintech company. Requires 2-9 years experience with AWS, distributed systems, CI/CD, and building platforms that accelerate engineering velocity.
Senior Software Engineer building Chainguard's internal Developer Platform "Factory", including monorepo CI/CD pipelines, Agentic AI platform for automated changes, and paved-road build infrastructure to reduce developer toil and accelerate secure artifact delivery.
Infrastructure Engineer building and securing Kubernetes-based platforms, cloud-native deployments, and air-gapped appliances for military planning software. Requires 5+ years production infrastructure experience, deep Kubernetes and cloud expertise, security fundamentals, and full-stack engineering skills in languages like Go or Python.
Senior Site Reliability Engineer owning production reliability and enterprise customer implementations for Kong's fast-growing Managed Gateways SaaS product across AWS, GCP, and Azure. Requires deep Kubernetes, cloud-native, and Golang expertise plus customer-facing technical leadership.
Design and operate secure, scalable ClickHouse Cloud platforms across regulated cloud, hybrid, on-premises, and disconnected environments. The role requires 6+ years of distributed-systems experience and strong Kubernetes, infrastructure automation, cloud, database, and security expertise.
Staff Software Engineer leading technical direction for Instacart's Bazel build system (remote execution, caching, performance) and Go platform (frameworks, libraries, patterns). Hands-on role driving initiatives to improve build times, CI reliability, and developer velocity for 1000+ engineers. Requires 10+ years experience with deep Bazel and Go expertise.
Build and operate scalable cloud-native infrastructure for ClickHouse’s serverless database platform, including Kubernetes-based management, metrics systems, and distributed data-plane capabilities. Requires 5+ years of software development experience and production expertise with cloud platforms and Go, C++, or Java.
Build and operate ClickHouse’s cloud-native database infrastructure, including Kubernetes-based management, metrics systems, and highly available distributed services. The role requires 5+ years of software development experience, production expertise in Go, C++, or Java, and experience with public cloud and data infrastructure.
Staff Engineer on the People Technology team building and maintaining scalable automations, agents, and integrations across Workday, Greenhouse, Slack, Workato, and GCP. Requires strong software engineering skills applied to HR systems with heavy use of AI/LLMs to eliminate manual work.
Leads a global 24×7 DevOps and Customer Enablement organization supporting highly available SaaS operations. The role requires 12+ years in DevOps, SRE, cloud operations, or production engineering, including substantial experience managing engineering teams and driving reliability, incident response, automation, and customer outcomes.
Senior platform engineer responsible for shipping and operating scalable infrastructure and services supporting financial products. Requires at least five years building non-trivial products or services, with backend experience and strong collaboration, ownership, and communication skills.
Senior Elasticsearch Engineer owning full lifecycle of massive-scale search and analytics platform at Chess.com: capacity planning, architecture, performance tuning, incident response, and Elasticsearch-to-OpenSearch migrations on bare-metal Kubernetes. Requires 7+ years operating Elasticsearch at scale with deep internals knowledge.
This role builds and operates scalable platform infrastructure, automates operational and database workflows, and improves reliability, observability, and deployment practices. It requires at least eight years of platform experience, production systems expertise, Kubernetes, cloud, Linux, and SQL datastore experience.
Senior frontend platform engineer strengthening React, TypeScript, and Vite foundations, optimizing testing, CI/CD, automation, and observability to enable fast, high-quality frontend development across the organization. Requires 6-8+ years experience, leadership of complex projects, and deep frontend tooling expertise.
Platform engineer responsible for designing, building, and operating cloud infrastructure on AWS and Kubernetes. Focus on developer tooling, automation with AI, performance tuning, security, observability, and on-call support for a SaaS platform serving private capital markets. Requires 5+ years in DevOps/SRE/Platform roles.
Software Engineer L2 responsible for evolving and maintaining Twilio's Compute infrastructure, including VM orchestration, AWS Auto Scaling Groups, hardened AMIs, secure container images, and automation of operational tasks in a remote-first environment.
Staff Platform Engineer building self-service infrastructure tools and platforms on Kubernetes and GCP. Requires 10+ years cloud/software experience, strong Go and Kubernetes expertise, and mentoring skills.
Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.
Own and scale AWS infrastructure, CI/CD pipelines, GitOps automation, observability, and cloud security for high-throughput payment systems. The role requires 5+ years of DevOps or infrastructure experience plus strong Kubernetes, automation, and SRE expertise.
Senior Platform Engineer owning AWS/EKS infrastructure, IaC with Terraform, GitOps, security, observability, and data systems for a fast-growing expert network marketplace. Requires 5+ years production Kubernetes/AWS experience, strong IaC and security skills.
Operate and expand banking/financial data integrations at Stripe while building and reviewing LLM-powered agentic automation. Requires strong technical problem-solving across code, data, logs, and partner coordination plus experience building with AI tooling.
Senior Site Reliability Engineer responsible for designing, implementing, and operating observability systems for complex cloud platforms. Focus on improving reliability, resilience, reducing toil through automation, and participating in on-call duties. Requires strong experience with IaC, cloud technologies, observability tools, and Linux engineering.
Senior Site Reliability Engineer building and scaling Teleport's secure SaaS cloud infrastructure. Focus on global scaling, observability, automation, incident response, and on-call in a security-first environment using Go and Kubernetes.
Senior Site Reliability Engineer partnering with development teams to design, build, and operate scalable GCP and Kubernetes infrastructure at high request volumes (75k RPS). Focus on observability, reliability, CI/CD, edge modernization, on-call, and mentoring to maintain platform performance and uptime.
Staff SRE embedding early with product/platform teams to design reliability, observability, and scalability into systems across multi-cloud environments. Define SLIs/SLOs, build golden paths with Terraform, lead incidents/postmortems, influence org-wide standards, and share real on-call rotation.
Senior Cloud Engineer responsible for designing, operating, and improving petabyte-scale Product Metrics systems built with Golang, Kubernetes, and ClickHouse. The role requires 5+ years of experience with scalable distributed systems and production ownership.
Build and operate shared Kubernetes (EKS) and AWS cloud infrastructure powering Upstart's product and ML workloads. Requires 3+ years Kubernetes production experience plus strong AWS, IaC, and GitOps skills.
Staff Software Engineer leading automation traffic management (bots/crawlers) and rate limiting at Pinterest's Edge (CDN, TLS, DNS, proxies). Design/implement Envoy-based L7 logic in C++/Go/Python; own roadmap, mentor, and drive reliability for 600M+ user platform.
Senior Platform Engineer building CI/CD pipelines, developer productivity tools, and Kubernetes infrastructure at Mux to accelerate engineering velocity from local dev to production. Requires strong Go engineering, production Kubernetes experience, networking fundamentals, and cross-team partnership.
Owns and automates multi-cloud, multi-region SaaS infrastructure, focusing on reliability, observability, performance, and incident response. Requires 5–7+ years of DevOps or infrastructure engineering experience, cloud tooling expertise, and business-level English and Japanese.
Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.
Staff engineer leading Pinterest's service communications platform. Architect and scale Envoy-based service mesh, mTLS identity, traffic optimization, and multi-language RPC frameworks for reliable, secure, high-volume service-to-service communication.