Latest remote DevOps / SRE jobs
Job results
The Senior Cloud Operations Engineer designs, operates, and improves highly available multi-cloud infrastructure, observability, automation, and distributed systems across AWS and GCP. The role requires 5–7 years of DevOps or SRE experience, strong Kubernetes and infrastructure-as-code expertise, and participation in on-call operations.
Senior security-focused DevOps engineer responsible for improving platform security through architecture, automation, secure defaults, and observability. The role partners with engineering teams across Google Cloud, GKE, Kubernetes, identity, networking, CI/CD, and software supply chain security.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Build and operate reliable, secure cloud infrastructure across AWS and Kubernetes while automating delivery, observability, disaster recovery, and cost optimization. The role requires 8+ years in DevOps, SRE, or infrastructure engineering and strong hands-on experience with Terraform, Kubernetes, AWS, and CI/CD.
Leads Firefox performance engineering by writing code, profiling bottlenecks, improving benchmarks, and guiding cross-functional teams. Requires 7+ years of experience, strong C++ and JavaScript skills, and expertise in performance-critical software, profiling, concurrency, and systems analysis.
Build and own Webflow’s corporate cloud foundation, including landing zones, networking, security, Infrastructure as Code, GitHub delivery pipelines, self-service deployment patterns, and observability. The role requires 5+ years of platform or cloud engineering experience and strong AWS, Azure, or GCP expertise.
The Senior Infrastructure Engineer designs and operates internal data platforms and production web-service environments, develops cloud and Linux integrations, and ensures capacity and security. The role requires strong Terraform, Kubernetes, Python, Linux, networking, and cloud-provider experience.
Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Hands-on DevOps platform engineer building product-platform tooling, containerized deployments, CI/CD, and developer-experience improvements. The role requires strong Linux, containers, Kubernetes, troubleshooting, and Python or TypeScript development skills, with L2 and L3 ownership scopes available.
Build and operate reliability and resilience capabilities for a multi-cloud platform, including chaos engineering, observability-driven validation, failover testing, and resilient distributed systems. The role requires 6–9 years of software engineering experience and hands-on expertise with Java, Kubernetes, cloud platforms, and CI/CD.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.
Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.
Builds and operates CI/CD platforms across GitHub Actions, Jenkins, Kubernetes, and cloud infrastructure. The role owns GitOps delivery, reusable developer tooling, observability, reliability, and migration initiatives across distributed engineering teams.
Build and lead the evolution of Twilio’s large-scale observability platform, including telemetry pipelines, query systems, developer tooling, and standards. The role requires expertise in observability systems, distributed systems, cloud infrastructure, and modern programming languages.
Builds and operates Twilio’s global corporate network, VPN, zero-trust access, and cloud connectivity while monitoring performance and resolving incidents. The role requires substantial experience with Cisco, Palo Alto Networks, AWS networking, security protocols, and enterprise troubleshooting.
Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Designs and operates highly available GCP infrastructure and developer platforms for trading-critical systems. The role requires 5+ years of DevOps, platform, infrastructure, or SRE experience, with strong Terraform, Kubernetes, networking, CI/CD, observability, and incident-management skills.
Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Leads Ireland-based Platform Developer Enablement and SRE teams, defining platform strategy, developer self-service, reliability objectives, and observability standards. Requires senior software, SRE, or platform engineering experience, management leadership, and expertise in cloud infrastructure, Kubernetes, Terraform, CI/CD, and distributed systems.
This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.
Operates and scales shared telemetry infrastructure spanning metrics, logs, traces, alerting, dashboards, and profiling. The role requires at least three years of production engineering experience, distributed-systems troubleshooting, Infrastructure as Code, container orchestration, incident response, and on-call participation.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
The Senior SRE will design, automate, and operate high-throughput AWS and Kubernetes infrastructure, improving reliability, observability, CI/CD, and cost efficiency. The role requires 4–6 years of production SRE, DevOps, or systems engineering experience and strong Terraform, Linux, scripting, and incident-response skills.
Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.
Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.
Builds and operates core platform services, Kafka-based event streaming, and developer frameworks for high-scale distributed systems in fintech. Requires 7+ years experience with AWS, Kubernetes, and cloud-native infrastructure.
Leads the architecture, reliability, observability, and operational excellence of secure cloud and hybrid infrastructure for a mission-critical collaboration platform. Requires 5+ years in SRE, DevOps, or cloud infrastructure, with expertise in Kubernetes, Terraform, AWS, and regulated environments.
Build and operate scalable cloud infrastructure and application platforms, enabling frequent deployments, resilient systems, observability, and self-healing capabilities. The role requires strong troubleshooting, Linux and cloud experience, networking knowledge, and expertise in one or more platform engineering focus areas.
Operates and scales Kraken’s core infrastructure platforms, with a focus on OpenStack, Ceph, Linux, distributed systems, and automation. The role requires 3+ years of infrastructure or software engineering experience and supports reliable compute and storage services across cloud and on-premises environments.
Owns and evolves Webflow’s highly available, multi-cloud infrastructure, including Kubernetes, networking, infrastructure as code, observability, and AI-powered automation. The role requires 5+ years operating customer-facing cloud infrastructure and deep AWS experience.
Leads Webflow’s deployment strategy and GitOps platform, improving CI/CD reliability, progressive delivery, and developer productivity. The role requires 7+ years in DevOps, SRE, or infrastructure engineering and deep experience with Kubernetes, AWS, Docker, and infrastructure as code.
Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.
Owns and automates production infrastructure on multi-region AWS with EKS clusters, focusing on scaling, reliability, and self-healing systems. Requires deep Kubernetes, Terraform, and Linux expertise for large-scale stateful workloads.
Build and operate large-scale, geographically distributed infrastructure supporting massive file metadata, data volumes, analytics, and concurrent connections. The role requires 5+ years of software development experience and expertise in backend systems, programming, operating systems, and distributed infrastructure.
Operates and improves reliability for a trading-critical brokerage platform across cloud infrastructure, Kubernetes, observability, messaging, and PostgreSQL. Requires 4+ years of production operations experience, strong PostgreSQL fundamentals, incident response expertise, and proficiency in Go or Python.