Latest remote DevOps / SRE jobs
Job results
Design, build, and operate scalable distributed systems and edge networks on AWS to handle Figma's growing customer traffic and services. Requires 4+ years building infrastructure at scale, experience with TypeScript or Go, and distributed/traffic systems.
Design, build, and operate petabyte-scale Product Metrics systems that process massive event volumes with strong reliability, performance, and availability. Requires 5+ years of distributed-systems experience, Golang expertise, cloud-platform experience, and familiarity with Kubernetes and infrastructure as code.
Senior Cloud Engineer responsible for designing, operating, and improving petabyte-scale Product Metrics systems built with Go, Kubernetes, and ClickHouse. The role requires 5+ years of experience with scalable distributed systems and strong reliability, performance, and production debugging skills.
Own stability and deployment of PostgreSQL products. Package software with Nix, manage upgrades, optimize CI/CD, and resolve production issues. Requires 3+ years PostgreSQL experience and Nix proficiency.
Build internal developer platform, tooling, and automation to accelerate engineering velocity. Focus on CI/CD pipelines, test infrastructure, build systems, and metrics to help engineers ship faster and more reliably.
Builds automation and testing infrastructure for a platform, enabling reliable validation across environments and configurations. The role requires 5+ years of software engineering experience, strong Python and TypeScript skills, and expertise in CI/CD, containers, cloud platforms, and developer tooling.
SRE embedded in Service Operations to establish reliability practices, frameworks, and feedback loops across engineering teams. Focus on SLOs/SLIs, ORR processes, incident-to-improvement pipelines, and influencing without authority in a distributed environment.
Operate and scale a cloud-native CTV advertising platform on AWS and Kubernetes. Focus on reliability, GitOps workflows, infrastructure automation, observability, and incident response.
Design, build, and operate AI-powered engineering tools and developer productivity platforms. Focus on AI pairing pipelines, automated workflows, and internal tooling to accelerate engineering velocity.
Lead technical direction for Komodo's core control plane (KMC/PSS, identity, subscriptions) and App Builder/Connector. Architect platform primitives, APIs, and AI tooling in a multi-tenant SaaS environment.
Senior Infrastructure Engineer building and operating AWS cloud infrastructure for healthcare data platform. Requires Python, Terraform, CI/CD expertise, and big data tools experience.
Build and operate foundational data infrastructure including Airflow, Flink, DynamoDB, and RDS using Terraform and Kubernetes. Requires 2-4 years of infrastructure/platform experience and strong Python skills.
Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.
Owns weekly releases and phased production rollouts for a globally distributed remote execution platform. The role focuses on Terraform-managed infrastructure, CI health, incident coordination, operational reliability, and continuous improvement of release processes.
Senior IC role owning reliability, performance, and scalability of PostgreSQL (Aurora), OpenSearch, Redis, and CDC pipeline to Snowflake. Sets standards for ORM usage, migration safety, and observability at scale.
Senior SRE owning reliability, monitoring, and automation for Coinbase's AI infrastructure on AWS and Kubernetes. Requires 5+ years cloud automation experience and strong incident response skills.
Staff SRE on the IT Operations team owning reliability, automation, and observability for Coinbase's AI infrastructure on AWS and Kubernetes. Requires 8+ years of cloud infrastructure experience and strong incident response leadership.
Leads the design and production adoption of self-service infrastructure platforms, multi-region cloud foundations, and reliable delivery workflows. The role requires 8+ years of hands-on software or platform engineering experience, strong Go or comparable programming skills, and deep expertise in a core infrastructure domain.
Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.
Principal Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 8+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.
Seasoned Platform Engineer designing and maintaining scalable distributed systems and infrastructure on AWS. Builds foundational patterns, IaC, CI/CD pipelines, and observability for Python services using Kubernetes and serverless. Requires 6+ years of platform engineering experience.
Design, deploy, and secure highly available ClickHouse Cloud platforms across regulated cloud, hybrid, on-premises, and disconnected environments. Requires 6+ years of distributed-systems experience plus expertise in Kubernetes, infrastructure automation, databases, cloud platforms, and Go or Python.
As a Senior Developer Productivity Engineer, you will own the build, test, and deployment processes for a 50+ person engineering team. You will improve monorepo productivity, drive excellence in testing, and support multi-cloud/multi-region infrastructure to enable fast and safe shipping.
As a Senior/Staff Distributed Systems Engineer, you will design and evolve core control, data, and observability systems for LiveKit's platform, focusing on latency, availability, and operational simplicity. You will implement resilient architectures and build tools to enhance reliability and developer velocity.
As a Senior Infrastructure Engineer, you will be responsible for architecting and maintaining scalable, reliable cloud infrastructure, leading incident management, and improving operational processes. This role requires strong proficiency in AWS, infrastructure-as-code, and experience with monitoring and observability tools.
As a Senior Infrastructure Engineer, you will design, implement, and maintain cloud infrastructure on GCP, focusing on CI/CD pipelines, Kubernetes, and Terraform. This role requires a strong background in DevOps/SRE and a passion for building foundational systems in a fast-paced environment.
Seeking a Senior DevOps/Site Reliability Engineer to build, operate, and scale reliable cloud-native infrastructure and distributed data platforms. This role requires expertise in Kubernetes, cloud infrastructure, observability, automation, CI/CD, and incident management.
Leads release management and cloud infrastructure initiatives for a highly available SaaS platform, mentoring the DevOps/Release team and improving automation, deployment, observability, security, and reliability. Requires 8+ years of enterprise SaaS development or operations experience and 6+ years with highly available cloud applications.
Hands-on technical lead owning cloud infrastructure, CI/CD pipelines, and deployment automation for a multi-tenant SaaS platform. Architect and build production systems using AWS, Terraform, CloudFormation, and Python in a GxP-regulated environment.
Design and operate multi-regional infrastructure for a high-traffic global gaming platform, owning on-call, monitoring, automation, and hybrid cloud migration to ensure reliability at massive scale.
Senior SRE responsible for building and operating reliable, scalable infrastructure on AWS with Kubernetes and Terraform. Focus on observability, incident response, automation, and mentoring engineers on SRE best practices.
Staff Software Engineer building and scaling Pinterest's observability platform (metrics, logs, traces) for massive distributed systems. Requires 7+ years distributed systems experience, strong data engineering skills, and expertise with modern observability tools.
Design, build, and operate Render's core networking stack across data centers and clouds, focusing on Kubernetes and Linux internals, traffic routing, and hybrid connectivity at scale.
Founding Senior Platform Engineer building and owning AWS cloud infrastructure, reliability, observability, security/compliance (SOC 2, Vanta), and release tooling for a fintech platform serving banks and credit unions.
Lead design and evolution of core platform services, APIs, and shared primitives that power every product surface and AI agent. Drive technical standards and architecture across SaaS, enterprise, and government environments while mentoring engineers.
The Senior Database Reliability Engineer will operate and improve large-scale PostgreSQL environments in AWS, focusing on availability, performance, security, backups, recovery, and incident response. The role requires 7+ years of database administration or engineering experience, strong Linux and SQL skills, and production cloud database expertise.
Owns and improves the Go-based Terraform provider for Supabase's developer platform, focusing on reliability, lifecycle management, schema evolution, and user migrations. Requires 5+ years experience with Go, deep Terraform expertise, and strong testing/CI/CD skills.
Develops and optimizes Supabase Edge Runtime, a Rust-based Deno host for global edge TypeScript functions. Evolves infrastructure for low-latency compute, integrates with Supabase stack, and improves developer tools. Requires 5+ years backend/systems experience with Rust, TypeScript, and scalable infra.
Builds and maintains AI platform infrastructure for agentic systems, including connectors, execution environments, governance, and self-service tools to enable safe, scalable AI use across engineering and business teams. Requires 8+ years experience with LLM agents, GCP, and cloud-native tech.
The Staff SRE Engineer will drive reliability, scalability, observability, and operational efficiency across highly available cloud and distributed systems. The role requires 5+ years of SRE, DevOps, or platform experience, advanced Kubernetes expertise, strong automation skills, and leadership in incident management.
Builds and maintains bare metal provisioning, orchestration engine, and internal tools for Railway's infrastructure platform. Optimizes fleet efficiency and develops resilient services using Golang/Rust, Ansible, and Terraform for distributed systems.
Develops and optimizes Kubernetes-based infrastructure for high-performance AI inference services, including deployment, scaling, debugging, and integration with ML workflows. Requires Master's in CS and 1+ year experience with Docker, Kubernetes, Python, and related tools.
Builds and maintains core backend systems, infrastructure, and automation for platform reliability and scalability. Owns troubleshooting across stack, integrations with external services, and observability. Requires strong Rust, systems engineering, and distributed systems experience.
This senior engineer will operate and optimize GPU and accelerator infrastructure for AI training, inference, evaluation, and experimentation. The role requires 5+ years of infrastructure experience, production GPU cluster operations, strong systems fundamentals, and expertise in serving, observability, reliability, and compute-cost optimization.
Builds and scales reliable cloud infrastructure, deployment systems, observability, and developer tooling to support mortgage market operations. Requires experience with strongly typed languages, PostgreSQL, Kubernetes, and major cloud providers.
Builds and scales platform infrastructure on AWS EKS with GitOps via ArgoCD, manages CI/CD with GitHub Actions, drives observability using Datadog/Sentry/CloudWatch, and ensures reliability through SLOs and incident response. Requires 3+ years SRE/DevOps experience and Kubernetes expertise.
Builds and maintains Kubernetes-based infrastructure for managed TimescaleDB cloud services, develops Go microservices and operators, automates database operations, and ensures platform scalability and reliability. Requires 3+ years experience with Go, Kubernetes, and PostgreSQL.
Senior SRE responsible for operating and improving highly available cloud platforms, distributed data systems, observability, incident response, and deployment automation. Requires 5+ years of SRE, DevOps, or platform engineering experience, advanced Kubernetes expertise, cloud proficiency, and strong Python and Bash skills.
Build high-scale observability pipelines and alerting engines handling 1M+ RPS for logs/metrics, develop Golang/Rust gRPC services and APIs, and manage immutable infrastructure with Terraform/Ansible in a distributed systems environment.
Owns GPU diagnostics, validation workflows, and automation for bare-metal infrastructure supporting AI/ML workloads. Requires 5+ years in systems engineering with strong Linux, Python, and NVIDIA tools expertise.