Latest DevOps / SRE jobs
Job results
Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
The DevOps Engineer builds and scales automation, Kubernetes environments, CI/CD pipelines, and operational tooling for research and trading platforms. The role requires at least two years of Linux-centric DevOps, infrastructure, or SRE experience, with expertise in Kubernetes, pipeline engineering, and observability.
Leads design, deployment, and operation of secure distributed cloud systems for public-sector and air-gapped environments. Requires active or obtainable TS/SCI clearance with polygraph, U.S. citizenship, and 7+ years of production experience.
Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
Build and operate Grafana’s physical infrastructure platform, including bare-metal environments, Kubernetes clusters, networking, scheduling, and autoscaling. The role requires datacenter and software-operations experience, with strong skills in Kubernetes and infrastructure automation using tools such as Go, Terraform, and Crossplane.
Build and maintain Linux-based NAS appliance software, focusing on Python systems development, storage and filesystem performance, packaging, automation, and production debugging. The role requires 7–10 years of systems or platform engineering experience and deep Linux expertise.
Leads the technical vision and development of a scalable observability platform, improving reliability, performance, incident response, and engineering productivity across Postman. Requires 10+ years of software engineering experience with distributed systems, cloud-native architectures, and production operations.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
The Network Engineer will design, operate, troubleshoot, and automate secure enterprise and cloud networks across offices, labs, and production services. The role combines network architecture and lifecycle planning with incident response, operational delivery, and automation using APIs, Infrastructure as Code, Git, testing, and CI/CD.
Build and operate Vercel’s low-level compute infrastructure, including storage, state, clusters, and distributed workloads. The role requires 5+ years of software engineering experience, strong Go skills, and deep expertise in Linux, virtualization, schedulers, and reliable distributed systems.
Build and operate delivery systems and infrastructure automation using Go and TypeScript/Node.js. The role combines production support, backend development, Kubernetes and cloud tooling, CI/CD, GitOps, and infrastructure-intelligence automation.
Builds and operates CI/CD platforms across GitHub Actions, Jenkins, Kubernetes, and cloud infrastructure. The role owns GitOps delivery, reusable developer tooling, observability, reliability, and migration initiatives across distributed engineering teams.
Build and operate production-critical GitOps deployment platforms, shared tooling, APIs, and durable infrastructure workflows. The role requires 8+ years of software engineering experience, strong Go or TypeScript skills, and expertise with cloud, Kubernetes, and delivery systems.
The role leads the design, automation, security, and reliability of multi-region cloud infrastructure and Kubernetes platforms. It requires extensive software and infrastructure engineering experience, strong AWS and infrastructure-as-code expertise, and proficiency in Python, Go, or Bash.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Backend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.
Build and lead the evolution of Twilio’s large-scale observability platform, including telemetry pipelines, query systems, developer tooling, and standards. The role requires expertise in observability systems, distributed systems, cloud infrastructure, and modern programming languages.
Builds and operates Twilio’s global corporate network, VPN, zero-trust access, and cloud connectivity while monitoring performance and resolving incidents. The role requires substantial experience with Cisco, Palo Alto Networks, AWS networking, security protocols, and enterprise troubleshooting.
Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Build scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build internal developer-platform tooling, automation, observability, and AI-based workflows that improve engineering productivity across AWS, Azure, and Google Cloud. The role requires professional Python and/or Go experience, cloud experience, and strong production-service fundamentals.
Build and operate reliable, observable multi-cloud infrastructure for the dbt platform across AWS, Azure, and Google Cloud. The role emphasizes automation, Kubernetes administration, infrastructure as code, cloud cost optimization, developer experience, and operational reliability.
Leads cross-team initiatives that improve developer productivity through internal platforms, self-service tooling, cloud automation, observability, and AI-assisted workflows. The role requires strong Python or Go experience, cloud-platform expertise, and a record of delivering ambiguous technical projects.
Owns the reliability, scalability, deployment automation, and incident response of production infrastructure for a large-scale data platform. The role requires expertise in managed Kubernetes, multi-cloud platforms, infrastructure as code, cloud networking, scripting, and Linux administration.
Staff Site Reliability Engineer responsible for the performance, scalability, deployment robustness, vulnerability management, and incident response of a large-scale cloud data platform. The role requires deep Kubernetes and multi-cloud expertise, infrastructure automation, scripting, Linux administration, and cloud networking experience.
Owns the reliability, scalability, deployment automation, and incident readiness of production infrastructure for a large-scale SaaS platform. Requires 9+ years of experience plus expertise in Kubernetes, cloud platforms, infrastructure tooling, scripting, Linux, networking, and databases.
Own the reliability, scalability, deployment automation, and incident response of Fivetran’s production infrastructure. The role requires 7+ years of SaaS-scale experience plus deep expertise in Kubernetes, cloud platforms, infrastructure as code, Linux, networking, and programming.
Owns the reliability, scalability, deployment automation, and incident response of production infrastructure for a SaaS data platform. The role requires 7+ years of experience, strong managed Kubernetes and cloud-platform expertise, and proficiency with infrastructure-as-code and scripting.
Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
Designs and operates highly available GCP infrastructure and developer platforms for trading-critical systems. The role requires 5+ years of DevOps, platform, infrastructure, or SRE experience, with strong Terraform, Kubernetes, networking, CI/CD, observability, and incident-management skills.
Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Build software, tooling, and operational frameworks that improve incident response, on-call practices, post-mortem learning, and reliability across Datadog. The role requires at least five years of software development experience, distributed-systems expertise, and strong cross-functional technical leadership.