Skip to content
VercelVercel

Software Engineer, Compute

Build and operate Vercel’s low-level compute infrastructure, including storage, state, clusters, and distributed workloads. The role requires 5+ years of software engineering experience, strong Go skills, and deep expertise in Linux, virtualization, schedulers, and reliable distributed systems.

About the job

Responsibilities

  • Manage and improve a fleet of clusters running hundreds of instances across customer deployment regions.
  • Write Go daily and use Terraform to provision infrastructure; work with Nomad as the workload scheduler.
  • Rethink infrastructure primitives involving virtual filesystems, Linux primitives, and low-level virtualization.
  • Own the reliability and performance of the compute platform, including on-call coverage.
  • Collaborate across teams to drive convergence of compute infrastructure.

Requirements

  • 5+ years of software engineering experience; Go experience is strongly preferred.
  • Deep experience with virtual machines, file systems, and Linux; familiarity with tcpdump, strace, and iptables.
  • Experience building and operating distributed systems at scale, with a focus on performance and reliability.
  • Experience with schedulers and orchestrators for containerized and non-containerized workloads, such as Nomad or Kubernetes.
  • Excellent problem-solving and communication skills, with enthusiasm for solving complex infrastructure problems.

Nice-to-haves

  • Experience with low-level virtualization or sandbox execution environments.
  • Product engineering experience and interest in developer-facing infrastructure impact.
  • Experience with on-call operations for large-scale distributed systems.

Compensation and Benefits

  • Competitive compensation package, including equity.
  • Inclusive healthcare package.
  • Mentorship and support for attending professional events.
  • Flexible time off.
  • Company-provided equipment and a work-from-home budget.

Skills

Go, Terraform, Nomad, Kubernetes, Linux, Virtual Machines, File Systems, Distributed Systems, Virtualization, Tcpdump, Strace, Iptables

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

GitLab

GitLab

United Kingdom

Site Reliability Engineer, Infrastructure Platforms
No salary listedRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Writer

Writer

London, United Kingdom

Infrastructure Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.