Software Engineer, Compute
Build and operate Vercel’s low-level compute infrastructure, including storage, state, clusters, and distributed workloads. The role requires 5+ years of software engineering experience, strong Go skills, and deep expertise in Linux, virtualization, schedulers, and reliable distributed systems.
About the job
Responsibilities
- Manage and improve a fleet of clusters running hundreds of instances across customer deployment regions.
- Write Go daily and use Terraform to provision infrastructure; work with Nomad as the workload scheduler.
- Rethink infrastructure primitives involving virtual filesystems, Linux primitives, and low-level virtualization.
- Own the reliability and performance of the compute platform, including on-call coverage.
- Collaborate across teams to drive convergence of compute infrastructure.
Requirements
- 5+ years of software engineering experience; Go experience is strongly preferred.
- Deep experience with virtual machines, file systems, and Linux; familiarity with tcpdump, strace, and iptables.
- Experience building and operating distributed systems at scale, with a focus on performance and reliability.
- Experience with schedulers and orchestrators for containerized and non-containerized workloads, such as Nomad or Kubernetes.
- Excellent problem-solving and communication skills, with enthusiasm for solving complex infrastructure problems.
Nice-to-haves
- Experience with low-level virtualization or sandbox execution environments.
- Product engineering experience and interest in developer-facing infrastructure impact.
- Experience with on-call operations for large-scale distributed systems.
Compensation and Benefits
- Competitive compensation package, including equity.
- Inclusive healthcare package.
- Mentorship and support for attending professional events.
- Flexible time off.
- Company-provided equipment and a work-from-home budget.
Skills
Go, Terraform, Nomad, Kubernetes, Linux, Virtual Machines, File Systems, Distributed Systems, Virtualization, Tcpdump, Strace, Iptables
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.