Skip to content
RenderRender

Software Engineer, Infrastructure

Builds and scales cloud infrastructure for Render's developer platform, focusing on container orchestration, networking, storage, and AI workloads. Requires 5+ years experience with Kubernetes, IaC tools like Terraform/Pulumi/Ansible, and production systems at scale.

About the job

What You'll Do

  • Own Render's core infrastructure across multiple data centers and regions.
  • Help offer unique capabilities to Render customers through infrastructure innovation.
  • Plan and architect for rapidly increasing scale.
  • Debug issues at all levels in our infrastructure stack.
  • Improve the performance and reliability of our infrastructure through increased observability, load testing, and chaos engineering.
  • Collaborate with other engineers to help keep our platform stable, predictable, and secure.
  • Participate in our on-call rotation, with the rest of the engineering team.

What We're Looking For

  • At least 5 years of experience building and scaling cloud infrastructure.
  • Experience developing, maintaining, and debugging production systems at scale.
  • Experience building, operating and scaling Kubernetes clusters or similar resource/container orchestration.
  • Experience with infrastructure-as-code tools like Terraform, Pulumi, and Ansible.

Nice-to-haves

  • Experience with Linux kernel and/or container optimization
  • Familiarity with observability tools like Datadog, Grafana, and OpenTelemetry.
  • Experience hosting PostgreSQL (or similar data stores) at scale.
  • Security hardening skills, especially in the context of untrusted workloads.

Skills

Kubernetes, Terraform, Pulumi, Ansible, Linux Kernel, Datadog, Grafana, OpenTelemetry, Postgres, Chaos Engineering

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.