Skip to content

Software Engineer, Infrastructure

Builds scalable data and ML infrastructure supporting multi-cloud and deployment models. Partners with founders and engineers on core platforms, tooling, and research features for reliable production systems.

About the job

What You'll Work On

  • Design and build the development and production platforms that power our products, enabling reliability and security at scale
  • Architect, build, and deploy our core infrastructure while supporting multiple cloud providers and various deployment models
  • Accelerate company productivity by empowering your fellow engineers & teammates with excellent tooling and systems, providing a best-in-case experience
  • Partner with researchers and engineers to bring new features and research capabilities to our customers

About You

  • Have meaningful experience in spearheading and constructing large-scale infrastructure
  • Proficiency in bash, Kubernetes, Python, and/or Terraform or similar technologies
  • Have experience working with AWS, other cloud platforms such as Azure or GCP and/or on-prem environments
  • Have expertise in debugging problems across the stack, such as networking issues, performance problems, hardware issues or memory leaks
  • Take pride in building and operating scalable, reliable, secure systems
  • Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed
  • Own problems end-to-end and are willing to pick up whatever knowledge you're missing to get the job done

We would love it if you had

  • Built out data infrastructure from, or nearly from, scratch at a fast-growing startup
  • Experience building ML/DL infrastructure and/or data infrastructure that feeds into training large ML models

Compensation

  • Base salary: $180,000 to $300,000
  • Significant equity
  • 100% covered health benefits (medical, vision, and dental)
  • 401(k) with 4% company match
  • Unlimited PTO
  • Annual $2,000 wellness stipend
  • Annual $1,000 learning stipend
  • Daily lunches and snacks
  • Relocation assistance

Skills

Kubernetes, Terraform, Python, Bash, AWS, Azure, GCP, ML Infrastructure, Data Infrastructure, Debugging

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.