Skip to content

Software Engineer, Cloud Infrastructure

Designs, builds, and operates scalable multi-cloud infrastructure powering ML training, inference, and data curation pipelines. Collaborates with teams on AWS-focused systems using Kubernetes and IaC tools like Terraform.

About the job

What You'll Work On

  • Architect and maintain our multi-cloud infrastructure (primarily AWS, potentially Azure/GCP), with a focus on reliability, security, and scalability
  • Define and implement infrastructure-as-code best practices using Terraform, CloudFormation, Pulumi (and similar technologies)
  • Design and manage Kubernetes-based systems for model training, inference, and data processing workloads
  • Optimize our CI/CD pipelines and streamline deployment of services across environments
  • Build monitoring, alerting, and logging systems to ensure high system availability and observability
  • Collaborate with research and engineering teams to provide infrastructure support for training large-scale ML models
  • Ensure our infrastructure supports various deployment models (cloud, on-prem, hybrid) for enterprise use cases
  • Drive cost-efficiency strategies across compute and storage resources
  • Respond to and resolve infrastructure-related incidents with a sense of ownership and urgency

About You

  • You've led or helped build robust infrastructure systems at a startup or fast-moving engineering organization
  • Deep experience working with cloud providers (especially AWS), and ideally exposure to multi-cloud or hybrid-cloud setups
  • Strong with Kubernetes, Terraform, and containerized architectures
  • Confident with systems-level debugging—networking issues, memory leaks, resource bottlenecks, etc.
  • Comfortable writing clean, maintainable scripts in Bash, Python, or Go
  • You care deeply about building secure and scalable systems and take pride in reliable infrastructure
  • You're collaborative, humble, and ready to own high-impact projects end-to-end

Nice to Have

  • Experience supporting infrastructure for ML workloads (training pipelines, inference clusters, GPU orchestration)
  • Built or scaled infrastructure for teams working with large-scale datasets
  • Exposure to cost monitoring and optimization tools in cloud environments
  • Background supporting compliance and security in enterprise deployments

Compensation

Salary ranges from $180,000 to $300,000.

Comprehensive benefits: 100% covered health benefits, 401(k) with 4% match, unlimited PTO, wellness and learning stipends, daily lunches/snacks, relocation assistance.

Skills

AWS, Kubernetes, Terraform, CI/CD, Python, Go, Bash, CloudFormation, Pulumi, GPU

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.