Skip to content
SesameSesame

SWE - Backend Infrastructure Engineer

Builds and scales core infrastructure including ML training/serving, Kubernetes clusters, and low-latency voice/audio pipelines. Requires 3+ years in infrastructure/ML systems, hands-on reliability engineering, and Kubernetes expertise.

About the job

Responsibilities

  • Design and build secure, maintainable, self-serve core infrastructure that engineering teams can rely on and operate independently
  • Architect and evolve a modern ML training infrastructure — scalable, reproducible, and built for rapid experimentation
  • Build and operate a modern model serving architecture with a focus on reliability, cost efficiency, and low latency
  • Own and scale the low-latency voice interface and audio processing pipeline — a technically demanding, performance-sensitive system at the core of Sesame's product
  • Build developer tooling, server infrastructure, and data infrastructure that is high leverage and low maintenance
  • Set technical direction within your domain, bring others along through clear communication and well-reasoned proposals, and raise the engineering bar across the team

Required Qualifications

  • A strong systems thinker who is equally comfortable setting direction and getting hands-on with implementation
  • Hands-on reliability engineering experience — you have well-formed convictions about observability, monitoring, deployment systems, and loosely coupled architectures, and you've put them into practice at scale
  • Proven track record of shipping services at scale, with all the operational complexity that comes with it
  • Kubernetes — significant production experience operating and scaling Kubernetes clusters
  • Experience designing and shipping flexible domain models and APIs — you think carefully about boundaries, contracts, and long-term maintainability
  • A default toward automation — you've consistently delivered efficiency gains through automation and have the track record to show it
  • Strong communication skills — you can set your own direction, write clearly about tradeoffs, and bring engineers and stakeholders along with you
  • 3+ years of software engineering experience, with significant time in infrastructure, platform, or ML systems roles

Preferred Qualifications

  • Infrastructure as Code at scale — significant IaC experience, preferably Terraform; CloudFormation, Pulumi, or Kubernetes-based approaches also welcome
  • ML infrastructure — PyTorch experience, especially model optimization for serving; ML training or serving experience; building ML serving and/or training infrastructure (TorchServe, Seldon, KServe, Ray Serve); large-scale distributed training and serving systems
  • Data engineering — pipeline design, dataset management, or data platform experience
  • Database design — complex schema design, query optimization, and hard data modeling decisions across relational and non-relational stores
  • Real-time communication systems — low-latency audio, video, or streaming infrastructure

Benefits

  • 401(k) max employer match: 3.5% of compensation
  • 100% employer-paid health, vision, and dental benefits for you and your dependents
  • Unlimited PTO and sick time
  • Flexible spending account with employer matching up to $1,650/year (medical FSA)
  • Guardian Employee Assistance Program (EAP)
  • Opportunity to share in the company's success with competitive stock options

Skills

Kubernetes, Terraform, PyTorch, ML Infrastructure, Infrastructure As Code, Observability, Monitoring, APIs, Automation, Torchserve, Seldon, Kserve, Ray Serve, Data Pipelines, Real-Time Systems

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.