Skip to content
OrbOrb

Software Engineer, Infrastructure - San Francisco HQ

Senior infrastructure engineer focused on building reliable, scalable systems for event ingestion, APIs, and billing. Requires 5+ years software engineering experience with 4+ years in infrastructure, strong debugging skills, and mentoring ability.

About the job

In this role you will:

  • Lead infrastructure resiliency efforts - recovery mechanisms, tenant isolation, load spike handling, etc
  • Improve observability and operability of our systems
  • Build performance-critical, user-facing infrastructure (eg. real-time event processing)
  • Plan scaling initiatives to handle large customer growth
  • Partner with other engineering teams to ensure we build reliable product features
  • Learn from a talented peer group, and share your expertise

About you:

  • You think deeply about edge cases, failure modes, bottlenecks, etc
  • You have a knack for investigating and debugging tricky errors and performance issues
  • You enjoy building scalable infrastructure for a high-growth product
  • You can effectively mentor fellow engineers on best practices (observability, rollout strategies, risk mitigation, etc)
  • You have 5+ years of experience in software engineering, and you’ve worked in in the infrastructure domain for 4+ years

Orb’s Tech Stack:

Frontend: Typescript + React + Tailwind CSS
Backend: Python
Datastores: PostgreSQL + Apache Druid + Clickhouse
Streaming Platforms: Kafka + Spark Streaming
Cloud Platform: AWS

Benefits:

  • Excellent medical, dental, and vision insurance
  • One Medical membership
  • Unlimited PTO plus an additional week off between Christmas and New Year’s
  • 401k plan
  • 16-week paid parental leave with equity vesting
  • Commuter stipend
  • Catered lunches in the office
  • Annual learning & development stipend
  • Meaningful equity in the form of stock options

Skills

Python, Postgres, Apache Druid, ClickHouse, Kafka, Spark Streaming, AWS, TypeScript, React, Kubernetes

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.