Skip to content
HebbiaHebbia

Software Engineer, Infrastructure

Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.

About the job

Responsibilities

  • Own Hebbia’s AWS footprint, including multi-account architecture, networking, IAM, and container orchestration.
  • Define and manage infrastructure entirely as code using application-code review, testing, and versioning standards.
  • Model cloud growth, right-size infrastructure, manage capacity, and automate cost controls.
  • Build and own the developer platform, including build pipelines, ephemeral and preview environments, and development tooling.
  • Improve CI/CD speed, safety, build times, and pipeline throughput.
  • Harden infrastructure for enterprise and financial-services security and compliance requirements.
  • Architect for scale readiness, including multi-region capability, network and account isolation, and enterprise deployment headroom.
  • Partner with product engineering to co-design infrastructure and application systems.
  • Evaluate and integrate infrastructure technologies that improve scalability, security, or engineering velocity.

Requirements

  • 5+ years building and operating cloud infrastructure at a venture-backed startup or top technology firm.
  • Deep expertise with infrastructure-as-code tooling and production estates managed entirely in code.
  • Production-depth AWS experience, including networking, IAM, and account architecture.
  • Strong command of containerization and orchestration technologies.
  • Production proficiency in Python, Go, or a similar scripting or backend language.
  • Strong CI/CD pipeline expertise and a record of improving developer velocity safely.
  • Working knowledge of operating-system and networking concepts, with the ability to debug infrastructure/application boundary issues.
  • Experience instrumenting infrastructure with monitoring and observability tooling.
  • Experience securing cloud infrastructure through least-privilege IAM, secrets management, hardening, and audit readiness.

Nice-to-haves

  • Experience managing cloud costs at scale or building tooling and guardrails that keep spending predictable.

Compensation and Benefits

  • Salary range: $160,000–$300,000 USD.
  • Unlimited PTO.
  • Medical, dental, and vision insurance.
  • 401(k).
  • Catered lunch daily and DoorDash dinner credit when staying late.
  • Parental leave: 3 months for non-birthing parents and 4 months for birthing parents.
  • $15,000 lifetime fertility benefit.
  • Competitive new-hire equity grant.

Skills

AWS, Infrastructure As Code, Networking, IAM, Kubernetes, Containers, Python, Go, CI/CD, Monitoring, Observability, Secrets Management, Cloud Security, Multi-Region Architecture, Cloud Cost Management

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.