Skip to content
RoboflowRoboflow

Infrastructure Engineer

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

About the job

Responsibilities

  • Secure, scale, and maintain cloud architecture, databases, file storage, search clusters, microservices, and machine-learning pipelines.
  • Operate and optimize high-availability machine-learning inference services.
  • Build and manage containerized applications at scale.
  • Automate infrastructure with infrastructure-as-code and develop cost-effective scaling solutions.
  • Monitor and scale large applications, define SLOs/SLAs, participate in incident response, and join an on-call rotation.
  • Improve observability, alerting, and reliability processes.
  • Identify and implement infrastructure cost optimizations.
  • Contribute Python and JavaScript code to product features.
  • Collaborate with customer security teams on secure integrations and onboarding.
  • Fix vulnerabilities and bugs and harden systems for SOC 2, HIPAA, and GDPR readiness.

Requirements

  • Production experience with Kubernetes.
  • Experience with infrastructure-as-code, including Terraform, Helm, Bash, and Python.
  • Experience operating, monitoring, and scaling applications in AWS and/or Google Cloud, particularly ML/AI workloads.
  • Proficiency with Node.js and Python.
  • Experience with ML and big-data infrastructure, including GPUs, Docker, and Kubernetes.
  • Familiarity with PyTorch or TensorFlow.
  • Experience with CI/CD tools such as GitHub Actions or Spacelift.
  • Understanding of cloud-operations security best practices.
  • Ability to use AI and LLM tools throughout the development lifecycle.
  • Willingness to work collaboratively across product, operations, engineering, and customer-facing projects.

Compensation and Benefits

  • Target base compensation: USD $165,000–$200,000 annually.
  • $4,000 annual travel stipend.
  • $350 monthly productivity stipend.
  • Relocation bonus and access to company hubs and coworking resources.

Skills

Kubernetes, Terraform, Helm, Bash, Python, AWS, GCP, Node.js, Docker, PyTorch, TensorFlow, GitHub Actions, Spacelift, LLMs

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Hebbia

Hebbia

New York, NY
Software Engineer, Infrastructure
$160k+/yrOn-site5+ YOEDevOps / SRE

Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.