# Infrastructure Engineer

**Company:** [Roboflow](https://hotfix.jobs/companies/roboflow)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $165k – $200k/yr
**Skills:** Kubernetes, Terraform, Helm, Bash, Python, AWS, GCP, Node.js, Docker, PyTorch, TensorFlow, GitHub Actions, Spacelift, LLMs
**Posted:** 2026-09-02

> Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

## Job Description

## Responsibilities
- Secure, scale, and maintain cloud architecture, databases, file storage, search clusters, microservices, and machine-learning pipelines.
- Operate and optimize high-availability machine-learning inference services.
- Build and manage containerized applications at scale.
- Automate infrastructure with infrastructure-as-code and develop cost-effective scaling solutions.
- Monitor and scale large applications, define SLOs/SLAs, participate in incident response, and join an on-call rotation.
- Improve observability, alerting, and reliability processes.
- Identify and implement infrastructure cost optimizations.
- Contribute Python and JavaScript code to product features.
- Collaborate with customer security teams on secure integrations and onboarding.
- Fix vulnerabilities and bugs and harden systems for SOC 2, HIPAA, and GDPR readiness.

## Requirements
- Production experience with Kubernetes.
- Experience with infrastructure-as-code, including Terraform, Helm, Bash, and Python.
- Experience operating, monitoring, and scaling applications in AWS and/or Google Cloud, particularly ML/AI workloads.
- Proficiency with Node.js and Python.
- Experience with ML and big-data infrastructure, including GPUs, Docker, and Kubernetes.
- Familiarity with PyTorch or TensorFlow.
- Experience with CI/CD tools such as GitHub Actions or Spacelift.
- Understanding of cloud-operations security best practices.
- Ability to use AI and LLM tools throughout the development lifecycle.
- Willingness to work collaboratively across product, operations, engineering, and customer-facing projects.

## Compensation and Benefits
- Target base compensation: **USD $165,000–$200,000 annually**.
- $4,000 annual travel stipend.
- $350 monthly productivity stipend.
- Relocation bonus and access to company hubs and coworking resources.

## Similar jobs

- [Software Engineer - Continuous Delivery](https://hotfix.jobs/jobs/0ceaa410-df08-45d6-80b3-68328c769297) - Baseten - San Francisco, CA - $165k – $330k/yr
- [TLM, Production Engineering](https://hotfix.jobs/jobs/0b915bd3-c030-4665-9806-d05334aa1351) - Ramp - New York, NY - $168k – $325k/yr
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/c5ea94a6-fb3e-4d29-9048-cd01579e7630) - Hebbia - New York, NY - $160k – $300k/yr
- [Software Engineer, Platform](https://hotfix.jobs/jobs/406361ec-fd77-45f6-8c6c-137d79ffa206) - Benchling - San Francisco, CA - $173k – $234k/yr

**Apply:** https://hotfix.jobs/jobs/a1531a31-7123-41d5-862f-d653b614ddd7
**Canonical:** https://hotfix.jobs/jobs/a1531a31-7123-41d5-862f-d653b614ddd7