Skip to content
Flux AutoFlux Auto

DevOps Engineer

DevOps Engineer builds and maintains full-stack systems for billing, authentication, onboarding, and integrations on Flux's AI hardware platform. Requires 5+ years SRE/DevOps experience, TypeScript, observability tools like Datadog/Sentry, and IaC with Pulumi across GCP/AWS/Firebase.

About the job

Responsibilities

  • Improve the reliability, availability, and operational health of production systems.
  • Set observability standards across services (metrics, logs, errors).
  • Set SLOs/SLIs, alerting, and on-call readiness with a focus on signal quality.
  • Partner with engineers to design resilient systems and reduce operational risk early.
  • Build internal tooling that improves system safety, debugging, and developer velocity.
  • Manage infrastructure via Pulumi across GCP, AWS, and Firebase.

Required Qualifications

  • 5+ years SRE, DevOps, or production operations experience.
  • At least 2 years of TypeScript web app development experience.
  • Proven experience operating and scaling production systems with uptime and latency goals.
  • Strong hands-on experience with observability stacks (Datadog, Sentry, or similar).
  • Experience defining SLOs/SLIs and building effective alerting strategies.
  • Proficiency with CI/CD systems and infrastructure-as-code.
  • Experience with cloud-native and serverless platforms (GCP, AWS).
  • Strong cross-system debugging and incident response skills.

Preferred Qualifications

  • Experience with multi-product distributed cloud tracing (stitching together Datadog, Sentry, GCP tracing).
  • Familiarity with containerized and serverless workloads (Cloud Run, Firebase, Lambda).
  • Experience supporting large TypeScript monorepos and build tooling (PNPM).
  • Background supporting AI-powered systems or high-variance workloads.
  • Startup experience or ownership of systems through rapid growth.

Skills

TypeScript, Pulumi, GCP, AWS, Firebase, Datadog, Sentry, CI/CD, Kubernetes, Serverless, Cloud Run, AWS Lambda, Pnpm

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.