Skip to content
FluidstackFluidstack

Production Engineer, Network

Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure at Fluidstack. Requires systems thinking, automation-first mindset, on-call ownership, and daily use of AI coding tools like Claude/Cursor alongside Go/Python and network protocols.

About the job

Responsibilities

  • Own network fleet health end to end: define realtime monitoring requirements, build alerting lifecycle, and ship dashboards for network state across all sites.
  • Build active debugging tooling including link diagnostics, remote command execution across the fleet, and repair visualization.
  • Turn repair into a pipeline: build automation from detection through parts management and return to service, including ticket integration, repair lifecycle pipelines, transceiver and optics tracking.
  • Own network qualification and validation: build frameworks that gate new sites and hardware into production.
  • Own end-to-end reliability, scalability, and operation of the network at-scale with aggressive automation, tooling, and incident discipline.

Requirements

  • Treat toil as a bug and build tools to eliminate manual work (e.g., replace manual SSH diagnostics with automation).
  • Think in systems: understand how faults like transceiver issues, misconfigured routes, and power events propagate and build tooling to distinguish them.
  • Move toward ambiguity, build maps, and communicate them clearly.
  • Learn at a steep slope and reach competence in unfamiliar domains quickly.
  • Comfortable carrying a pager, running incidents, writing postmortems, and fixing systemic causes.
  • Fluent with AI tooling: LLM APIs, MCP servers, agentic frameworks; drive Claude Code, Cursor, or similar daily.
  • Shipped production network tooling or automation that other teams depend on; comfortable in any language using AI coding tools.
  • Developed automation tools in Go and Python.
  • Experienced in link diagnostics, optical networks, and network monitoring (gNMI, gRPC, NETCONF, SONiC).

Nice-to-Haves

  • RMA and repair lifecycle automation.
  • Large-scale datacenter fabric (BGP, ECMP, spine-leaf).
  • Out-of-band network management.

Skills

Go, Python, Gnmi, gRPC, Netconf, Sonic, LLM APIs, Ai Coding Tools, Network Monitoring, Link Diagnostics, Optical Networks, BGP, Ecmp, Spine-Leaf, Rma Automation

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.