Distributed Systems Engineer
Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.
About the job
Role Scope
Own the observability platform. Build and operate the data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible — from site down to device and link. No other team should need to scrape production directly to answer a question.
Define and build the API surface for infrastructure. Design the contracts between production infrastructure and every tool that touches it. All other teams at Fluidstack use your tooling to manage and operate our hyperscale fleet.
Build the production control plane. Unified machine management, actual state inspection, distributed command execution — and the Kubernetes-based infrastructure that underpins it all.
Own fleet state as source of truth. SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms. What the system says about itself should match reality, and you're accountable when it doesn't.
Land new hardware into the platform cleanly. ZTP, DHCP, DNS, artifacts — every new XPU generation and site integration goes through IaaS before production.
What We're Looking For
- You treat toil as a bug. If something requires a human to do it twice, you build the thing that makes it not require a human.
- You design APIs that age well. You've felt the pain of a leaky abstraction at scale and you don't repeat it.
- You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
- You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
- You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
- You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
- You've shipped production services that other teams depend on at scale, and you're comfortable in any language using AI coding tools.
Bonus:
- Distributed systems and data pipeline engineering.
- Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics).
- API design and versioning at scale.
- Workflow and orchestration engines (Temporal, Cadence).
- BMC/Redfish or hardware telemetry.
- Go, Python, and Postgres.
Skills
Kubernetes, Prometheus, Thanos, Victoriametrics, Go, Python, Postgres, Temporal, Cadence, Redfish, API Design, Distributed Systems, Data Pipelines
Similar jobs
DevOps / SRE jobsBuild and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.