Leads reliability strategy and hands-on implementation for a large-scale, low-latency AI inference service. The role requires 7+ years in backend, infrastructure, or reliability engineering, strong backend programming skills, and deep expertise in distributed-system reliability.
Salary not listed
On-site7+ YOEDevOps / SRE
About the role
Responsibilities
Define and drive reliability strategy, including establishing SLOs and aligning engineering teams.
Design and implement reliability mechanisms for fault detection, graceful degradation, failover, throttling, and recovery across multiple regions and data centers.
Lead large-scale incident management, including postmortems, root-cause analysis, and prevention loops.
Architect for reliability and observability by influencing system design for redundancy, durability, and debuggability.
Develop internal tooling and frameworks for chaos testing, load simulation, and distributed fault injection.
Collaborate with software, infrastructure, and hardware teams to embed reliability across the inference service.
Build dashboards and alerts to monitor service health and communicate actionable reliability metrics.
Mentor engineers and establish best practices for designing, testing, and operating reliable large-scale systems.
Requirements
Bachelor's or master's degree in computer science or a related field.
7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
Strong programming skills in at least one backend programming language, such as Python, C++, Go, or Rust.
Deep experience with reliability principles, including SLO/SLI/SLA design, incident response, and postmortem culture.
Excellent communication and cross-functional leadership skills.
Nice-to-haves
Experience building large-scale AI infrastructure systems.
Benefits
Work on a breakthrough AI platform beyond the constraints of GPU-based systems.
Publish and open-source cutting-edge AI research.
Work on one of the world's fastest AI supercomputers.
Enjoy job stability with startup vitality.
Participate in a non-corporate work culture that respects individual beliefs.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Systems Engineer, DevTools
CloudflareAtlanta, GA +3
Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
200k – 281k/yrHybrid7+ YOEDevOps / SRE
Principal Operations Engineer, Controls
FluidstackUnited States
As Principal Operations Engineer, Controls, you will be the senior technical authority for operational building automation and control systems across hyperscale AI data centers. You will lead site assessments, drive technical readiness, review designs, and ensure the continuous improvement of control systems.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Operations Engineer, Electrical
FluidstackUnited States
Fluidstack is seeking a Principal Operations Engineer, Electrical to be the senior technical authority for electrical infrastructure across their hyperscale AI data center portfolio. This role involves leading site assessments, driving technical readiness, reviewing designs, and feeding operational learnings back into the design and manufacturing organization.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Classified Systems Architect, Okta Federal
OktaWashington, DC
Leads architecture of air-gapped classified (SIPR/JWICS) developer platforms for DoD compliance, designs IaC for disconnected ops, integrates hardened tools like Big Bang/Iron Bank, and ensures secure, scalable Kubernetes infrastructure. Requires 12+ years experience with 5+ in classified DoD environments.