DevOps Engineer embedded in platform teams to build and operate scalable AI infrastructure on GCP/AWS. Own CI/CD, Kubernetes, observability, reliability (SLIs/SLOs), automation, and performance at petabyte scale. Requires 4-5 years production DevOps/SRE experience.
150k – 170k/yr
On-site4+ YOEDevOps / SRE
About the role
Responsibilities
Own and continuously improve CI/CD pipelines and deployment processes; partner with developers to review infrastructure changes, streamline releases, and champion DevOps best practices.
Design, deploy, and maintain cloud infrastructure on GCP and AWS; manage Kubernetes clusters, networking, and storage at petabyte scale using infrastructure-as-code.
Build, guide, and review automation and internal tooling to drive developer productivity and eliminate manual toil.
Profile and optimize services for large-scale data pipelines; perform capacity planning for storage and compute-intensive workloads; establish performance benchmarks.
Define and own SLIs/SLOs/SLAs for critical services; build alerting, runbooks, and incident response processes; lead blameless postmortems.
Instrument services with distributed tracing, logging, and metrics (Prometheus, Grafana, OpenTelemetry, GCP Dashboards); define observability best practices and ensure services are observable before production.
Requirements
4–5 years of hands-on DevOps, platform engineering, or SRE experience in a production environment.
Strong experience building and maintaining CI/CD pipelines and deployment automation at scale.
Proven experience with infrastructure-as-code tools (e.g., Terraform, Pulumi) and configuration management.
Strong fundamentals in designing, building, and maintaining resilient distributed and/or high performance systems.
Hands-on experience with Kubernetes and containerised workloads in cloud environments (GCP and/or AWS).
Solid understanding of networking, operating systems, and database technologies.
Experience with observability fundamentals: metrics, logs, traces, and alerting.
Nice-to-Haves
Experience with Python, TypeScript, React, PyTorch, CUDA, or Ray.
Openness to learning new technologies (company is technology agnostic).
Compensation and Benefits
Competitive salary and equity in a hyper growth startup.
Flexible PTO, 18 paid vacation days + federal holidays.
Annual learning and development budget.
Health, dental, and vision insurance.
Opportunities for travel, bi-annual off-sites, and monthly socials.
Strong in-person culture (4-5 days/week in North Beach office).
Skills
KubernetesTerraformGCPAWSCI/CDPrometheusGrafanaOpenTelemetryPythonInfrastructure As Code
Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.
150k – 250k/yr
Hybrid5+ YOEDevOps / SRE
Network Engineer, BMS/EPMS Networks
FluidstackNew York, NY
Own and design OT facility networks for BMS, EPMS, and controls traffic in AI data centers. Requires experience building industrial/OT networks, deep knowledge of controls protocols like BACnet/IP, and implementing segmentation that works with operations teams.
150k – 203k/yr
On-site5+ YOEDevOps / SRE
Site Reliability Engineer
RunpodUnited States
Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.
150k – 200k/yr
Remote5+ YOEDevOps / SRE
DevSecOps Engineer
Turion SpaceIrvine, CA
DevSecOps Engineer designing, developing, and maintaining reliable C++/Python software systems with a focus on CI/CD automation, build systems, and developer tooling for aerospace and defense applications. Requires 3+ years experience, proficiency in modern C++, CMake, and CI/CD pipelines.
150k – 213k/yr
Hybrid3+ YOEDevOps / SRE
Software Engineer
Clear StreetNew York, NY
Platform Engineer building self-service internal developer platforms, reusable services, intelligent CI/CD, and automation to improve developer productivity and software delivery at a capital markets fintech. Requires 5+ years experience, strong software engineering skills in Go/Python/Java, deep cloud-native and Kubernetes expertise, and a product mindset for internal tools.