Skip to content
IdmeIdmeMountain View, CA

Software Engineer III - Data Platform

Builds reliability tooling, observability systems, infrastructure automation, and safe deployment processes for scalable cloud services. Requires a bachelor’s degree and 3–5 years of SRE, DevOps, or infrastructure experience, with cloud and programming expertise.

173k – 201k/yr
On-site3+ YOEDevOps / SRE

About the role

Responsibilities

  • Build and maintain automated reliability tooling, infrastructure as code, and observability systems that enhance uptime and service performance.
  • Develop monitoring, logging, and alerting frameworks such as Prometheus, Grafana, and OpenTelemetry.
  • Implement automated architectural reviews and reliability guardrails for agent-developed applications.
  • Partner with engineering teams to design and implement scalable, fault-tolerant systems that meet defined SLIs and SLOs.
  • Automate repetitive operational tasks and develop self-healing and auto-remediation mechanisms.
  • Participate in on-call rotations and lead incident response efforts, including post-incident reviews and systemic improvements.
  • Improve deployment and release processes using CI/CD pipelines and progressive delivery techniques.
  • Champion observability, reliability, and operational readiness reviews.
  • Collaborate with Security and Compliance teams to meet FedRAMP, NIST, and internal policy requirements.
  • Contribute to documentation, runbooks, and internal tooling.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related technical field.
  • 3–5 years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • 2+ years of hands-on experience managing and scaling services in AWS, Google Cloud, or Azure.
  • 1+ year of proficiency in at least one modern programming language such as Java, Go, Python, Ruby, or JavaScript.

Nice-to-haves

  • Strong understanding of Docker and Kubernetes.
  • Experience implementing and maintaining CI/CD pipelines and automation frameworks.
  • Working knowledge of observability systems, including metrics, tracing, logging, and alerting.
  • Experience building automated recovery, failover, or chaos-engineering systems.
  • Familiarity with event-driven architecture and asynchronous processing systems.
  • Knowledge of distributed systems design, load balancing, and performance optimization.
  • Exposure to Terraform, Pulumi, Ansible, and GitOps practices.
  • Understanding of FedRAMP, SOC 2, or NIST 800-53.
  • Strong analytical and troubleshooting skills across network and application layers.
  • Excellent communication and documentation skills.
  • Experience using AI agentic coding assistants and deploying custom AI agents or automated workflows into production.

Compensation and Benefits

  • Annual base salary: $172,528–$201,375 USD.
  • Compensation may vary based on qualifications, experience, skills, education, training, geographic location, and other job-related factors.
  • Benefits include medical, dental, vision, HSA, FSA, commuter benefits, life and disability insurance, 401(k) with company match, parental leave, paid time off, company holidays, employee assistance, pet insurance, wellbeing and childcare discounts, learning and development benefits, and other programs.

Skills

AWSGCPAzurePythonGoJavaScriptDockerKubernetesPrometheusGrafanaOpenTelemetryTerraformPulumiAnsibleCI/CD

Similar roles

DevOps / SRE jobs
Watershed

Software engineer, cloud infrastructure

WatershedNew York, NY +1

Builds and maintains cloud infrastructure systems on Google Cloud to support engineering teams in deploying, scaling, and observing production workloads. Requires 3+ years experience in infrastructure engineering.

174k – 242k/yrHybrid3+ YOEDevOps / SRE
Fluidstack

Software Engineer, GPU Infrastructure

FluidstackSan Francsisco, CA +3

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.

175k – 300k/yrOn-site5+ YOEDevOps / SRE
Fluidstack

Software Engineer, Cloud Infrastructure

FluidstackSan Francsisco, CA +3

Build and own the observability platform, control plane APIs, and fleet state management for a hyperscale GPU infrastructure powering AI compute at 10-100s of GW scale. Requires production service ownership at scale, comfort with AI coding tools, and on-call incident response.

175k – 300k/yrOn-site5+ YOEDevOps / SRE
Fluidstack

Production Engineer, Network

FluidstackAustin, TX

Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure at Fluidstack. Requires systems thinking, automation-first mindset, on-call ownership, and daily use of AI coding tools like Claude/Cursor alongside Go/Python and network protocols.

175k – 300k/yrOn-site5+ YOEDevOps / SRE
Fluidstack

Production Engineer, Compute

FluidstackSan Francisco, CA +3

Own end-to-end health, repair automation, and qualification of a hyperscale GPU/TPU compute fleet. Build metrics pipelines, firmware tooling, and self-healing repair workflows across Kubernetes and bare metal.

175k – 300k/yrHybrid5+ YOEDevOps / SRE