Skip to content
IntusCareIntusCare

Director of SRE (FTE)

Leads the reliability, observability, incident management, QA, and release engineering functions for a cloud-native healthcare SaaS platform. Requires 12+ years of SRE, infrastructure, or platform engineering experience and significant engineering leadership experience.

About the job

Responsibilities

  • Own and execute the SRE strategy and multi-quarter roadmap across reliability, observability, incident management, QA maturity, and release engineering.
  • Define, measure, and improve SLAs, SLOs, error budgets, uptime, performance, and operational health metrics across products and services.
  • Lead production reliability, including monitoring, alerting, on-call operations, incident response, root cause analysis, and MTTR reduction.
  • Establish release-readiness standards, deployment safety controls, and quality gates.
  • Manage external SRE vendors and partners, including service delivery, SLA governance, escalations, performance reviews, and compliance expectations.
  • Lead QA engineering strategy focused on automation, regression prevention, test coverage, and reducing escaped defects.
  • Partner with Security and Engineering to ensure cloud infrastructure, CI/CD pipelines, and operational tooling meet HIPAA, SOC 2, and internal security standards.
  • Oversee Azure AKS environments, Kubernetes, GitOps workflows, CI/CD pipelines, GitHub Actions, secrets management, access controls, and audit readiness.
  • Drive observability maturity using Grafana, Prometheus, logging platforms, tracing tools, and automated alerting frameworks.
  • Collaborate across Product, Platform, and Engineering to embed reliability and quality practices throughout the software development lifecycle.
  • Build, mentor, and scale SRE and QA teams.
  • Drive AI-enabled automation and intelligent tooling to reduce toil and improve operational excellence.

Requirements

  • 12+ years of SRE, infrastructure, or platform engineering experience, including 5+ years in engineering leadership roles.
  • Proven experience owning site reliability for complex, multi-tenant SaaS platforms with demanding availability requirements.
  • Experience defining SLA and SLO frameworks, error budgets, and incident management processes at scale.
  • Experience managing managed-infrastructure or SRE service vendors, including SLA governance and performance management.
  • Experience leading QA or quality engineering functions, test automation maturity, and release gate ownership.
  • Strong communication and cross-functional influence skills.
  • Hands-on experience with Microsoft Azure, preferably including AKS, networking, storage, IAM, and security services.
  • Deep expertise in Kubernetes, containerized workloads, and production-scale distributed systems.
  • Experience with CI/CD tooling such as GitHub Actions, ArgoCD, and Terraform.
  • Strong background in monitoring, logging, tracing, and observability platforms.
  • Experience with scripting and automation using Python, Bash, PowerShell, or similar languages.
  • Understanding of release engineering, automated testing frameworks, QA tooling, and shift-left quality practices.
  • Experience supporting SaaS applications with uptime, scalability, and security requirements in regulated industries.
  • Knowledge of HIPAA, SOC 2, vulnerability management, access controls, and infrastructure security.
  • Familiarity with databases, APIs, networking, and troubleshooting modern web application stacks.

Nice-to-haves

  • Healthcare technology or HIPAA-compliant environment experience.
  • Familiarity with FHIR-native or EMR/EHR platform architectures.
  • Experience implementing AI-assisted SRE automation, including runbook generation, anomaly detection, or incident triage.
  • Experience with Playwright or equivalent test automation frameworks in a QA leadership capacity.
  • Experience building internal SRE capabilities alongside a managed services provider.

Compensation and Benefits

  • Base salary range: $175,000–$200,000.
  • Final compensation may include a variable component and stock options.
  • Fully remote role based in the United States.
  • This position is not eligible for sponsorship.

Skills

Microsoft Azure, Azure Aks, Kubernetes, GitHub Actions, Argo CD, Terraform, Grafana, Prometheus, Datadog, Splunk, Python, Bash, PowerShell, Playwright, CI/CD

Fluidstack

Fluidstack

Remote

Principal Operations Engineer, Mechanical
$150k+/yrRemote10+ YOEDevOps / SRE

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

Cloudflare

Cloudflare

Atlanta, GA
Principal Systems Engineer, DevTools
$200k+/yrHybrid7+ YOEDevOps / SRE

Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.

Fluidstack

Fluidstack

United States

Principal Operations Engineer, Network
$258k+/yrRemote7+ YOEDevOps / SRE

This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.

Snowflake

Snowflake

Menlo Park, CA

Principal Software Engineer - Performance Engineering
$264k+/yrOn-site12+ YOEDevOps / SRE

Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.

AllSpice

AllSpice

Boston, MA
Principal / Staff / Senior Infrastructure Engineer
No salary listedHybrid8+ YOEDevOps / SRE

Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.