Skip to content
IroncladIronclad

Senior Staff Site Reliability Engineer

Ironclad is seeking a Senior Staff Site Reliability Engineer to provide technical leadership and strategic direction for the SRE team, champion engineering excellence, and drive architectural resilience for their cloud platform.

About the job

Roles & Responsibilities:

  • Provide technical leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud Platform
  • Define and champion SRE best practices, setting the standard for engineering excellence across the entire organization
  • Solve the whole problem. Architecture for resiliency, identify risks, and make it happen.
  • A proven track record of designing and driving an 'automate-everything' culture (build, test, deploy, monitor)
  • Preference for collaboration, open communication and reaching across functional borders
  • Thorough understanding of backup/recovery systems, cloud storage architecture, and distributed systems
  • Be on an on-call rotation to respond to incidents that impact Ironclad’s availability, and provide support with internal or customer-facing incidents
  • Translate the near, mid, and long-term strategic needs of the business into a scalable, resilient platform roadmap
  • Drive critical architectural decisions with a relentless focus on security, scalability, and high performance
  • Be a mentor, multiply our team’s output with leadership and guidance

Key Skills:

  • 8+ years of DevOps / SRE experience
  • 5+ years of coding experience
  • Expert knowledge of Kubernetes and Google Cloud Platform (or similar provider)
  • Ability to build resilient infrastructure
  • Modern GitOps - Experience with tools like Terraform/Pulumi, CircleCI, ArgoCD
  • Experience with modern AI enabled tools such as Claude Code, Cursor, Zed
  • Troubleshooting and analytical skills, can PR review human and AI generated code
  • Strong technical aptitude and exceptional communication skills (written and verbal)
  • Desire for helping customers, and the ability to dive deep and learn a new product.
  • Experience and desire to work cross-functionally
  • Team and goal-oriented.
  • High output; low ego

Bonus Points if you have:

  • Experience with multi-region support
  • Expert Database Management Experience
  • Experience managing AI Infrastructure
  • Typescript Experience

Base Salary Range:

  • Staff Site Reliability Engineer: $220,000 - $235,000
  • Senior Staff Site Reliability Engineer: $245,000 - $270,000

US Full-Time Employee Benefits at Ironclad:

  • 100% health coverage for employees (medical, dental, and vision), and 75% coverage for dependents with buy-up plan options available
  • Market-leading leave policies, including gender-neutral parental leave and compassionate leave
  • Family forming support through Maven for you and your partner
  • Paid time off - take the time you need, when you need it
  • Monthly stipends for wellbeing, hybrid work, and (if applicable) cell phone use
  • Mental health support through Modern Health, including therapy, coaching, and digital tools
  • Pre-tax commuter benefits (US Employees)
  • 401(k) plan with Fidelity with employer match (US Employees)
  • Regular team events to connect, recharge, and have fun
  • And most importantly: the opportunity to help build the company you want to work at

Skills

Kubernetes, GCP, Terraform, Pulumi, CircleCI, Argo CD, Claude Code, TypeScript, Database Management

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.