Site Reliability Engineer (SRE)
Builds automation, observability, and tooling for Mithril's multi-cloud GPU orchestration platform, ensuring reliability, SLOs, and capacity management. Requires 3+ years SRE experience, Kubernetes proficiency, cloud expertise, and Python/Go coding skills.
About the job
Core Responsibilities
Reliability & SLOs
- Implement and own SLIs and SLOs across Mithril's API layer and internal orchestration services to ensure customer commitments are met.
- Partner with Product and Platform teams to ensure new features are designed for operability, reliability, and performance from the start.
- Participate in an on-call rotation; drive root cause analysis (RCA) for production incidents and implement durable fixes to prevent recurrence.
Observability & Monitoring
- Build and maintain dashboards, alerts, and distributed tracing within Mithril's monitoring stack to provide high-granularity visibility into our multi-cloud infrastructure.
- Own the tooling that surfaces supply-side signal — GPU availability, provider health, reservation fill rates — to both engineering and operations teams.
Infrastructure as Code & Automation
- Develop and maintain Terraform/Pulumi modules and Kubernetes configurations to manage Mithril's growing multi-cloud provider footprint.
- Write clean, maintainable Python (or Go) to automate repetitive operational tasks — from provider API reconciliation to automated health checks and capacity rebalancing.
Capacity Support
- Assist in managing GPU capacity across providers, ensuring the marketplace can dynamically respond to supply fluctuations and customer demand signals.
- Surface capacity constraints and failure modes early; contribute to the tooling that enables Mithril to make fast, data-driven supply decisions.
Requirements
- 3+ years of experience in SRE, Production Engineering, or Infrastructure roles at a high-growth technology company.
- Hands-on Kubernetes experience: comfortable managing clusters, deployments, and troubleshooting production incidents in a multi-tenant environment.
- Cloud proficiency in at least one major provider (AWS, GCP, or Azure), including practical understanding of cloud networking fundamentals (VPC, DNS, load balancing, security groups).
- Coding ability: proficiency in Python or equivalent (Go, Rust, etc.) — you build tools and services, not just scripts. Willing to pick up new languages as needed.
- Linux fundamentals: strong command of Linux systems, TCP/IP networking, and security best practices.
- Disciplined troubleshooter: calm under pressure during production incidents, with a rigorous approach to RCA and long-term remediation.
- Clear communicator: able to document processes and explain technical trade-offs to engineering and non-engineering teammates.
Nice to Have
- Experience with GPU/TPU-accelerated workloads or AI/ML infrastructure.
- Exposure to multi-cloud deployments or niche/specialized cloud providers (e.g., CoreWeave, Lambda Labs, Nebius).
- Familiarity with distributed systems concepts — service discovery, circuit breakers, consensus protocols.
- Production experience with Prometheus, Grafana, or OpenTelemetry.
- Prior experience in a high-growth startup environment where infrastructure scope expands faster than headcount.
Benefits
- Health, dental, and vision coverage for you and your dependents
- 401k Plan with 4% company match
- 21 days of PTO & 14 company holidays; including 2 floating holidays
Salary Range: $170,000 - $230,000
Skills
Kubernetes, Terraform, Pulumi, Python, Go, AWS, GCP, Azure, Prometheus, Grafana, OpenTelemetry, Linux
Similar jobs
DevOps / SRE jobsLeads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.