Software Engineer, Site Reliability
Seasoned SRE owning reliability of Kubernetes-based production infrastructure at scale for a generative AI platform. Responsibilities include operating clusters, CI/CD, SLOs, monitoring, automation with AI, and driving improvements via chaos engineering. Requires 5+ years production experience with deep Kubernetes and observability expertise.
About the job
Key Responsibilities
- Own and operate Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads.
- Build and maintain CI/CD pipelines and deployment infrastructure.
- Leverage AI to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability.
- Build dashboards, alerting, and anomaly detection across systems.
- Define and enforce SLOs and build out incident response processes.
- Manage and improve networking, load balancing, and service mesh configurations.
- Drive reliability improvements across the stack through automation, runbooks, and chaos engineering.
Requirements
- 5+ years experience in managing critical production systems and software development workflows.
- Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible).
- Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS.
- Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD).
- Proficiency in Python and either Go or Bash for tooling and automation.
- Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog).
- Excellent communication and ability to drive technical decisions across teams.
- Self-starter who executes quickly, takes ownership, and constantly seeks improvement.
Nice-to-Haves
- Experience with managing GPU and AI/ML workloads.
- Experience with kernel-based monitoring and routing (eBPF, XDP).
- Experience with security tooling (Falco, Coroot, SIEM).
- Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB).
- Experience with distributed storage systems (Ceph, Longhorn, etc.).
Compensation
$180,000-250,000 plus equity + benefits (Range is based across 3 levels Mid, Senior and Staff).
Skills
Kubernetes, Terraform, Ansible, Python, Go, Bash, Prometheus, Grafana, Loki, Thanos, Victoriametrics, Datadog, Fluxcd, Argo CD, Linux Networking
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.