Lead/Manager Site Reliability Engineering Team
Leads, coaches, and develops an SRE team while owning infrastructure reliability, scalability, monitoring, incident response, and operational processes. The role requires 7+ years of SRE or related experience, strong Ansible, Terraform, and Kubernetes expertise, and a bachelor's degree or equivalent experience.
About the job
Responsibilities
- Participate in an on-call PagerDuty rotation to respond to incidents impacting availability.
- Manage, develop, and coach the SRE team.
- Build and run infrastructure with Ansible, Terraform, and Kubernetes to enable scaling to a massive number of concurrent users.
- Build monitoring systems to ensure high-quality service for customers.
- Design and implement operational processes, including deployments and upgrades.
- Debug production issues across services and all levels of the stack.
- Identify improvements to product architecture from reliability, performance, and availability perspectives.
- Plan the growth of Together AI's infrastructure.
Requirements
- 7+ years of professional SRE or related experience.
- Ideally 2 years of experience as a Lead SRE.
- Bachelor's degree in Computer Science or a related field, or equivalent work experience.
- Expert knowledge of Ansible, including roles and playbooks, Terraform, and Kubernetes.
- Proficiency in programming and scripting languages.
- Direct experience with monitoring and observability practices.
- Advanced knowledge of cloud services.
- Ability to thrive in a collaborative environment involving different stakeholders and subject matter experts.
Skills
Ansible, Terraform, Kubernetes, Pagerduty, Monitoring, Observability, Cloud Services, Programming Languages, Scripting Languages, Distributed Systems
Similar jobs
Engineering Management jobsLeads the team responsible for building and operating a reliable managed PostgreSQL service, combining people leadership, technical direction, roadmap ownership, and operational accountability. Requires 8+ years of software engineering experience, including 3+ years managing infrastructure, database, or developer-platform engineers.
Leads engineering teams building scalable infrastructure and products for distributed data and machine-learning workloads. The role emphasizes hiring and developing engineers, establishing technical excellence, roadmap planning, and cross-team execution.
Leads and develops Dandy’s Sales Engineering team, partnering with Account Executives to support revenue growth through technical sales, product demonstrations, and customer advisory. Requires at least five years of dental-industry Sales Engineering experience and strong people leadership skills.
Leads technical strategy and a multidisciplinary engineering team responsible for billing, accounting, eligibility, and enrollment systems. The role requires 5+ years of engineering management experience, strong distributed-systems expertise with Python or Go, and experience developing engineering leaders.
Leads the engineering team and cross-organizational initiatives responsible for secure, consistent Unity Catalog runtime enforcement across Databricks compute engines and clouds. The role requires extensive distributed-systems experience, security expertise, operational maturity, and technical leadership.