Manager - Site Reliability Engineering
Leads and develops an SRE team responsible for reliable, secure, large-scale cloud production systems. The role requires extensive engineering leadership experience, strong security expertise, incident response capability, and knowledge of AWS, Linux, networking, and CI/CD.
About the job
Responsibilities
- Mentor, manage, and lead a team of site reliability engineers with a broad range of expertise and experience.
- Advocate for security best practices and lead initiatives that strengthen the security posture of critical infrastructure.
- Respond to production incidents, drive rapid remediation, and identify preventative improvements.
- Triage and troubleshoot complex production issues to ensure reliability and performance.
- Collaborate with stakeholders across the organization to balance reliability, security, and delivery velocity.
- Partner with recruiting and people operations to hire and retain engineering talent.
- Monitor vulnerability scanning and security posture, cloud spend, RPO and RTO, and toil overhead; ensure projects improve these metrics.
- Support a 24x7 online environment as part of an on-call rotation.
Requirements
- 4+ years of experience managing SRE or software engineering teams, ideally in a cloud-native environment.
- 13+ years of experience, with strong leadership, communication, and project management skills.
- Experience managing teams operating large-scale production Java/Tomcat and containerized services in AWS or other cloud providers.
- Strong security background and knowledge.
- Deep knowledge of CI/CD principles, Linux fundamentals, OS hardening, networking concepts, and IP protocols.
- Bachelor's degree in computer science or equivalent experience.
Skills
Site Reliability Engineering, Java, Tomcat, AWS, Amazon Ec2, Amazon Ecs, Aws Kms, Amazon Kinesis, Amazon Rds, CI/CD, Linux, Os Hardening, Networking, Ip Protocols, Containerization
Similar jobs
Engineering Management jobsLeads architecture, hands-on development, and team execution for MongoDB’s modernization Validation and Eval product suite. Requires 8+ years of software development experience, JVM and Python expertise, database modernization knowledge, and at least 2 years of engineering leadership.
Leads and builds a new Bengaluru engineering team responsible for delivering secure, reliable identity-platform experiences. The role combines people management, hiring, roadmap collaboration, architectural leadership, and hands-on engineering, requiring 10+ years of experience and at least 3 years in engineering management.
Leads engineering delivery for secure workforce and customer authentication experiences, guiding technical strategy, architecture, implementation, and mentoring. Requires 10+ years of software development experience, including substantial enterprise Java experience, plus AWS, distributed systems, Agile, and security expertise.
Leads the SIW Platform engineering team building secure, scalable web authentication experiences. The role combines people management, technical roadmap ownership, cross-functional delivery, and expertise in modern web development, enterprise software, security, and identity protocols.
Manages and grows a team of frontend and backend engineers while contributing to technical design and implementation. The role requires 10+ years of software engineering experience, at least 2 years of engineering management, and strong Python or Java expertise.