Manager - Site Reliability Engineer
Leads a team responsible for internal developer platforms, SRE tooling, automation, and production reliability in cloud environments. Requires people-management experience, strong security and infrastructure expertise, and a bachelor's degree or equivalent experience.
About the job
Responsibilities
- Lead the architecture, design, and rollout of an internal developer platform spanning CI/CD, tooling, infrastructure-as-code integrations, and modernization efforts.
- Enhance fleet-management automation and build features for the infrastructure platform.
- Lead initiatives that improve engineer productivity by identifying and mitigating development-flow bottlenecks.
- Mentor, manage, and lead software engineers and site reliability engineers.
- Triage and troubleshoot complex production issues to maintain reliability and performance.
- Partner with stakeholders to balance reliability, security, and delivery velocity.
- Partner with recruiting and people operations to hire and retain engineering talent.
- Monitor vulnerability scanning, security posture, cloud spend, RPO, RTO, and operational toil, and ensure projects improve these metrics.
- Support a 24x7 online environment through an on-call rotation.
Requirements
- 3+ years of experience managing software engineering or SRE teams, ideally in a cloud-native environment.
- Proven experience building and scaling internal developer platforms using a platform-as-product mindset.
- Experience leading teams that build automated tooling for CI/CD orchestration, Kubernetes operators, or self-service infrastructure provisioning.
- Experience replacing manual, toil-heavy operations with automated, code-driven workflows in high-scale cloud production environments.
- Experience managing large-scale production Java/Tomcat and containerized services in AWS or another cloud provider.
- Deep knowledge of CI/CD principles, Linux fundamentals, OS hardening, networking concepts, and IP protocols.
- Strong leadership, communication, project management, and security skills.
- Bachelor's degree in computer science or equivalent experience.
Nice-to-haves
- Experience with AWS EC2, ECS, KMS, Kinesis, and RDS.
- Experience with infrastructure as code, cloud-native platforms, fleet management, and developer productivity engineering.
Compensation and Benefits
- Benefits include well-being support, social-impact opportunities, talent development, and community connection.
Skills
CI/CD, Infrastructure As Code, Kubernetes, AWS, Java, Tomcat, Linux, Os Hardening, Networking, Ip Protocols, Amazon Ec2, Amazon Ecs, Amazon Kms, Amazon Kinesis, Amazon Rds
Similar jobs
Engineering Management jobsLeads a technical engineering team responsible for large-scale distributed services and GitLab Dedicated infrastructure. The role requires strong operational expertise, experience managing high-performing remote teams, and the ability to guide architecture, incident response, automation, and AI-enabled productivity.
Leads and develops Dandy’s Sales Engineering team, partnering with Account Executives to support revenue growth through technical sales, product demonstrations, and customer advisory. Requires at least five years of dental-industry Sales Engineering experience and strong people leadership skills.
Engineering manager leading the Client Foundations team responsible for libraries, tooling, security, and infrastructure across device platforms. Requires significant software development experience, mobile or macOS expertise, and engineering people-management experience.
Leads a globally distributed team responsible for deployment tooling, upgrades, reliability, and operational simplicity across GitLab deployment models. Requires platform or SRE leadership experience, strong Kubernetes and Helm expertise, and the ability to guide distributed teams through complex operational challenges.
Leads a distributed backend engineering team responsible for PostgreSQL health and a major database migration-system replacement. The role requires engineering management, production PostgreSQL and capacity-planning experience, async-first written communication, hiring and team-building capability, and experience managing long-horizon infrastructure programs.