Engineering Manager, Cloud Infrastructure
Leads a senior team managing Brex's AWS cloud infrastructure, driving Business Continuity and Disaster Recovery initiatives with multi-region failover. Requires 3+ years managing engineering teams and deep hands-on AWS, Kubernetes, Postgres experience.
About the job
Responsibilities
- Lead and Grow: Manage and support a senior team of 6 ICs, fostering a culture of high ownership, autonomy, and technical excellence.
- Drive BCDR Strategy: Oversee the execution of the multi-year BCDR initiative, ensuring multi-region infrastructure readiness and automated failover capabilities.
- Technical Stewardship: Maintain and scale Brex’s core cloud infrastructure, including Kubernetes (EKS), RDS/Aurora/Postgres, networking (VPC), and AWS service integrations.
- Cross-Functional Collaboration: Partner with product and engineering teams across Brex to drive adoption of new infrastructure capabilities and ensure successful migrations.
- Operational Excellence: Improve team hygiene by establishing and improving best practices for work tracking (Linear), documentation, and incident response.
- Resource Management: Effectively allocate resources to balance critical project delivery with long-term infrastructure stability.
Requirements
- Leadership Experience: 3+ years of experience directly managing software or infrastructure engineering teams.
- Technical Depth: Deep hands-on experience with cloud infrastructure foundations, specifically AWS, Kubernetes, and Postgres.
- Architectural Insight: Proven ability to design and implement complex infrastructure projects such as multi-region deployments or disaster recovery frameworks.
- Collaborative Mindset: Strong communication skills with the ability to influence technical decisions across multiple autonomous teams.
- Analytical Problem Solving: Ability to dive deep into technical issues while maintaining a high-level strategic view of organizational goals.
- Engineering Background: 5+ years of experience in software or systems engineering.
Nice-to-Haves
- Experience with Infrastructure as Code tools, specifically Terraform.
- Background in managing cloud infrastructure cost optimization programs.
- Experience with our tech stack components: Go, Kotlin, Python, or Elixir.
- Passion for leveraging AI and LLM-assisted workflows to increase engineering velocity.
Compensation
Expected salary range: $240,000 - $300,000. Starting base pay depends on location, skills, experience, market demands, and internal pay parity. Equity and other compensation may be provided.
Skills
AWS, Kubernetes, Postgres, Terraform, EKS, Rds, Aurora, Vpc, Go, Kotlin
Similar jobs
Engineering Management jobsLeads architecture and development of machine-learning and generative-AI systems for network analysis, including agentic workflows and conversational interfaces. Manages and mentors ML engineers while coordinating cross-functional delivery, production support, and roadmap execution.
Leads multiple engineering teams, combining people management with hands-on technical leadership, architecture, and coding. The role requires at least four years of software engineering experience and a demonstrated ability to build high-performing teams in complex startup environments.
Leads technical strategy and a multidisciplinary engineering team responsible for billing, accounting, eligibility, and enrollment systems. The role requires 5+ years of engineering management experience, strong distributed-systems expertise with Python or Go, and experience developing engineering leaders.
Leads and scales a customer-facing Applied AI Engineering team serving high-growth startups, guiding customers from experimentation to production while building repeatable deployment mechanisms. Requires technical depth in AI/ML platforms and experience leading teams in startup-focused, ambiguous environments.
Leads and builds a platform team responsible for infrastructure, developer productivity, and SRE/production support. The role combines hands-on technical work with team building and requires cloud, infrastructure troubleshooting, large-scale operations, and SRE experience.