Engineering Manager, Site Reliability
Lead and mentor a distributed SRE team responsible for Radar's high-throughput production infrastructure, observability, multi-region availability, and incident management on AWS/EKS/Terraform. Requires hands-on production experience, Kubernetes, sharded MongoDB, and high-growth startup background.
About the job
What you’ll do
- Manage, grow, and mentor a globally distributed SRE team
- Be responsible for our observability, monitoring, and incident management practices
- Work on core Radar infrastructure using Terraform all deployed to AWS via EKS
- Work with our product, data, platform, and security engineers to ensure the development of highly-available, scalable, and reliable software
- Champion automation and self-service capabilities over manual processes, and have grounded perspectives on AI-assisted tooling
- Drive critical company-level initiatives, for example, 99.99+% availability, multi-region deployment, and cloud cost-saving activities
- Ensure compliance and security by default in our processes and infrastructure by partnering with our GRC and security teams
- Have your work be used by 100's of millions of devices
- Be part of the on-call rotation
- Talk to Radar customers and prospects, hear their feedback, incorporate it into your work and make them successful
You should
- Be able to get your hands dirty and dive deep into production issues
- Have experience managing a production AWS environment via Terraform
- Have experience with high availability multi-region production infrastructure running on Kubernetes
- Have experience managing large sharded Mongo clusters with heavy read-write workloads
- Have experience at a high growth startup
- Be interested in talking to customers or prospects and making them successful
Bonus points if you
- Are a former technical co-founder
- Have experience with high throughput data intensive applications
What we offer
- Competitive salary
- Meaningful stock options in a fast-growing company
- 401(k) plan with 4% match
- New HQ in Flatiron, NYC
- Top-notch equipment
- Catered lunches
- Unlimited PTO
- Health, dental, and vision insurance with 100% coverage for employees
- 12 weeks of paid parental leave
- Commuter and fitness benefits
Skills
Terraform, AWS, EKS, Kubernetes, MongoDB, Observability, Monitoring, Incident Management, CloudWatch, Grafana, Pagerduty, CircleCI, TypeScript, Rust, Airflow
Similar jobs
Engineering Management jobsLeads a platform engineering team responsible for shared systems, identity, permissions, session management, and messaging infrastructure. The manager owns delivery and architecture, develops engineers, embeds AI into development workflows, and partners closely with Product and Design.
Engineering manager leading a data platform team that allocates cloud costs to products and customers. The role requires experience with reproducible batch pipelines, versioned financial metrics, cross-functional delivery, and cloud cost, billing, revenue, or financial data systems.
Leads a hands-on team building and maintaining third-party security and enterprise integrations. The manager coaches engineers, guides technical decisions and delivery, and contributes to backend systems, APIs, and production troubleshooting.
Leads the engineering team building an AI-powered referral coordination product integrating document processing, voice automation, and EHR systems. The role combines people management, hands-on technical leadership, scalable architecture, customer collaboration, and healthcare interoperability expertise.
Leads and develops a 4–5 person engineering squad while continuing to write production code and owning delivery, architecture, and operational health. The role focuses on safely shipping LLM-powered financial experiences and requires people-management experience, AI-native development fluency, and hands-on software engineering skills.