# Engineering Manager, GPU Infrastructure

**Company:** [Cohere](https://hotfix.jobs/companies/cohere)
**Location:** San Francisco, CA
**Role:** Engineering Management
**Experience:** 5+ years
**Skills:** Kubernetes, PyTorch, TensorFlow, JAX, Terraform, Argo CD, Prometheus, Grafana, gpu infrastructure, hpc
**Posted:** 2026-07-21

> Lead and mentor a team building and optimizing GPU clusters and HPC infrastructure that powers Cohere's frontier AI models. Requires deep ML/HPC expertise, Kubernetes at scale, and experience managing distributed engineering teams.

## Job Description

## Responsibilities
- Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
- Manage performance, career development, and hiring for team members.
- Conduct regular 1:1s and team meetings to ensure alignment and address challenges.
- Provide technical guidance and support to team members on complex infrastructure problems.
- Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
- Oversee the implementation of topology-aware scheduling, hardware fault detection, and performance optimization systems.
- Collaborate with cloud providers to validate and deploy new GPU architectures.
- Ensure infrastructure reliability, scalability, and security across all GPU environments.
- Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
- Work with the Foundations team on training software stack adaptation for new GPU architectures.
- Coordinate with Capacity EPM on delivery timelines and resource planning.
- Interface with Legal and Security teams on compliance requirements.
- Collaborate with other infrastructure teams on shared goals and dependencies.
- Establish observability and monitoring frameworks for GPU utilization, performance, and reliability.
- Implement infrastructure-as-code practices and automation for cluster provisioning.
- Drive cost optimization initiatives while maintaining performance standards.
- Manage vendor relationships and contract negotiations for hardware and cloud services.
- Ensure documentation is comprehensive, up-to-date, and accessible to stakeholders.

## Requirements
- Experience managing engineering teams with a focus on technical mentorship and growth.
- Strong communication skills to translate complex technical concepts for diverse audiences.
- Ability to make data-informed decisions under pressure.
- Experience working in remote, distributed teams.
- Commitment to fostering an inclusive and collaborative team culture.
- Deep expertise in ML/HPC infrastructure: GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing environments.
- Proven experience with Kubernetes at scale: deployment, management, and troubleshooting cloud-native clusters for AI workloads in multi-cloud environments.
- Knowledge of infrastructure monitoring tools (Prometheus, Grafana).
- Familiarity with Terraform, ArgoCD, or other IaC tools.
- Experience with cost optimization and capacity planning for GPU infrastructure.
- Track record of collaborating with AI researchers or ML engineers to solve infrastructure challenges.
- Strong problem-solving abilities with a data-driven approach.
- Passion for enabling AI research through robust infrastructure.
- Collaborative mindset with a focus on cross-team success.
- Willingness to learn and adapt in a fast-paced, evolving environment.

## Nice-to-Haves
- None explicitly listed beyond the requirements.

## Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent.
- Full health and dental benefits, including a separate budget for mental health.
- RRSP matching, 401K, Pension Scheme.
- 100% Parental Leave top-up for up to 6 months, for either parent.
- Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
- Education & learning stipend for conferences, courses, and coaching.
- 6 weeks of paid vacation (30 working days!).
- Budget for traveling to other offices if you are remote, plus an annual company offsite.
- $500 home office stipend.
- Co-working benefit for those not near an office.

## Similar roles

- [Engineering Manager](https://hotfix.jobs/jobs/b0a49b7e-287e-4a91-9159-24a97f25cb3b) - WorkWhile - San Francisco, CA - $250k – $340k/yr
- [Engineering Manager, Foundation Model Inference](https://hotfix.jobs/jobs/6e6fe53a-525e-43b3-b91c-6f902a2541b0) - Databricks - Mountain View, CA - $190k – $261k/yr
- [Engineering Manager, AI Tooling](https://hotfix.jobs/jobs/0dc7d879-b2d2-41ab-bab0-c055f7cc34e3) - Airtable - San Francisco, CA - $281k – $366k/yr
- [Software Engineer - Notebooks](https://hotfix.jobs/jobs/fd20f44e-9761-4fb0-acdc-b3083ea83cbf) - Snowflake - Bellevue, WA - $236k – $339k/yr
- [Engineering Manager, Support & Stability](https://hotfix.jobs/jobs/b6a96911-a67b-4795-bea4-245240ce0d55) - Loop Returns - Remote - $150k – $190k/yr

**Apply:** https://hotfix.jobs/jobs/bcdd4677-ef1c-4b49-b17b-7c80bbe2e4cf
**Canonical:** https://hotfix.jobs/jobs/bcdd4677-ef1c-4b49-b17b-7c80bbe2e4cf