Skip to content
CohereCohereSan Francisco, CA

Engineering Manager, GPU Infrastructure

Lead and mentor a team building and optimizing GPU clusters and HPC infrastructure that powers Cohere's frontier AI models. Requires deep ML/HPC expertise, Kubernetes at scale, and experience managing distributed engineering teams.

Salary not listed
On-site5+ YOEEngineering Management

About the role

Responsibilities

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Manage performance, career development, and hiring for team members.
  • Conduct regular 1:1s and team meetings to ensure alignment and address challenges.
  • Provide technical guidance and support to team members on complex infrastructure problems.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Oversee the implementation of topology-aware scheduling, hardware fault detection, and performance optimization systems.
  • Collaborate with cloud providers to validate and deploy new GPU architectures.
  • Ensure infrastructure reliability, scalability, and security across all GPU environments.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
  • Work with the Foundations team on training software stack adaptation for new GPU architectures.
  • Coordinate with Capacity EPM on delivery timelines and resource planning.
  • Interface with Legal and Security teams on compliance requirements.
  • Collaborate with other infrastructure teams on shared goals and dependencies.
  • Establish observability and monitoring frameworks for GPU utilization, performance, and reliability.
  • Implement infrastructure-as-code practices and automation for cluster provisioning.
  • Drive cost optimization initiatives while maintaining performance standards.
  • Manage vendor relationships and contract negotiations for hardware and cloud services.
  • Ensure documentation is comprehensive, up-to-date, and accessible to stakeholders.

Requirements

  • Experience managing engineering teams with a focus on technical mentorship and growth.
  • Strong communication skills to translate complex technical concepts for diverse audiences.
  • Ability to make data-informed decisions under pressure.
  • Experience working in remote, distributed teams.
  • Commitment to fostering an inclusive and collaborative team culture.
  • Deep expertise in ML/HPC infrastructure: GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing environments.
  • Proven experience with Kubernetes at scale: deployment, management, and troubleshooting cloud-native clusters for AI workloads in multi-cloud environments.
  • Knowledge of infrastructure monitoring tools (Prometheus, Grafana).
  • Familiarity with Terraform, ArgoCD, or other IaC tools.
  • Experience with cost optimization and capacity planning for GPU infrastructure.
  • Track record of collaborating with AI researchers or ML engineers to solve infrastructure challenges.
  • Strong problem-solving abilities with a data-driven approach.
  • Passion for enabling AI research through robust infrastructure.
  • Collaborative mindset with a focus on cross-team success.
  • Willingness to learn and adapt in a fast-paced, evolving environment.

Nice-to-Haves

  • None explicitly listed beyond the requirements.

Compensation and Benefits

  • Weekly lunch stipend of $75/£75 or equivalent.
  • Full health and dental benefits, including a separate budget for mental health.
  • RRSP matching, 401K, Pension Scheme.
  • 100% Parental Leave top-up for up to 6 months, for either parent.
  • Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
  • Education & learning stipend for conferences, courses, and coaching.
  • 6 weeks of paid vacation (30 working days!).
  • Budget for traveling to other offices if you are remote, plus an annual company offsite.
  • $500 home office stipend.
  • Co-working benefit for those not near an office.

Skills

KubernetesPyTorchTensorFlowJAXTerraformArgo CDPrometheusGrafanagpu infrastructurehpc
WorkWhile

Engineering Manager

WorkWhileSan Francisco, CA +2

Lead and grow a small team of engineers while staying hands-on with architecture, technical strategy, and product initiatives. Requires 2+ years managing engineers plus strong experience with modern web technologies in a B2B SaaS or marketplace environment.

250k – 340k/yr
Hybrid5+ YOEEngineering Management
Databricks

Engineering Manager, Foundation Model Inference

DatabricksMountain View, CA

Lead and grow a team of infrastructure engineers building Databricks' Foundation Model Inference and APIs for serving partner and self-hosted LLMs at scale. Shape roadmaps for real-time, provisioned, and batch inference while maintaining high reliability, operational excellence, and technical quality.

190k – 261k/yr
On-site5+ YOEEngineering Management
Airtable

Engineering Manager, AI Tooling

AirtableSan Francisco, CA

Lead an 8-12 person engineering team building AI Tooling and Field Agents at Airtable to drive adoption, usage, and first-party AI token consumption. Requires strong people management, product sense, experimentation rigor, and cross-functional partnership with PM, Design, Data Science, and GTM teams.

281k – 366k/yr
Hybrid5+ YOEEngineering Management
Snowflake

Software Engineer - Notebooks

SnowflakeBellevue, WA

Lead a 10-engineer team building Snowflake's Developer Experience for data engineering workloads, including IDEs (Workspaces, Cortex Code), CLI, Pipeline Builder, dbt integration, and DCM. Requires 3-5 years management experience, strong product instincts, hands-on coding, and cross-functional collaboration.

236k – 339k/yr
Hybrid5+ YOEEngineering Management
Loop Returns

Engineering Manager, Support & Stability

Loop ReturnsColumbus, OH +4

Lead the Support & Stability engineering team responsible for highest-risk application domains, proactive reliability, AI-driven automation of support issues, and first-line triage. Requires 3+ years engineering management experience with emphasis on building stability programs, using AI as default, and hands-on technical depth.

150k – 190k/yr
Remote5+ YOEEngineering Management