Skip to content

Manager, Platform Engineering

Leads a hands-on platform engineering team responsible for AWS infrastructure, Kubernetes deployment paths, developer self-service, CI/CD governance, reliability, and audit readiness. The role requires deep infrastructure experience, Terraform expertise, production Kubernetes operations, and people leadership.

About the job

Responsibilities

  • Lead the platform engineering team, owning delivery, reliability, team health, and engineer development while remaining hands-on.
  • Make infrastructure auditable through environment separation, change control, access management, and audit logging.
  • Migrate incrementally to a standardized Kubernetes deployment path while maintaining operational continuity.
  • Build developer tooling, paved roads, golden paths, and self-service deployment workflows.
  • Own CI/CD pipelines and controls, including branch protection, peer review, security and quality testing, and gated deployment.
  • Establish Terraform infrastructure-as-code standards and bring unmanaged infrastructure under version control and change management.
  • Reduce operational toil through automation and replatforming.
  • Own infrastructure reliability, monitoring, alerting, and incident response.
  • Manage cloud cost and capacity while balancing performance and reliability.
  • Set standards for infrastructure code quality, testing, documentation, and review.
  • Partner with application-platform and domain teams and influence shared infrastructure and security practices.

Requirements

  • Deep, hands-on platform and infrastructure engineering experience across AWS, Linux, and container-first workflows.
  • Production experience running Kubernetes and providing standardized deployment paths for engineering teams.
  • Strong Terraform infrastructure-as-code experience, including managing existing or unmanaged infrastructure.
  • Experience designing and owning CI/CD pipelines with security, quality, and approval gates.
  • Experience building internal developer tooling and self-service deployment workflows.
  • Experience designing auditable infrastructure and change processes, including access control, audit logging, and SOC 2 evidence.
  • Strong observability, monitoring, alerting, and infrastructure incident-response skills.
  • Experience with authentication and authorization platforms such as Okta or Auth0.
  • Ability to improve unfamiliar, legacy, or poorly maintained infrastructure through troubleshooting and root-cause analysis.
  • People leadership experience, including managing, mentoring, and growing engineers while remaining technically hands-on.
  • Ability to influence without direct authority and establish standards adopted by other teams.
  • Ability to lead incremental migrations and build durable engineering practices.
  • Ability to communicate technical concepts and tradeoffs clearly to non-technical stakeholders.
  • Strong English communication skills.

Compensation and Benefits

  • Competitive compensation.
  • Flexible work options.
  • Visa sponsorship is not available.
  • International remote work is not supported.

Skills

AWS, Linux, Kubernetes, Terraform, CI/CD, Infrastructure As Code, Docker, Observability, Monitoring, Incident Response, Okta, Auth0, SOC 2, Access Control, Audit Logging

Crusoe

Crusoe

Denver, CO

Senior Manager, Commissioning
$160k+/yrOn-site10+ YOEEngineering Management

Leads commissioning programs across multiple data center projects, managing commissioning teams, third-party agents, and stakeholder coordination from pre-functional testing through turnover. Requires 10+ years of mission-critical commissioning experience, 5+ years of leadership, engineering knowledge, and a bachelor’s degree.

Shield AI

Shield AI

San Diego, CA
Manager, Software Engineering
$170k+/yrOn-site7+ YOEEngineering Management

Leads a hands-on Shared Services Engineering team building and operating reusable services, SDKs, APIs, and customer-facing systems. Requires 7+ years of software engineering experience, engineering management experience, and strong technical judgment across distributed and full-stack systems.

Forge

Forge

New York, NY
Manager, Site Reliability Engineer
$150k+/yrOn-site8+ YOEEngineering Management

Leads Forge’s Site Reliability Engineering team, overseeing availability, incident response, observability, production operations, and reliability practices. Requires substantial infrastructure or software engineering experience, people leadership, and expertise with cloud platforms and operational tooling.

BuildOps

BuildOps

San Francisco, CA
Senior Engineering Manager, Financial Platform
$172k+/yrHybrid10+ YOEEngineering Management

Leads the architecture and development of BuildOps’ integration and data platform, driving technical strategy, reliability, APIs, databases, and engineering standards across multiple teams. Requires 10+ years of software engineering experience and deep expertise in TypeScript, Node.js, PostgreSQL, cloud infrastructure, and distributed systems.

Crusoe

Crusoe

Shakopee, MN

Senior Manager, Data Center Facility Operations
$175k+/yrOn-site5+ YOEEngineering Management

Leads 24/7 critical facility operations for a 20 MW AI data center expanding to 40 MW, overseeing electrical, mechanical, safety, maintenance, staffing, and operational performance. Requires at least five years of data center operations management experience and expertise in infrastructure reliability and team scaling.