Technical Program Manager, Infrastructure
Coordinates complex infrastructure programs across developer productivity, reliability, and cross-functional teams at Anthropic, driving scaling, security, and AI research support. Requires 5+ years TPM experience in ML/AI or distributed systems with deep technical knowledge.
About the job
What you’ll do:
Developer Productivity & Tooling
- Drive cross-functional programs to improve developer environments, CI/CD infrastructure, and release processes that enable rapid innovation while maintaining high security standards
- Coordinate large-scale migrations and platform modernization efforts across engineering teams
- Partner with teams to measure and improve developer productivity metrics, identifying bottlenecks and driving systematic improvements
- Lead initiatives to integrate AI tools into development workflows, helping Anthropic be at the forefront of AI-assisted research and engineering
Infrastructure Reliability & Operations
- Drive programs to establish and achieve reliability targets across training infrastructure and production services
- Coordinate incident response improvements, post-mortem processes, and on-call rotations that help teams operate effectively
- Establish metrics and dashboards to track infrastructure health, capacity utilization, and operational excellence
Cross-functional Coordination
- Serve as the critical bridge between infrastructure teams, research, and product, translating technical complexities into clear updates for a variety of audiences
- Consult with stakeholders to deeply understand infrastructure, data, and compute needs, identifying solutions to support frontier research and product development
- Drive alignment on priorities and timelines across teams with competing constraints
You May Be a Good Fit If You
- Have 5+ years of technical program management experience, with a track record of successfully delivering complex infrastructure programs in ML/AI systems or large-scale distributed systems
- Have deep technical understanding of infrastructure systems—enough to engage substantively with engineers, identify technical risks, and add value beyond project tracking
- Excel at creating structure and processes in ambiguous environments, bringing clarity to complex cross-team initiatives
- Have strong stakeholder management skills and can build trust with both technical and non-technical partners
- Are comfortable navigating competing priorities and using data to drive technical decisions
- Have experience with developer productivity initiatives, CI/CD systems, or infrastructure scaling
- Thrive in fast-paced environments and can balance strategic planning with tactical execution
- Are obsessed with reliability, scalability, security, and continuous improvement
- Have a passion for supporting internal partners like research to understand their unique needs
- Are passionate about AI infrastructure and understand the unique challenges of building and operating systems at frontier scale
- Experience with Kubernetes, cloud platforms (AWS, GCP, Azure), and ML infrastructure (GPU/TPU/Trainium clusters)
- Background working with research teams and translating their needs into concrete technical requirements
- Experience driving adoption of AI tools to improve engineering productivity
- Familiarity with observability tooling and practices
Annual Salary: $290,000 — $365,000 USD
Skills
Kubernetes, AWS, GCP, Azure, CI/CD, ML Infrastructure, Gpu Clusters, Tpu, Trainium, Observability
Similar jobs
Technical Program Management jobsLeads end-to-end software programs for AI accelerator systems, coordinating internal teams, strategic partners, and vendors from design through data center deployment. The role requires experience managing data center system products and complex software programs at scale.
Leads cross-functional programs that translate model and product needs into scalable, reliable data platform, database, and storage infrastructure. The role requires deep cloud storage and data-stack expertise, architectural judgment, and experience driving complex production programs through adoption.
Leads cross-functional programs that improve developer velocity and deployment excellence across ChatGPT, including CI, testing, release infrastructure, progressive rollouts, and operational validation. Requires strong technical program leadership, systems thinking, and data-driven execution across engineering teams.
Leads cross-functional programs for Chat capacity planning and model deployment, connecting demand forecasting, serving capacity, launch readiness, rollout coordination, and post-deployment learning. Requires technical program leadership in infrastructure, distributed systems, model serving, or large-scale deployment environments.
Leads technical strategy and cross-functional execution for enterprise products and AI workflows, influencing architecture, launch readiness, security, governance, and customer adoption. The role requires strong technical fluency, enterprise software experience, product judgment, and influence across engineering and business functions.