# Member of Technical Staff - Training Platform

**Company:** [Prime Intellect](https://hotfix.jobs/companies/prime-intellect)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $150k – $300k/yr
**Skills:** Kubernetes, Helm, Python, FastAPI, React, TypeScript, Next.js, GCP, Terraform, Prometheus, Grafana, Keda, Nccl, Sqlalchemy, GitOps
**Posted:** 2026-07-08

> Build and operate a hosted AI training platform spanning Kubernetes GPU orchestration, Python control-plane services, developer-facing APIs, and monitoring interfaces. The role requires depth across AI infrastructure, distributed training, cloud operations, and full-stack platform development.

## Job Description

## Responsibilities
- Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets.
- Build and maintain Helm charts for reproducible training stacks comprising trainers, inference servers, environment servers, and supporting services.
- Develop Python control-plane agents that monitor pods, report run state, and synchronize clusters.
- Implement scheduling and autoscaling for heterogeneous GPU hardware using KEDA, LeaderWorkerSet, taints/tolerations, and gang scheduling.
- Maintain GitOps workflows through pull requests, Helm values, and continuous integration.
- Build node-local model caches, checkpoint pipelines, and shared storage for fast cold starts.
- Operate observability systems and debug GPU clusters.
- Build developer-facing hosted-training features, including job submission, live run monitoring, logs, metrics, model and adapter management, and comparisons.
- Develop FastAPI backend services and REST APIs connecting the platform to running clusters.
- Build real-time monitoring and debugging tools, including streaming logs, step-level metrics, and failure analysis.
- Ship product UI using Next.js, React, TypeScript, Tailwind, shadcn, tRPC, and TanStack Query.
- Interface with RL trainers, inference servers, and environment servers.
- Productize new model architectures, reinforcement-learning algorithms, and training modes.

## Requirements
- Strong knowledge of modern AI stacks, open model families, fine-tuning methods, and inference engines.
- Familiarity with GPU hardware tradeoffs, including H100, H200, B200, NVLink, interconnects, and memory hierarchy.
- Understanding of distributed training fundamentals, including data, tensor, pipeline, and expert parallelism, NCCL, and multi-node scheduling.
- Strong Kubernetes operations experience, including Helm, CRDs, operators, KEDA, gang scheduling, and GPU Operator.
- Experience debugging production clusters with kubectl, pod lifecycle analysis, node troubleshooting, and networking.
- Cloud platform experience; GCP, GCS, GKE, Cloud Run, and Cloud Tasks are preferred.
- Experience with infrastructure automation and GitOps.
- Observability experience with Prometheus, Grafana, Loki, OpenTelemetry, and DCGM.
- Linux fundamentals, including networking, namespaces, and performance tuning.
- Strong Python backend development experience with FastAPI, asynchronous programming, and SQLAlchemy.
- Experience building Python control-plane agents that interact with Kubernetes APIs.
- Modern frontend development experience with TypeScript and React/Next.js.
- Experience designing REST and tRPC APIs.
- Experience building developer tools, dashboards, and live-monitoring interfaces.

## Compensation and Benefits
- Cash compensation of $150,000–$300,000 with significant equity.
- Flexible work arrangement with remote work or a San Francisco office option.
- Visa sponsorship and relocation support.
- Professional development budget for courses and conferences.
- Team off-sites and conference attendance.

## Similar jobs

- [AI Engineer, Enablement](https://hotfix.jobs/jobs/6ae315a1-a605-45dc-b1ee-88fa8e9dee24) - LangChain - New York, NY - $150k – $195k/yr
- [Member of Technical Staff — Frontier Data](https://hotfix.jobs/jobs/20efe17f-aa7c-401b-8add-7086d4571fba) - Roboflow - Remote - $150k – $300k/yr
- [Algorithm Engineer](https://hotfix.jobs/jobs/3bef67e8-d4e7-4866-a9b8-6c7863fe2e96) - Beacon Biosignals - Remote - $150k – $170k/yr
- [Software Engineer - Prediction and Planning ML](https://hotfix.jobs/jobs/21b9c778-e1ae-4695-b26d-fec68ea8a8cc) - Applied Intuition - Sunnyvale, CA - $151k – $258k/yr
- [Software Engineer, AI Platform](https://hotfix.jobs/jobs/7dcee5ac-38bb-4da0-b399-0b5896976722) - Fab2 - Austin, TX - $140k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/8a2ba710-ad93-4211-98f9-5771d2b9dad0
**Canonical:** https://hotfix.jobs/jobs/8a2ba710-ad93-4211-98f9-5771d2b9dad0