Skip to content

Member of Technical Staff - Training Platform

Build and operate a hosted AI training platform spanning Kubernetes GPU orchestration, Python control-plane services, developer-facing APIs, and monitoring interfaces. The role requires depth across AI infrastructure, distributed training, cloud operations, and full-stack platform development.

About the job

Responsibilities

  • Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets.
  • Build and maintain Helm charts for reproducible training stacks comprising trainers, inference servers, environment servers, and supporting services.
  • Develop Python control-plane agents that monitor pods, report run state, and synchronize clusters.
  • Implement scheduling and autoscaling for heterogeneous GPU hardware using KEDA, LeaderWorkerSet, taints/tolerations, and gang scheduling.
  • Maintain GitOps workflows through pull requests, Helm values, and continuous integration.
  • Build node-local model caches, checkpoint pipelines, and shared storage for fast cold starts.
  • Operate observability systems and debug GPU clusters.
  • Build developer-facing hosted-training features, including job submission, live run monitoring, logs, metrics, model and adapter management, and comparisons.
  • Develop FastAPI backend services and REST APIs connecting the platform to running clusters.
  • Build real-time monitoring and debugging tools, including streaming logs, step-level metrics, and failure analysis.
  • Ship product UI using Next.js, React, TypeScript, Tailwind, shadcn, tRPC, and TanStack Query.
  • Interface with RL trainers, inference servers, and environment servers.
  • Productize new model architectures, reinforcement-learning algorithms, and training modes.

Requirements

  • Strong knowledge of modern AI stacks, open model families, fine-tuning methods, and inference engines.
  • Familiarity with GPU hardware tradeoffs, including H100, H200, B200, NVLink, interconnects, and memory hierarchy.
  • Understanding of distributed training fundamentals, including data, tensor, pipeline, and expert parallelism, NCCL, and multi-node scheduling.
  • Strong Kubernetes operations experience, including Helm, CRDs, operators, KEDA, gang scheduling, and GPU Operator.
  • Experience debugging production clusters with kubectl, pod lifecycle analysis, node troubleshooting, and networking.
  • Cloud platform experience; GCP, GCS, GKE, Cloud Run, and Cloud Tasks are preferred.
  • Experience with infrastructure automation and GitOps.
  • Observability experience with Prometheus, Grafana, Loki, OpenTelemetry, and DCGM.
  • Linux fundamentals, including networking, namespaces, and performance tuning.
  • Strong Python backend development experience with FastAPI, asynchronous programming, and SQLAlchemy.
  • Experience building Python control-plane agents that interact with Kubernetes APIs.
  • Modern frontend development experience with TypeScript and React/Next.js.
  • Experience designing REST and tRPC APIs.
  • Experience building developer tools, dashboards, and live-monitoring interfaces.

Compensation and Benefits

  • Cash compensation of $150,000–$300,000 with significant equity.
  • Flexible work arrangement with remote work or a San Francisco office option.
  • Visa sponsorship and relocation support.
  • Professional development budget for courses and conferences.
  • Team off-sites and conference attendance.

Skills

Kubernetes, Helm, Python, FastAPI, React, TypeScript, Next.js, GCP, Terraform, Prometheus, Grafana, Keda, Nccl, Sqlalchemy, GitOps

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.

Roboflow

Roboflow

San Francisco, CA

Member of Technical Staff — Frontier Data
$150k+/yrRemoteML Engineering

Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.

Beacon Biosignals

Beacon Biosignals

Boston, MA
Algorithm Engineer
$150k+/yrRemote4+ YOEML Engineering

Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

Fab2

Fab2

Austin, TX
Software Engineer, AI Platform
$140k+/yrOn-siteML Engineering

Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.