Skip to content
TetraScienceTetraScienceUnited States

Lead Software Platform Engineer, MLOps

Leads the architecture and operation of a multi-tenant AI/ML platform supporting production models, LLMs, and agents in regulated scientific environments. Requires 10+ years in distributed cloud-native systems, strong TypeScript and Python skills, production LLM/RAG experience, and technical leadership.

Salary not listed
Remote10+ YOEML Engineering

About the role

Responsibilities

  • Own the technical architecture of the AI/ML platform and the service and API surface used to run models and agents against scientific data.
  • Own the end-to-end model and prompt lifecycle across Databricks MLflow and AWS Bedrock, including registration, versioning, asset bundles, staged promotion, rollback, and multi-model serving.
  • Design inference infrastructure for real-time and batch workloads, including routing, batching, caching, concurrency control, GPU and accelerator capacity planning, large binary inputs, and graceful degradation.
  • Integrate AI models and LLMs into production systems using retrieval-augmented generation (RAG), tool and function calling, MCP-based tooling, and agent runtimes.
  • Design security controls for guardrails, prompt-injection and tool-abuse defenses, PII and PHI handling, and tenant data boundaries.
  • Build evaluation and quality infrastructure, including offline and online evaluation harnesses, golden datasets, CI regression gates, A/B and shadow deployments, and drift and hallucination detection.
  • Establish monitoring, alerting, logging, and distributed tracing, and define SLI, SLO, and SLA practices for probabilistic systems.
  • Design reproducibility and lineage for validated environments through versioned data, code, prompts, and model artifacts and auditable trails.
  • Contribute to infrastructure-as-code and deployment automation using CloudFormation and AWS CDK, including multi-tenant infrastructure, online upgrades, and on-demand compute allocation.
  • Own production readiness, performance, reliability, cost efficiency, incident response, and runbooks for the AI platform.
  • Lead design reviews, write reference architectures and technical documentation, mentor engineers, and evaluate emerging AI infrastructure and build-versus-buy decisions.

Requirements

  • 10+ years of professional experience in software and infrastructure engineering, designing, building, and scaling distributed cloud-native systems in production.
  • Experience as a technical leader or architect accountable for system design, scalability, performance, and cost optimization.
  • Experience designing security into multi-tenant platforms, including tenant authorization boundaries, PII and PHI handling, prompt injection, and tool-abuse risks.
  • Extensive experience building and maintaining production AI/ML infrastructure as a multi-tenant product for external users.
  • Hands-on experience taking LLM systems to production, including RAG, retrieval and embedding design, prompt and model versioning, and tool or function calling.
  • Expert coding skills in TypeScript and Python for robust APIs and backend services.
  • Production experience with model registries and serving stacks, ideally Databricks MLflow.
  • Experience using AI evaluation as a release gate, including evaluation harnesses, regression gates, and drift or quality monitoring.
  • Proficiency in API-first design, REST, and OpenAPI.
  • Working knowledge of AWS and Docker, plus infrastructure-as-code such as CloudFormation or AWS CDK.
  • Experience with CI/CD pipelines, deployment automation, monitoring, alerting, distributed tracing, and SLI/SLO/SLA practices.
  • Strong communication, cross-functional influence, technical direction, and mentoring skills.

Skills

TypeScriptPythondatabricks mlflowaws bedrockAWSDockerCloudFormationaws cdkREST APIsopenapiRAGmcpdistributed tracingCI/CDmodel serving

Similar roles

ML Engineering jobs
Airbnb

Senior Machine Learning Engineer, Price Modeling

AirbnbUnited States

Senior Machine Learning Engineer building and refining reinforcement learning models for Airbnb's host pricing recommendations. Requires 5+ years ML engineering experience with strong Python skills; hospitality or pricing background is a plus.

196k – 227k/yrRemote5+ YOEML Engineering
Shield AI

Senior Software Engineer, Autonomy Capabilities - Space

Shield AIBoston, MA +3

Build and deploy autonomy software for satellites and missile-defense systems, spanning optimization, tasking, scheduling, track fusion, and motion planning. The role requires strong C++ and Python skills, robotics or unmanned-systems experience, simulation expertise, and the ability to obtain a SECRET clearance.

160k – 240k/yrOn-site5+ YOEML Engineering
Pinterest

Senior Manager, Machine Learning Engineering-Applied Research

PinterestUnited States

Leads a team of machine learning researchers and engineers developing web-scale recommendation systems, guiding strategy, research-to-production execution, and cross-functional delivery. Requires 7+ years of post-graduate academic and industry experience, 3+ years of people management, advanced education, and strong ML publication credentials.

228k – 469k/yrRemote8+ YOEML Engineering
Baselayer

Senior AI Engineer, Agentic Data Enrichment

BaselayerSan Francisco, CA

Build and own production LLM-driven agents that enrich business identities using web discovery, evidence extraction, classification, and risk signals. The role requires strong asynchronous Python, browser automation, multi-provider LLM experience, evaluation methodology, and production agent ownership.

230k – 340k/yrHybrid5+ YOEML Engineering
Airbnb

Senior Machine Learning Engineer

AirbnbUnited States

Build and deploy cutting-edge Agentic AI and LLM systems to transform Airbnb's customer service experience, including Chat and Voice AI assistants. Requires 6+ years experience with production ML/AI systems at scale.

196k – 227k/yrRemote6+ YOEML Engineering