Build and scale cloud infrastructure for ML training, inference, and data pipelines at a healthcare AI startup. Requires 4+ years in distributed systems, strong Python/TypeScript backend skills, AWS experience, and interest in ML infra (prior MLOps not required).
200k – 230k/yr
Hybrid4+ YOEML Engineering
About the role
Responsibilities
Architect, build, and scale the cloud infrastructure behind our ML training, inference, and data pipelines.
Design resilient systems for model deployment, evaluation, and monitoring that stay reliable as traffic grows.
Own observability across the stack - logging, metrics, tracing, and alerting.
Troubleshoot production issues and continuously improve performance and efficiency.
Collaborate with ML engineers, backend engineers, and cross-functional teams to integrate models cleanly with data pipelines and products.
Requirements
4+ years building and scaling infrastructure in production-distributed systems, cloud platforms, or data engineering.
Strong backend software engineering fundamentals, with proficiency in Python and TypeScript.
Hands-on experience with AWS, and PostgreSQL.
Solid grasp of observability, reliability, and production incident response.
Comfortable with ambiguity and high ownership; you move fast and drive projects from idea to production in a startup environment.
Interested in growing into ML infrastructure.
Nice-to-Haves
Exposure to inference engines (vLLM, SGLang, TensorRT), or Kubernetes.
Build state-of-the-art end-to-end speech and audio generation systems with a focus on joint audio-video modeling. Own audio representations (VAEs, neural codecs), generative backbones (diffusion/flow-matching transformers), conditioning, alignment for voice cloning and sync with video, plus data flywheel, evaluation, and inference optimization for large-scale multimodal models.
200k – 220k/yrRemote7+ YOEML Engineering
Senior Software Engineer, Build
AstronomerNew York, NY
Build and scale Astronomer's AI-powered context layer for data platforms, focusing on semantic search, retrieval, code generation, and applied AI using LLMs and embeddings. Requires 5+ years software engineering experience with Python or Go, strong interest in data/AI tools, and comfort with ambiguity in early-stage product development.
200k – 230k/yrHybrid5+ YOEML Engineering
Senior Software Engineer, Backend
SiftstackMarina Del Rey, CA +1
Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for engineering teams. Own product areas end-to-end, from customer conversations to designing tools, APIs, distributed execution on Kubernetes, evaluation frameworks, and integrating frontier models. Requires 8+ years software engineering experience with backend services and distributed systems.
200k – 250k/yrHybrid8+ YOEML Engineering
Senior Software Engineer
TrabaNew York, NY +1
Build and own production AI agent systems (harnesses, evals, orchestration) on frontier LLMs for industrial supply chain workflows at Traba. Requires 5+ years software engineering with 1+ year shipping LLM/agent features, strong Python/TS, and high-agency in ambiguous customer environments.
200k – 240k/yrHybrid5+ YOEML Engineering
Machine Learning Research Manager
Rad AISan Francisco, CA
Lead and mentor a team of applied and clinical researchers as a player-coach. Guide ML research in NLP, LLMs, and clinical applications for radiology, translating ideas into production systems while partnering with clinicians and engineers. Requires MS/PhD and 6+ years applied ML research experience.