Senior ML Operations (MLOps) Engineer
Build and operate scalable ML infrastructure for deploying models to IoT sleep devices. Own end-to-end pipelines, optimize performance, and collaborate cross-functionally. Requires 5+ years in ML ops, Python, AWS, and production ML deployment.
About the job
Responsibilities
- Pioneer cutting-edge ML technologies, integrating them into products and processes for health monitoring.
- Own design and operation of robust ML infrastructure, building scalable data, model, and deployment pipelines.
- Partner with R&D, firmware, data, and backend teams for reliable ML inference at scale.
- Optimize compute, storage, and deployment resources for training and inference.
- Develop tooling, microservices, and frameworks for data processing, experimentation, and deployment.
Requirements
- 5+ years software engineering experience focused on ML infrastructure, distributed systems, or large-scale data processing in Python (e.g., PyTorch, TensorFlow).
- Hands-on experience with ML workflow orchestration and CI/CD pipelines for model deployment.
- Experience shipping ML models to production at scale, handling telemetry, monitoring, and feedback loops.
- Strong experience with AWS (Lambda, ECS, DynamoDB, CloudWatch) or equivalent for serving and monitoring ML systems.
- Adaptive problem-solving in fast-paced, collaborative environments.
Nice-to-Haves
- Expertise in real-time ML workflows and streaming systems (e.g., Kinesis, Kafka, Flink).
- Optimizing ML infrastructure for efficiency, latency, and cloud cost.
- Understanding of secure ML operations, privacy, and compliance for health/IoT data.
- Familiarity with health, wellness, or IoT domains, especially wearables or medical-grade devices.
Skills
Python, PyTorch, TensorFlow, AWS, AWS Lambda, ECS, DynamoDB, CloudWatch, CI/CD, Ml Workflows, Kubernetes, Kafka, Kinesis, Flink
Similar jobs
ML Engineering jobsBuild and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Leads production deployment and optimization of large language models on GPU hardware, focusing on quantization, inference engines, serving infrastructure, and performance benchmarking. Requires a bachelor's degree and 5+ years of software engineering experience in ML infrastructure, LLM inference, or model optimization.
Build and ship autonomous, agentic software development lifecycle capabilities, including AI agents, orchestration, and safety guardrails. The role requires senior software engineering experience, proficiency in Ruby, Go, or Python, distributed systems knowledge, and experience with AI/ML applications.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.