Senior MLOps Engineer - Edge
Build and operate edge MLOps infrastructure for smart-camera machine-learning systems, including model deployment, TensorRT compilation, fleet updates, telemetry, and reliability. The role requires production MLOps experience, embedded inference optimization, and strong collaboration with data-science and embedded-engineering teams.
About the job
Responsibilities
- Build scalable edge infrastructure to deploy machine-learning models to fleets of devices.
- Own the model-compilation platform that converts trained models into optimized, hardware-specific inference engines.
- Manage TensorRT compilation, FP16/INT8 precision trade-offs, calibration, and engine validation.
- Collaborate with Data Scientists, Embedded Engineers, and Product Managers to integrate complex features.
- Implement automation for testing candidate models on production devices.
- Build telemetry pipelines to monitor model drift, thermal impact, and inference latency.
- Develop resilient update mechanisms for low-bandwidth environments, limited storage, and network failures.
- Establish best practices for Python tooling, infrastructure as code, and CI/CD; mentor and guide the team.
Requirements
- Production MLOps experience building and operating model-deployment pipelines.
- Deep experience with CI/CD, Docker, and Linux systems.
- Hands-on experience compiling and optimizing models for embedded hardware, ideally with TensorRT.
- Understanding of precision, quantization, and inference-engine validation at scale.
- Ability to design architectures with graceful failure handling, canary releases, and safe rollbacks.
- Strong collaboration and communication skills with research and embedded-engineering teams.
- Initiative and a bias toward solving problems and filling gaps.
Nice-to-haves
- Experience with NVIDIA's edge ecosystem, including Jetson Orin, DeepStream SDK, and TensorRT.
- Familiarity with video pipelines, GStreamer, or FFmpeg.
- Experience with AWS IoT Greengrass, Balena, or custom OTA and fleet-management solutions.
- Interest in sports technology, video analytics, or performance metrics.
Benefits
- Flexible vacation time, company-wide holidays, meeting-free days, and remote-work options.
- Autonomy and an open, supportive work culture.
- Professional-development resources and career-growth opportunities.
- Well-equipped offices and technology for both office-based and remote work.
- Location-dependent medical and retirement benefits, plus an Employee Assistance Program and employee resource groups.
Skills
Python, Docker, Linux, CI/CD, TensorRT, Jetson Orin, Deepstream Sdk, Fp16, Int8, Quantization, Gstreamer, Ffmpeg, Aws Iot Greengrass, Balena, Infrastructure As Code
Similar jobs
ML Engineering jobsBuild and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Build and deploy embedding, sequence, and language-model representations for Reddit Ads, taking ML projects from requirements and experimentation through production. The role requires 5+ years of end-to-end industry ML experience, with expertise in NLP or computer vision and deep-learning frameworks.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.
Build full-stack AI agent fleets, APIs, workflows, and internal services that automate complex business processes. The role requires at least five years of engineering experience, hands-on LLM framework experience, production AWS expertise, Kubernetes, and strong API and database skills.