Research Engineer, Machine Learning Systems
Builds scalable ML training systems and infrastructure for speech AI models (STT/TTS), prototypes novel ideas with researchers, and creates internal tools for cross-functional teams. Requires strong ML research pipeline experience, especially in speech domains, plus orchestration tools expertise.
About the job
Key Responsibilities
- Scalable Model Training: Architect and manage horizontally scalable systems to accelerate end-to-end training lifecycle for STT and TTS models, including data preparation, high-throughput pipelines, distributed infrastructure, and automated evaluation tooling.
- Tooling & Accessibility: Design and implement internal UIs and tools making ML systems accessible to non-technical stakeholders.
- Infrastructure & Tools: Oversee training tooling, job orchestration, experiment tracking, and data storage.
It's Important to Us That You Have
- Strong experience with the machine learning research pipeline, particularly in STT or related speech domains, experimenting with architectures, and implementing large-scale training systems.
- Proficiency with orchestration and infrastructure tools like Kubernetes, Docker, and Prefect.
- Familiarity with ML lifecycle tools such as MLflow.
- Experience building internal tools or dashboards for non-technical users.
- Hands-on experience with data engineering practices for unstructured audio and text data.
- Comfortable working in cross-functional teams.
Skills
Machine Learning, Speech-To-Text, Stt, Text-To-Speech, Tts, Kubernetes, Docker, Prefect, MLflow, Data Engineering, Distributed Training, Ml Pipelines
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.