Senior Software Engineer, ML Infrastructure
Leads design and development of scalable, real-time API and data infrastructure for AI model evaluations, processing large-scale event streams with low latency. Requires 5+ years in infrastructure or ML systems, expertise in distributed systems, stream processing, and backend architecture.
About the job
Responsibilities
- Architect and scale high-performance, real-time API and data systems
- Design and implement low-latency pipelines to process and analyze large-scale event streams
- Ensure reliability through robust data integrity, availability, and consistency mechanisms
- Mentor and guide engineers on infrastructure best practices, architecture, and performance tuning
- Collaborate cross-functionally with AI researchers, product leaders, and engineers to anticipate evolving infrastructure needs and deliver resilient, extensible systems
Requirements
- 5+ years of experience in software engineering, with a focus on infrastructure or large-scale data and ML systems
- Deep expertise in distributed systems, stream processing, and scalable backend architecture
- Proven ability to design and operate low-latency, high-throughput, and fault-tolerant systems
- Strong foundation in systems design, performance tuning, and building reliable, fault-tolerant services
- Comfortable in a dynamic, high-ownership, fast-growth environment
Nice-to-haves
- Prior experience with PyTorch model development
Skills
Distributed Systems, Stream Processing, Scalable Backend, Low-Latency Systems, High-Throughput Systems, Fault-Tolerant Systems, System Design, Performance Tuning, PyTorch, ML Infrastructure
Similar jobs
ML Engineering jobsBuild and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.