Software Engineer, Machine Learning Infrastructure
Builds scalable ML infrastructure services for experimentation, model training, serving, and LLM applications. Requires 2+ years of software development experience plus experience with distributed systems, production ML platforms or MLOps, and high-availability operations.
About the job
Responsibilities
- Design and build scalable, reliable, and secure services for notebooks, ML model training, experimentation, serving, and LLM applications across multiple regions.
- Create services and libraries that enable ML engineers to transition seamlessly from experimentation to production across company systems.
- Work directly with product teams and ML engineers to improve day-to-day productivity.
- Own and solve technical and product challenges across diverse systems, processes, and technologies.
Requirements
- 2+ years of professional software development experience.
- Strong background in service-oriented architecture and large-scale distributed systems.
- Experience across the full software development lifecycle, including user discovery, design, implementation, testing, deployment, and operations.
- Experience working on production ML platforms, MLOps solutions, or LLM applications.
- Experience operating high-availability, low-latency systems.
- Experience partnering with other teams to drive business outcomes.
- Pragmatic approach to selecting and adjusting solutions.
Nice-to-haves
- Experience building and shipping production AI agents.
- Familiarity with LLMs and LLM frameworks.
- Experience training and shipping machine learning models to production for critical business problems.
Skills
Machine Learning, MLOps, LLMs, Llm Frameworks, AI Agents, Distributed Systems, Service-Oriented Architecture, Software Development, Model Training, Model Serving, Notebooks, Production Operations
Similar jobs
ML Engineering jobsBuilds secure, vendor-agnostic agent platform primitives, evaluation tooling, safety controls, retrieval systems, and production workflows. Requires 2+ years of software engineering experience plus strong foundations in LLMs, agentic AI, GPU computing, model serving, and distributed systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate LLM-powered agents that automate business workflows, integrating systems of record through MCP and validating behavior with structured evaluations. Requires 5+ years of software development experience, strong Python and SQL, and hands-on experience shipping agents to production.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.