Member of Technical Staff, Cloud Infrastructure
Build and optimize backend infrastructure that powers high-performance generative AI workloads. The role requires experience scaling enterprise machine-learning systems and working with ML infrastructure such as PyTorch, Vertex AI, or SageMaker.
About the job
Responsibilities
- Design and build core backend software components, ensuring efficiency, scalability, and stability of system resources.
- Conduct design and code reviews and collaborate with cross-functional teams.
- Continuously analyze and optimize infrastructure efficiency for AI workloads, including compute, storage, and networking.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
- 3+ years of experience working in machine learning infrastructure, such as PyTorch, Vertex AI, or SageMaker.
- Experience building, scaling, and optimizing enterprise-grade machine learning systems.
Nice-to-Haves
- Experience in AI or large-scale infrastructure.
- Master’s or PhD degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
Benefits
- Solve challenging problems at the forefront of AI infrastructure, including low-latency inference and scalable model serving.
- Work with emerging technology that shapes how businesses and developers use AI.
- High ownership and direct impact in a fast-growing team.
- Collaborate with experienced engineers and AI researchers.
Skills
Machine Learning Infrastructure, PyTorch, Google Vertex Ai, Amazon Sagemaker, Artificial Intelligence, Machine Learning Systems, Backend Software, Cloud Infrastructure, Compute, Storage, Networking, Model Serving
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.
Conduct research and engineering on next-generation video generation models, including post-training, evaluation, multimodal data systems, and scalable ML infrastructure. The internship requires current bachelor's or master's study and practical experience with Python, data pipelines, or distributed systems.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.