Infrastructure Engineer
Designs and operates high-performance inference and training infrastructure for ML models, focusing on GPU scheduling, distributed systems, and cloud optimization. Requires 2+ years experience with cloud platforms like AWS/GCP/Azure.
About the job
Responsibilities
- Architect and manage the infrastructure powering our ultra-fast inference and training stack.
- Build reliable, efficient systems for deploying and scaling ML workloads globally.
- Work on GPU scheduling, distributed systems, and high-performance cloud deployments.
- Optimize performance and cost across compute, networking, and storage layers.
- Collaborate with world-class engineers to push the limits of what small models can do.
Requirements
- 2+ years of experience writing high-quality production code
- Strong experience with cloud infrastructure (AWS, GCP, Azure, or equivalent)
- Experience with data science and systems optimization
- Familiarity with ML infrastructure, GPUs, etc. a plus
Skills
AWS, GCP, Azure, Kubernetes, Docker, Gpus, Distributed Systems, ML Infrastructure, Terraform, Linux
Similar jobs
DevOps / SRE jobsSummer 2027 internship on a Site Reliability Engineering team, building software and automation for deployment, operations, monitoring, and reliability. Requires a software engineering foundation, programming experience, and strong problem-solving and collaboration skills.
Customer-facing DevOps Engineer helping organizations implement secure, compliant cloud infrastructure through the DuploCloud platform. Requires 2–3 years of cloud or DevOps experience, containerization expertise, public cloud knowledge, and strong customer communication skills.
Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.
Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.