Skip to content
RelaceRelace

Infrastructure Engineer

Designs and operates high-performance inference and training infrastructure for ML models, focusing on GPU scheduling, distributed systems, and cloud optimization. Requires 2+ years experience with cloud platforms like AWS/GCP/Azure.

About the job

Responsibilities

  • Architect and manage the infrastructure powering our ultra-fast inference and training stack.
  • Build reliable, efficient systems for deploying and scaling ML workloads globally.
  • Work on GPU scheduling, distributed systems, and high-performance cloud deployments.
  • Optimize performance and cost across compute, networking, and storage layers.
  • Collaborate with world-class engineers to push the limits of what small models can do.

Requirements

  • 2+ years of experience writing high-quality production code
  • Strong experience with cloud infrastructure (AWS, GCP, Azure, or equivalent)
  • Experience with data science and systems optimization
  • Familiarity with ML infrastructure, GPUs, etc. a plus

Skills

AWS, GCP, Azure, Kubernetes, Docker, Gpus, Distributed Systems, ML Infrastructure, Terraform, Linux

Immuta

Immuta

Columbus, OH

Platform & Site Reliability Engineering Internship
$52k+/yrHybridDevOps / SRE

Summer 2027 internship on a Site Reliability Engineering team, building software and automation for deployment, operations, monitoring, and reliability. Requires a software engineering foundation, programming experience, and strong problem-solving and collaboration skills.

DuploCloud

DuploCloud

United States

DevOps Engineer
$80k+/yrRemote2+ YOEDevOps / SRE

Customer-facing DevOps Engineer helping organizations implement secure, compliant cloud infrastructure through the DuploCloud platform. Requires 2–3 years of cloud or DevOps experience, containerization expertise, public cloud knowledge, and strong customer communication skills.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff, Systems Infrastructure
$200k+/yrOn-siteDevOps / SRE

Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.

Fab2

Fab2

Austin, TX
Infrastructure Software Engineering Intern
$114k+/yrOn-siteDevOps / SRE

Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.

Ontic

Ontic

Austin, TX

Associate DevOps Engineer
$100k+/yrHybridDevOps / SRE

Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.