Skip to content

Software Engineer, Inference Platform

Software engineer building and operating the orchestration layer for a globally distributed, high-performance AI inference platform on custom wafer-scale hardware.

About the job

Responsibilities

  • Help shape the technical direction for the Inference Platform, including k8s custom resource definitions, failure domains, service boundaries, and system evolution over time; own the roadmap for major technical areas.
  • Architect active-active systems with rapid failover, graceful degradation, and clear SLOs. Drive system-level improvements in latency, throughput, capacity efficiency, and resilience under unpredictable demand.
  • Write and review production code in the most important parts of the platform. Make high-consequence architectural decisions within your area and set the technical bar through design reviews, code reviews, and sound engineering judgment.
  • Lead on the hardest production issues and cross-system bottlenecks. Drive observability, incident response, capacity planning, and post-incident improvement with a high standard for operational rigor.
  • Partner with ML, Product, Infrastructure, and Cloud teams to translate product and business requirements into scalable system designs, and drive alignment on shared technical decisions within your domain and adjacent platform surfaces.

Requirements

  • 3+ years of experience in software engineering, with experience building and operating large-scale distributed systems or cloud infrastructure.
  • Experience in distributed systems, ideally with Kubernetes.
  • Experience building highly available, latency-sensitive systems at scale.
  • Experience with security (certificates, TLS, mTLS).
  • Experience optimizing latency, throughput, and efficiency in high-QPS systems.
  • Strong proficiency in backend or systems languages such as Go, C++.

Nice-to-Haves

  • Experience with TTFT and tail-latency reduction.
  • Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads.

Skills

Kubernetes, Go, C++, Distributed Systems, Tls, Mtls, Observability, High Availability, Latency Optimization, Cloud Infrastructure

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.