Software Engineer, Platform Infrastructure
Builds and scales control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, Kubernetes deployments, and accelerator integration. Requires 3+ years of production software experience, distributed-systems expertise, and proficiency in Go and Python.
About the job
Responsibilities
- Design, build, and scale services that orchestrate Ray clusters across cloud and on-premises environments, supporting VM-based and Kubernetes-based deployments.
- Optimize control-plane components for large-scale distributed AI/ML workloads.
- Build intelligent scheduling and resource-management systems for heterogeneous compute clusters.
- Develop features that improve the reliability, performance, scalability, and observability of managed Ray workloads.
- Support and optimize accelerator integration, including GPUs and TPUs.
- Handle container image management and dependency resolution for distributed workloads.
- Participate in code reviews and design and architecture discussions.
- Provide on-call support and collaborate with customer and field teams to troubleshoot infrastructure issues.
- Collaborate with distributed-systems and machine-learning experts to advance AI infrastructure.
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years of experience writing high-quality production code.
- Hands-on experience building and maintaining highly available, scalable, and performant distributed systems.
- Expertise in cloud-native technologies and Kubernetes-based deployments.
- Strong understanding of networking, security, and authentication mechanisms in cloud environments.
- Familiarity with observability stacks such as Prometheus and Grafana.
- Proficiency in Go and Python.
- Knowledge of low-level operating-system foundations, including the Linux kernel, file systems, and containers.
Skills
Ray, Kubernetes, AWS, Azure, GCP, Go, Python, Prometheus, Grafana, Linux, Containers, Distributed Systems, Networking, Cloud Security, Gpus
Similar jobs
Backend Engineering jobsDevelop and improve Ray Core’s C++ distributed-systems backend, focusing on performance, reliability, fault tolerance, and scalability. The role requires at least five years of experience with distributed systems, C/C++, low-level operating systems, algorithms, and system design.
Design and lead secure, scalable wallet infrastructure spanning custody, key management, signing, authorization, and recovery. The role requires deep expertise in wallet or cryptographic security infrastructure, distributed systems architecture, and modern blockchain account models.
Build and operate production backend features, APIs, integrations, and services while collaborating asynchronously across product and engineering teams. The role requires strong backend fundamentals, independent feature ownership, and practical experience with modern backend languages and AI coding tools.
Build and operate production backend features, APIs, integrations, and services for a large-scale DevSecOps platform. The role requires professional backend development experience, strong database and testing fundamentals, independent feature ownership, and clear asynchronous communication.
Build and operate a globally distributed CI/CD execution platform, including branch orchestration, ephemeral build environments, scheduling, caching, and pipeline observability. Requires Go and TypeScript expertise plus at least 3 years building cloud infrastructure or distributed systems.