Senior Site Reliability Engineer, Platform Infrastructure
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
About the job
Responsibilities
- Design, build, and scale services that orchestrate Ray clusters across cloud and on-premises environments, supporting VM- and Kubernetes-based deployments.
- Optimize control-plane components for large-scale distributed AI/ML workloads.
- Build intelligent scheduling and resource-management systems for heterogeneous compute clusters.
- Improve the reliability, performance, scalability, and observability of managed Ray workloads.
- Support and optimize accelerator integrations, including GPUs and TPUs.
- Manage container images and dependency resolution for distributed workloads.
- Participate in code reviews and design and architecture discussions.
- Provide on-call support and troubleshoot infrastructure issues with customer and field teams.
- Collaborate with distributed-systems and machine-learning experts on AI infrastructure.
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years of experience writing high-quality production code.
- Experience building and maintaining highly available, scalable, and performant distributed systems.
- Expertise in cloud-native technologies and Kubernetes-based deployments.
- Deep understanding of networking, security, and authentication mechanisms in cloud environments.
- Familiarity with observability stacks such as Prometheus and Grafana.
- Proficiency in Go and Python.
- Knowledge of Linux kernel foundations, file systems, and containers.
Skills
Kubernetes, AWS, Azure, GCP, Go, Python, Distributed Systems, Prometheus, Grafana, Linux, Containers, Networking, Cloud Security, Authentication, Gpus
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.