What You’ll Do
- Design, build, and operate backend services in Go or Python that power Lightning AI's managed infrastructure platform.
- Develop control plane services that provision, orchestrate, and manage Kubernetes and Slurm clusters across large-scale GPU infrastructure.
- Build distributed systems that automate cluster lifecycle management, workload scheduling, infrastructure provisioning, and platform operations.
- Develop platform capabilities using Kubernetes APIs, controllers, operators, and other cloud-native technologies.
- Improve the reliability, scalability, security, and observability of our managed platform through automation and operational excellence.
- Diagnose and resolve complex production issues across Kubernetes, distributed systems, networking, and cloud infrastructure.
- Collaborate with infrastructure, AI, and platform engineering teams to shape the future of our cloud platform.
- Contribute to technical design, architecture, mentoring, engineering best practices, and on-call operations.
What You’ll Need
Required Qualifications
- Significant professional experience designing, building, and operating production backend systems using Go or Python.
- Deep hands-on experience with Kubernetes or Slurm, including operating large-scale production environments.
- Strong understanding of distributed systems, cloud-native architectures, and production infrastructure.
- Experience designing and building scalable backend services, APIs, and automation for infrastructure or platform operations.
- Strong understanding of networking, storage, and cloud infrastructure fundamentals.
- Familiarity with observability, CI/CD, testing, production operations, and incident response.
- Ability to own complex technical projects while collaborating effectively across engineering teams.
Ideal Experience
- Kubernetes platform development, including operators, controllers, CRDs, or control plane components
- Slurm administration, scheduling, or HPC environments
- Infrastructure-as-Code or GitOps tooling such as Terraform, Crossplane, Helm, Argo CD, Flux, or Kustomize
- GPU infrastructure, AI platforms, or large-scale compute environments
- Kubernetes networking, storage, or multi-tenant platform architecture
- Event-driven systems, gRPC, or distributed messaging
- Building CI/CD systems for production environments
- Contributions to cloud-native or open source infrastructure projects
Compensation
The anticipated annual base salary range for this role is: $180,000–$250,000 USD
In addition to base salary, our total rewards package for eligible roles includes a discretionary bonus, a meaningful equity component, and comprehensive benefits.