Lead Member of Technical Staff, Inference Infrastructure
Leads architecture and strategy for deploying optimized NLP models in high-throughput, low-latency production environments using Kubernetes and cloud platforms. Mentors engineers and designs custom customer deployments with 8+ years infrastructure experience.
About the job
Responsibilities
- Provide technical leadership across multiple teams, driving architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, high availability environments.
- Serve as key point of contact for customers, leading design of customized deployments.
- Mentor engineers to raise technical bar.
Requirements
- 8+ years engineering experience running production infrastructure at large scale, with technical leadership track record.
- Experience leading architecture/design of large, highly available distributed systems with Kubernetes and GPU workloads.
- Deep expertise with Kubernetes dev/production coding/support, setting team-wide standards.
- Extensive experience across GCP, Azure, AWS, OCI, multi-cloud/on-prem/hybrid environments.
- Lead design, deployment, support, troubleshooting of complex Linux-based computing environments at scale.
- Own compute/storage/network resource and cost management at organizational level.
- Expertise in computational characteristics of accelerators (GPUs, TPUs, custom), leveraging for latency/throughput improvements.
- Deep knowledge of distributed systems, establishing patterns/practices.
- Proficiency in Golang, C++ or similar for high-performance scalable servers.
Nice-to-Haves
- Exceptional collaboration, communication, mentoring, cross-functional leadership.
- Grit and adaptability for complex technical challenges.
Skills
Kubernetes, GCP, Azure, AWS, Oci, Go, C++, Linux, Gpus, Tpus, Distributed Systems
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.