Member of Technical Staff - Research Infrastructure Engineer
Build and operate large-scale research infrastructure for generative AI training, including GPU clusters, distributed systems, telemetry, and reliability tooling. The role requires deep cloud infrastructure expertise, Kubernetes, infrastructure as code, and experience operating large-scale training platforms.
About the job
Responsibilities
- Maintain research infrastructure and optimize application- and infrastructure-level components for peak performance.
- Scale infrastructure for growing research demands while maintaining reliability and performance.
- Collaborate with research teams to understand infrastructure needs and design cost-efficient, high-performance solutions.
- Analyze distributed systems at scale to identify and resolve performance bottlenecks and capacity hotspots.
- Build and evolve telemetry and monitoring systems covering infrastructure performance, utilization, and costs across cloud and datacenter fleets.
- Participate in on-call rotations and incident response.
Requirements
- Experience building or operating large-scale training platforms.
- Experience with large-scale GPU compute clusters.
- Ability to debug performance and reliability issues across large distributed fleets.
- Strong problem-solving skills and ability to work independently.
- Strong communication and collaboration skills.
- Deep knowledge of modern cloud infrastructure, including Kubernetes, infrastructure as code, AWS, and GCP.
- Experience with SLURM.
Technical Focus
- Python
- Bash
- Go
- Kubernetes
- NVIDIA GPU drivers and operators
- OpenTelemetry
- Prometheus
Compensation
- EU base salary: €100,000–€230,000 plus equity.
- US base salary: $150,000–$300,000 plus equity.
Skills
Python, Bash, Go, Kubernetes, Nvidia Gpu Drivers, Nvidia Gpu Operators, OpenTelemetry, Prometheus, AWS, GCP, Infrastructure As Code, Slurm, Distributed Systems, Gpu Clusters
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.