Leads the architecture and delivery of hyperscale AI infrastructure spanning bare metal, virtualization, containers, kernel I/O, and high-performance networking. Requires 12+ years of infrastructure experience, deep Linux and virtualization expertise, GPU/NIC optimization, and significant technical leadership or industry contributions.
260k – 340k/yr
On-site12+ YOEBackend Engineering
About the role
Responsibilities
Architect systems for Bare-Metal-as-a-Service (BMaaS), delivering GPU throughput through InfiniBand/RDMA fabrics for large-scale training.
Design optimized virtualization layers using KVM or custom micro-VMs, providing enterprise-grade isolation with minimal virtualization overhead.
Build a high-performance container substrate using Kubernetes or Slurm for AI workloads across heterogeneous GPU nodes.
Lead architectural design for SR-IOV, RDMA, and virtualized GPU scheduling.
Lead R&D workstreams to productionize novel approaches to memory, networking, and compute management.
Draft white papers and RFCs defining the compute and networking stack roadmap.
Debug complex I/O-path race conditions and optimize kernel-level memory pinning for GPU clusters.
Represent the company in open-source communities and industry forums.
Requirements
12+ years of experience designing and shipping core infrastructure at a major hyperscaler or specialized HPC cloud.
Authoritative knowledge of the Linux kernel, virtualization internals, KVM, QEMU, Firecracker, and high-performance networking including RoCE v2 and InfiniBand.
Experience designing software that maximizes NVIDIA or AMD GPU and high-speed NIC performance.
Experience leading cross-functional teams through ambiguous projects and delivering production-ready, mission-critical systems.
Significant contributions to the field through patents, major open-source contributions, or published distributed-systems research.
Excellent communication skills, including explaining systems concepts to technical and executive audiences.
Bachelor's or master's degree in Computer Science, Computer Engineering, or a related analytical field, or equivalent professional experience.
Nice-to-haves
Patents related to network virtualization, GPU scheduling, or distributed file systems.
Maintainer status or significant contributions to the Linux kernel, Kubernetes, or specialized HPC projects.
Experience optimizing infrastructure for large language model training and inference at scale.
Peer-reviewed publications at systems venues such as OSDI, SOSP, NSDI, or SC.
IETF RFCs, significant open-source contributions, or industry white papers on distributed systems, high-performance networking, or AI infrastructure.
Compensation and Benefits
Compensation range: $260,000–$340,000, plus significant equity and bonus.
Restricted Stock Units.
Paid time off and paid holidays.
Comprehensive health, dental, and vision insurance.
Employer HSA contributions.
Paid parental leave.
Paid life insurance and short- and long-term disability coverage.
Professional development and tuition reimbursement.
Mental health and wellness support.
Commuter benefits for parking and transit.
Cell phone stipend.
401(k) retirement plan with company match up to 4% of salary.
Volunteer time off.
Skills
linux kernelkvmqemufirecrackerKubernetesslurmInfiniBandrdmaroce v2sr-iovgpu schedulingnvidia gpusamd gpusC++
Designs and builds fault-tolerant, scalable distributed systems for Snowflake's metadata, Snowgrid, and data sharing features. Requires 15+ years experience, strong CS fundamentals, Java fluency, and ability to mentor juniors while solving performance and scale challenges.
264k – 380k/yrOn-site15+ YOEBackend Engineering
Principal Software Engineer I - Snowhouse Foundation
SnowflakeMenlo Park, CA
Leads design and implementation of highly available distributed platforms and pipelines for Snowflake's petabyte-scale data warehouse. Requires 15+ years in distributed systems and data infrastructure, with cloud expertise in AWS, Azure, GCP.
264k – 380k/yrOn-site15+ YOEBackend Engineering
Principal Software Engineer, Applied AI
Invisible TechNew York, NY
Principal Software Engineer builds scalable backend systems for AI/ML operations, deploys production ML models, and collaborates with clients on use case discovery and AI solutions. Requires 8+ years experience in ML engineering, Python, cloud platforms, and client-facing work.
250k – 366k/yrHybrid8+ YOEBackend Engineering
Principal Software Engineer, Money Infrastructure
ReplitFoster City, CA
Leads design and implementation of core money infrastructure including pricing, billing, payments, and monetization for Replit's platform. Requires 8+ years backend experience, expertise in financial systems like subscriptions or usage-based billing.
250k – 340k/yrHybrid8+ YOEBackend Engineering
Principal Software Engineer
AstronomerNew York, NY
Provides principal-level technical leadership for a self-contained, cloud-agnostic platform that provisions and orchestrates Apache Airflow in enterprise and air-gapped environments. Requires 10+ years of software engineering experience, control-plane and API architecture expertise, Kubernetes proficiency, and strong distributed-systems design skills.