Principal Software Engineer - Performance Engineering
Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.
About the job
Responsibilities
- Lead evaluation and enablement of new AWS, Azure, and Google Cloud instance types, including Graviton, AMD Turin, Azure ARM/Cobalt, and GCP Axion.
- Set performance strategy, technical roadmaps, and measurement methodologies; translate benchmark results into rollout and adoption recommendations.
- Partner with engineering, capacity, finance, cloud providers, and silicon vendors to qualify new hardware and influence future instance designs.
- Analyze CPU microarchitecture, memory bandwidth, PMU counters, and I/O behavior to identify workload bottlenecks and drive resolution.
- Develop price/performance models supporting hardware transitions.
- Build automated performance-validation workflows and per-provider, per-instance scorecards for day-zero hardware readiness.
- Mentor engineers, expand benchmarking capabilities, and promote performance-engineering education.
Requirements
- 12+ years of experience in performance engineering, systems engineering, or infrastructure engineering, with principal-level technical leadership.
- Deep expertise in cloud infrastructure performance across at least one major cloud service provider, including instance types, pricing models, and capacity constraints.
- Strong hardware and systems fundamentals, including CPU microarchitecture, memory bandwidth, I/O subsystems, and profiling tools such as PMU counters.
- Experience building or operating large-scale benchmarking systems and converting benchmark output into pricing and rollout decisions.
- Ability to build quantitative models connecting technical performance metrics to business outcomes.
- Experience leading cross-functional and cross-company initiatives and serving as a technical point of contact for cloud and hardware partners.
- Bias toward automation, including scorecards, dashboards, and validation pipelines.
Nice to Have
- Experience with data warehouse or distributed database performance.
- Experience partnering with cloud providers or silicon vendors on early-access hardware programs.
- Experience building hardware-emulation systems.
Skills
AWS, Microsoft Azure, GCP, Cpu Microarchitecture, Memory Bandwidth, I/O Subsystems, Pmu Counters, Benchmarking, Performance Modeling, Price/Performance Modeling, Distributed Databases, Automation, Dashboards, Validation Pipelines
Similar jobs
DevOps / SRE jobsThis principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.
Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.