Skip to content
NuroNuro

Senior Software Engineer, Performance Tooling and Infrastructure

Owns performance simulation platform infrastructure for validating autonomy code changes on robot hardware, including benchmarking orchestration, data pipelines, observability, and fleet management. Requires 5+ years experience in Python/C++, Linux systems, data engineering, with technical leadership.

About the job

Responsibilities

  • Develop and maintain job orchestration layer for scheduling, executing, and validating autonomy performance benchmarks across physical bench-top systems, integrated into CI/CD pipelines.
  • Build monitoring, alerting, and self-healing automation for the bench fleet; track utilization, failure rates, and capacity trends.
  • Design and build end-to-end data pipelines capturing performance metrics (CPU/GPU utilization, memory bandwidth, E2E latency, scheduling jitter) and surface insights via dashboards and regression detection.
  • Collaborate with Data Science on statistical analysis and experimentation for non-deterministic workloads.
  • Guide SRE on OS and system-level configuration of bench hardware (Linux kernel tuning, boot infrastructure, networking).
  • Own planning lifecycle for benchmarking fleet, negotiate hardware allocation, and present trade-off recommendations.
  • Partner with Hardware Engineering, NPI, SRE, Perception, Behavior, and Data Science teams.

Requirements

  • 5+ years of industry software engineering experience.
  • Strong proficiency in Python and working proficiency in C++.
  • Experience building data pipelines, ingestion, transformation, storage, and visualization; familiarity with SQL.
  • Deep knowledge of Linux systems (kernel configuration, boot issues, systemd, bare-metal infrastructure, networking, storage).
  • Technical leadership: setting vision, roadmap, driving stakeholder alignment, briefing leadership.
  • AI-native workflow using agentic tooling.
  • Bias for action in ambiguous environments.

Nice-to-Haves

  • Performance engineering tools: perf, Perfetto, pprof, eBPF, NVIDIA Nsight Systems, NVIDIA CUPTI.
  • Experience in robotics or AV, especially NVIDIA DriveOS.

Compensation

  • Base pay: $183,000 - $275,000 (depending on experience, qualifications, education, location, skills).
  • Eligible for annual performance bonus, equity, and competitive benefits.

Skills

Python, C++, Kubernetes, GCP, BigQuery, Grafana, Linux, SQL, Ebpf, Nvidia Nsight Systems

Duolingo

Duolingo

New York, NY
Senior Site Reliability Engineer
$183k+/yrOn-site5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale distributed systems, infrastructure, reliability, and incident response. Requires 5+ years of SRE or DevOps experience plus programming and container orchestration expertise.

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.