Skip to content
BasetenBaseten

Software Engineer, Model Performance Tooling

Builds performance benchmarking, diagnostic, and optimization tools for LLM inference on GPU clusters. Early-career role requiring Python proficiency, systems curiosity, and interest in AI hardware—no prior experience needed.

About the job

The Opportunity

Early-career Software Engineers to build automated performance benchmarking and diagnostic tools for AI infrastructure at the intersection of high-performance computing and LLM engineering. Focus on tearing apart models to optimize performance on hardware.

Responsibilities

  • Performance Benchmarking: Run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse).
  • Infrastructure Validation: Create automated acceptance tests for new GPU clusters across x86 and ARM systems, measuring GPU memory bandwidth, networking throughput, and multi-node networking performance.
  • Model Dev Experience: Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces).
  • Tool Development: Build and contribute to tools such as InferenceMAX and genai-bench to automate model evaluation and optimization.
  • Deep Hardware Profiling: Use PyTorch Profiler and NVIDIA Nsight Systems to collect performance profiles, identify bottlenecks, and debug the NVIDIA compute/networking stack.
  • Monitoring & Observability: Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance.
  • Continuous Integration: Automate performance testing via CI/CD pipelines to catch regressions in model setups before they hit production.
  • Optimization Automation: Build tools to find the "Pareto frontier"—identifying the absolute best configuration (latency vs. cost vs. quality) for a given model and workload.

What We're Looking For

Fresher-friendly role emphasizing curiosity and technical depth over experience.

  • Love for systems & hardware (GPU memory, InfiniBand).
  • Automation mindset and passion for stress-testing.
  • Mathematical curiosity in Transformers, FLOPs, memory.
  • Interest in optimization (quantization, speculative decoding).
  • Technical Toolkit: Python; eagerness for NVIDIA stack; C++ nice-to-have.

Benefits

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents.
  • Generous PTO policy including company wide Winter Break.
  • Paid parental leave.
  • Company-facilitated 401(k).

Skills

Python, PyTorch, Nvidia Nsight Systems, GPU, InfiniBand, CI/CD, LLMs, Benchmarking, Kubernetes, C++

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff
$160k+/yrOn-siteML Engineering

Build, deploy, and optimize AI applications and machine learning models for customer use cases while contributing to an internal ML platform. This new graduate role requires a technical master’s degree, hands-on ML or LLM experience, and strong customer communication skills.

Nuro

Nuro

Mountain View, CA

Software Engineer, ML Inference Platform
$160k+/yrOn-site1+ YOEML Engineering

Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.

Nuro

Nuro

Mountain View, CA

Software Engineer, ML Infrastructure Platform
$160k+/yrOn-site1+ YOEML Engineering

Build and operate the infrastructure powering large-scale machine-learning training for autonomous-driving systems. The role requires Python proficiency, Kubernetes production experience, distributed-systems expertise, and ownership of reliability, observability, and operational maturity.

Garner Health

Garner Health

New York, NY

Applied Scientist II
$158k+/yrHybrid2+ YOEML Engineering

Build and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.

Ambral

Ambral

New York, NY
Member of Technical Staff
$165k+/yrOn-site1+ YOEML Engineering

Build replayable enterprise environments, evaluation systems, graders, and post-training workflows for AI agents. The role spans machine-learning research and production engineering and requires 1–7 years of software or ML systems experience.