Staff+ Software Engineer, Developer Productivity
Leads technical strategy and builds scalable developer infrastructure including build systems, CI/CD pipelines, and tooling for large monorepo environments. Requires 3+ years leading complex projects, proficiency in Python/Rust/Go, and experience with container orchestration.
About the job
Responsibilities
- Own the technical strategy and roadmap for your area, translating team-level goals into concrete execution plans
- Define infrastructure architecture, ensuring the hardest problems get solved — whether by you directly or by working through others
- Design and build scalable, reliable distributed infrastructure and shared libraries that support high-volume workloads across all engineering teams
- Own and evolve build environments, package management, and dependency systems to enable fast, reproducible builds
- Define and implement language ecosystem standards, tooling, and frameworks that drive developer productivity across research and production workloads
Requirements
- 3+ years (not including internships or co-ops) of experience leading large scale, complex projects or teams as an engineer or tech lead
- Deep experience with build systems, CI/CD pipelines, and/or developer tooling in a large monorepo environment
- Strong proficiency in Python, Rust and/or Go
- Obsessed with developer productivity and reducing friction in the software development lifecycle
- Experience with container orchestration and infrastructure at scale
- Excellent communication skills and enjoy supporting internal partners to improve their development experience
- Excited about designing foundational systems and comfortable working independently on ambiguous, high-impact technical challenges
Nice-to-Haves
- 15+ years (not including internships or co-ops) of experience in a Software Engineer role, building and operating large-scale developer infrastructure
- Experience with CI orchestration tools (Buildkite, Jenkins, GitHub Actions, or similar) and merge queue management at scale
- Experience building or operating remote build execution systems (Bazel Remote Execution API, BuildBarn, BuildBuddy, or similar)
- Experience with Nix/NixOS/Docker and managing large image / package sets at scale
- Experience building CLI tools, developer-facing services, and GitHub API and automation workflows
Skills
Python, Rust, Go, CI/CD, Build Systems, Bazel, Buildkite, Jenkins, GitHub Actions, Nix, Docker, Kubernetes, Monorepo, Remote Build Execution
Similar jobs
DevOps / SRE jobsBuild and operate portable infrastructure that enables Claude to run reliably across multiple cloud providers and accelerator platforms. The role requires 8+ years of distributed-systems experience, multi-cloud architecture expertise, production programming, Kubernetes, and Infrastructure as Code proficiency.
Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.