[Pipeline] Staff+ Software Engineer, Experimentation
Staff+ Software Engineer owning the strategy, architecture, and development of Anthropic's configuration management, feature flagging, and large-scale experimentation platforms to enable safe, data-driven changes and boost developer productivity.
About the job
Responsibilities
- Own the technical strategy and roadmap for config and experimentation infrastructure, translating team-level goals into concrete execution plans and partnering with teams focused on deployment, testing, and more.
- Maintain and enhance the architecture for feature flagging, dynamic configuration, and experimentation systems, ensuring the hardest problems get solved — whether by you directly or by working through others.
- Design and build scalable, reliable distributed infrastructure and shared libraries that support high-volume experimentation and config workloads across all engineering teams.
- Own and evolve the platforms and tooling that let engineers and researchers safely ship config changes, run experiments, and measure impact.
- Define standards, tooling, and frameworks for experimentation and configuration management that drive developer productivity across research and production workloads.
Requirements
- Deep experience with configuration management, feature flagging, and/or experimentation platforms in a large-scale environment.
- Strong proficiency in Python, Rust and/or Go.
- Obsessed with developer productivity and reducing friction in how teams configure, test, and ship changes.
- Experience with container orchestration and infrastructure at scale.
- Excellent communication skills and enjoy supporting internal partners to improve their development experience.
- Excited about designing foundational systems and comfortable working independently on ambiguous, high-impact technical challenges.
- Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role.
Nice-to-Haves
- 15+ years (not including internships or co-ops) of experience in a Software Engineer role, building and operating large-scale developer infrastructure.
- 3+ years (not including internships or co-ops) of experience leading large scale, complex projects or teams as an engineer or tech lead.
- Experience with config or experimentation platforms such as Statsig, GrowthBook, LaunchDarkly, Optimizely, Unleash, or similar (including building such systems in-house).
- Experience designing experimentation frameworks — randomization, assignment, metrics pipelines, and statistical analysis for A/B testing at scale.
- Experience with dynamic configuration systems and safe rollout mechanisms (gradual rollouts, kill switches, targeting rules).
- Experience building CLI tools, developer-facing services, and APIs/automation workflows that integrate config and experimentation into existing CI/CD pipelines.
Skills
Python, Rust, Go, Feature Flagging, Configuration Management, Experimentation Platforms, Container Orchestration, Distributed Infrastructure, A/B Testing, CI/CD
Similar jobs
DevOps / SRE jobsBuild and operate portable infrastructure that enables Claude to run reliably across multiple cloud providers and accelerator platforms. The role requires 8+ years of distributed-systems experience, multi-cloud architecture expertise, production programming, Kubernetes, and Infrastructure as Code proficiency.
Staff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.