Staff+ Software Engineer, Caching
Build and lead Anthropic's managed caching infrastructure as a foundational service, including a scalable Redis fleet, client libraries, and CDC-driven invalidation. Requires deep distributed systems and caching expertise to optimize latency and consistency across hot paths for Claude.
About the job
Key Responsibilities
- Drive the technical direction for caching infrastructure used across Product and Research.
- Design, build, and operate a managed Redis fleet that scales to support millions of users across Claude's product ecosystem.
- Build client libraries and developer-facing abstractions that make correct caching the default for Anthropic engineers.
- Design and operate CDC-driven cache invalidation that keeps cached data consistent with source-of-truth databases.
- Architect caching solutions that operate across GCP, AWS, first-party deployments, and other environments.
- Optimize latency, hit rates, reliability, and cost efficiency on Anthropic's hottest paths.
- Build observability and tooling that makes cache behavior easy to understand and debug.
- Partner with product and research teams to understand access patterns and build infrastructure that accelerates their work.
- Make build-vs-buy decisions for caching technologies.
Minimum Qualifications
- Significant experience as a software engineer building and operating production distributed systems.
- Deep knowledge of caching architectures, including invalidation strategies, consistency tradeoffs, and failure modes.
- Experience operating Redis, Memcached, or similar in-memory data stores in production.
- Proficiency in at least one systems programming language (e.g., Go, Rust, Java, C++) or Python at scale.
- Track record of leading large, complex infrastructure projects as an engineer or tech lead.
- Ability to balance moving quickly with the reliability needs of production systems.
- Strong technical leadership and cross-functional collaboration skills.
Preferred Qualifications
- 10+ years building and scaling distributed infrastructure, with 3+ years leading large-scale projects or teams.
- Experience building managed infrastructure platforms or internal services consumed by many engineering teams.
- Experience with change data capture (Debezium or similar) or streaming data infrastructure.
- Experience operating Redis Cluster, Valkey, ElastiCache, Memorystore, or similar managed offerings at scale.
- Experience designing client libraries or SDKs for internal infrastructure.
- Experience scaling infrastructure through periods of rapid growth at high-growth companies.
- Experience with multi-cloud or hybrid cloud deployments.
- Contributions to caching systems, database internals, or related open source projects.
Compensation and Benefits
- Annual compensation range: $320,000—$485,000 USD.
- Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours.
- Minimum education: Bachelor’s degree or equivalent.
- Location-based hybrid policy: expect staff to be in one of our offices at least 25% of the time.
Skills
Distributed Systems, Redis, Caching Architectures, Cache Invalidation, Cdc, Go, Rust, Java, C++, Python, Debezium, Redis Cluster, Valkey, Elasticache, Memorystore
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.