Staff / Senior Software Engineer, Infrastructure
Builds and operates scalable infrastructure systems including Kubernetes clusters, distributed databases, and cloud services to support AI music platform at consumer scale. Requires 5+ years experience in infrastructure engineering with strong ownership and scaling expertise.
About the job
What You’ll Do
- Architect and build services to handle massive consumer traffic, data, and usage
- Design systems that are performant, secure, scalable, and easy to observe
- Own systems end-to-end — from design and implementation through deployment, monitoring, and operational excellence
- Lead by example on engineering excellence — code quality, system design, documentation, and operational maturity
- Collaborate with engineering teams across the company to understand their needs and build the right abstractions
- Communicate proactively and with high bandwidth — concise and information-dense async updates, stakeholder alignment, low entropy
- Operate with ambiguity — scope what matters, make tradeoffs, and drive projects forward independently
What You’ll Need (Required)
- 5–7+ years of infrastructure, backend, or systems engineering experience
- Experience building and operating systems at significant scale in production
- Strong understanding of distributed systems, cloud services (AWS/GCP), and modern infrastructure patterns
- Experience with some combination of: Kubernetes, Docker, infrastructure as code (Pulumi/Terraform/CDK), databases (Postgres, distributed relational databases), caching systems, or container orchestration
- Ability to reason through hard scaling, reliability, and performance problems with clear technical judgment
- High ownership — you drive projects end-to-end without waiting for direction
- Strong communication skills — you keep stakeholders informed and reduce ambiguity for the teams you serve
- An obsession with engineering excellence, iterating and learning rapidly, and working hard
Strong Plusses
- Deep experience with Kubernetes at scale — cluster management, control plane scaling, multi-tenancy
- Experience with large-scale databases, distributed data layers, or storage systems
- Experience with ML infrastructure — inference serving, ML data pipelines, MLOps, GPU infrastructure
- Experience on a platform or developer experience team where your primary customers were other engineers
- Experience building internal systems 0→1 (auth, notifications, CDN, DevEx tooling, or similar)
- Strong oncall instincts — triage, debug, and resolve incidents across a distributed stack
- Hands-on familiarity with AI tooling and the current landscape of AI for software engineering — models, agents, coding assistants, and agentic workflows
- Golang or Rust experience, especially for large-scale systems
- Experience with websockets, CDNs, streaming traffic patterns, and audio/video delivery
- Security best practices in building and scaling infrastructure
- Technical leadership or management experience
Skills
Kubernetes, Docker, AWS, GCP, Terraform, Pulumi, Postgres, Go, Rust, Distributed Systems
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.