Staff Software Engineer - Cloud Infrastructure
Staff Software Engineer designs and scales cloud infrastructure, managing Kubernetes fleets, multi-region recovery, and distributed systems to support massive growth. Requires 8+ years experience with deep expertise in AWS, Kubernetes, and backend languages like Python/Go/Java.
About the job
What you will do
- Design and implement the next generation of our cloud infrastructure to support a 4x increase in system load, ensuring resilience and performance at scale.
- Own the evolution of our K8s ecosystem, ensuring our core monolith and growing microservices environment remain performant and developer-friendly.
- Leverage your product development background to build infrastructure primitives that make it easier for product teams to ship complex features seamlessly.
- Tackle high-order problems related to global consistency, database contention, and high-throughput background processing.
- Work deeply with infrastructure to manage complex distributed systems that power every Rippling product.
- Drive technical consensus on complex, cross-functional initiatives, acting as a bridge between infrastructure and product engineering.
- Engineer high-efficiency resource strategies, optimizing Kubernetes utilization and AWS spend to ensure our infrastructure scales sustainably with our business growth.
- Act as a force multiplier for the Cloud team, setting the standard for code quality and architectural patterns that will define the next decade of Rippling.
What you will need
- 8+ years of software engineering experience, with a proven track record of architecting and scaling infrastructure.
- Background in product development with an understanding that infrastructure exists to empower developers and serve the end-user.
- Deep experience with distributed systems and a strong grasp of where standard architectural patterns break under extreme load.
- Exceptional communication skills with the ability to influence technical strategy across multiple engineering teams.
- Hands-on expertise with Kubernetes orchestration, including custom controllers, networking, and sophisticated scaling strategies.
- Ability to approach a massive, complex codebase as a fascinating engineering puzzle that requires world-class architectural thinking.
- High proficiency in modern backend languages (Python, Go, or Java), with a desire to master Python and the ability to dive deep into any stack to solve a bottleneck.
Skills
Kubernetes, AWS, Python, Go, Java, Distributed Systems, Microservices
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.