Staff Software Engineer, Cloud Sandboxes
Design and operate scalable, secure cloud sandbox infrastructure powering Docker's agentic platform. Requires 10+ years building large-scale distributed systems, strong Go/Java skills, and deep Kubernetes experience.
About the job
Responsibilities
- Design, implement, and operate core services that power Docker’s Cloud Sandboxes platform
- Build scalable systems for microVM orchestration, workload scheduling, and lifecycle management
- Develop high-performance APIs and control plane components for managing multi-tenant workloads
- Ensure system reliability, observability, and performance across Docker’s Cloud Sandbox infrastructure
- Collaborate with product, platform, and security teams to deliver customer-focused capabilities
- Participate in architectural discussions, code reviews, and design documents
- Contribute to automation and CI/CD improvements across the deployment pipeline
- Debug and resolve production issues across distributed systems in cloud environments
- Take part in on-call rotation for your team; respond to incidents, debug production issues, and drive continuous improvement of system reliability
Requirements
- 10+ years of backend software engineering experience building large-scale cloud or distributed systems
- Strong proficiency in Go and/or Java
- Deep understanding of container orchestration, Kubernetes, and microservices architecture
- Experience designing and operating highly available, secure, and observable production systems
- Strong understanding of cloud infrastructure (AWS, Azure, or GCP) and related scalability patterns
- Familiarity with CI/CD pipelines, monitoring, and infrastructure-as-code tooling
- Excellent problem-solving and debugging skills in distributed environments
- Strong communication skills and ability to collaborate across remote, cross-functional teams
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
Nice-to-Haves
- Experience contributing to cloud-scale compute platforms or container infrastructure products
- Knowledge of service mesh, networking, or policy enforcement systems
- Experience with observability stacks (Prometheus, OpenTelemetry, Grafana, etc.)
- Familiarity with security best practices for multi-tenant cloud systems
- Prior experience in developer infrastructure, cloud platforms, or hyperscale environments
Benefits
- Freedom & flexibility; fit your work around your life
- Designated quarterly Whaleness Days plus end of year Whaleness break
- Home office setup
- 16 weeks of paid Parental leave (after 6 months of employment)
- Technology stipend equivalent to $100 USD net/month
- PTO plan that encourages you to take time to do the things you enjoy
- Training stipend for conferences, courses and classes
- Equity
- Docker Swag
- Medical benefits, retirement and holidays vary by country
- Remote-first culture, with offices in Seattle and Paris
Skills
Go, Java, Kubernetes, AWS, Azure, GCP, CI/CD, Prometheus, OpenTelemetry, Grafana, Microservices, Container Orchestration, Microvm Orchestration, Infrastructure-As-Code, Observability
Similar jobs
Backend Engineering jobsDesigns and operates Internet-scale scanning, DNS, attribution, and data pipelines, with deep ownership of distributed backend systems and production reliability. Requires 10+ years of software engineering experience, strong Go expertise, cloud and streaming infrastructure knowledge, and the ability to mentor engineers.
Staff backend engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production operations experience, distributed-systems expertise, strong communication, and familiarity with Go or willingness to learn it.
Leads architecture and development of mission-critical backend systems for spacecraft command, telemetry, mission planning, and operations. Requires 8+ years of software development experience, distributed-systems and cloud-native expertise, and technical leadership.
Leads the technical direction and hands-on development of scalable backend platforms for Pinterest’s AI-driven products. Requires extensive backend and distributed-systems experience, strong Python skills, production LLM or ML experience, and cross-functional technical leadership.
Leads the technical strategy, architecture, and operation of Pinterest’s large-scale storage infrastructure supporting SQL and graph workloads. Requires 8+ years of backend distributed-systems experience, technical leadership, and expertise in production reliability, performance tuning, and modern database technologies.