Staff Software Engineer - Developer Productivity
Leads cross-team initiatives that improve developer productivity through internal platforms, self-service tooling, cloud automation, observability, and AI-assisted workflows. The role requires strong Python or Go experience, cloud-platform expertise, and a record of delivering ambiguous technical projects.
About the job
Responsibilities
- Lead end-to-end projects that improve developer tooling, environments, and operational workflows for speed, reliability, and safety.
- Design and operate internal tooling and services in Python and/or Go.
- Build tooling for the internal developer platform, including self-service workflows, shared patterns, and developer-facing interfaces.
- Enable developer workflows across AWS, Azure, and Google Cloud.
- Develop AI-based workflows that improve developer experience and reduce repetitive manual work, including infrastructure-operations efficiency.
- Improve developer feedback loops through observability using metrics, logs, and traces.
- Design, build, and document paved paths and self-service capabilities.
- Improve reliability through automation and observability, and drive incident follow-ups to reduce recurrence.
- Set technical direction through design documents, tradeoff analysis, and sequencing; drive delivery across stakeholders.
- Mentor engineers through design and code reviews.
Requirements
- Strong experience building and operating production services in Python and/or Go.
- Experience with at least one cloud provider—AWS, Azure, or Google Cloud—and core primitives such as IAM, networking, and compute.
- Experience building developer productivity tooling, including CI systems, build tooling, developer portals, or workflow automation.
- Track record leading ambiguous, cross-team technical projects to completion.
- Strong communication and engineering judgment.
Nice to Have
- Kubernetes or platform-abstraction experience.
- Infrastructure as code with Terraform or OpenTofu, including reusable modules, safe state practices, and modernization of shared platform foundations.
- Experience building internal developer platforms, developer portals, or self-service tooling.
Perks and Benefits
- Employer-paid medical insurance.
- Paid time off, paid sick time, inclusive parental leave, holidays, and volunteer days.
- RSU stock grants.
- Professional development and training opportunities.
- Virtual happy hours, free food, and team-building activities.
- Monthly cell phone stipend.
- Mental health support platform with therapy, coaching, and mindfulness resources.
Skills
Python, Go, AWS, Microsoft Azure, GCP, IAM, Kubernetes, Terraform, Opentofu, CI/CD, Developer Portals, Observability, Infrastructure Automation, Networking, Compute
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.