Staff Software Engineer - Infrastructure
Build and operate reliable, observable multi-cloud infrastructure for the dbt platform across AWS, Azure, and Google Cloud. The role emphasizes automation, Kubernetes administration, infrastructure as code, cloud cost optimization, developer experience, and operational reliability.
About the job
Responsibilities
- Design, operate, and support infrastructure systems with parity across single-tenant and multi-tenant models and AWS, Azure, and Google Cloud environments.
- Work with engineering teams to deploy services consistently across cloud environments.
- Apply cloud infrastructure expertise to strengthen and scale multi-cloud capabilities.
- Partner with Architecture, Release Engineering, Product Engineering, and Security teams.
- Design and build automation to eliminate manual toil and streamline infrastructure operations at scale.
- Implement infrastructure optimizations that reduce cloud spend without sacrificing reliability.
- Participate in a balanced on-call rotation, continuously improve tooling, and reduce operational toil.
Requirements
- 10+ years of experience with the clouds, tools, and languages in the technology stack, especially AWS, Azure, Google Cloud, Terraform, Kubernetes, Python, and Bash.
- Solid experience with declarative infrastructure as code, ideally Terraform, or experience with CloudFormation or ARM/Bicep.
- Experience automating tasks.
- Experience working asynchronously as part of a fully remote, distributed team.
- Excellent communication and writing skills.
- Prior experience in a multi-cloud environment.
- Extensive Kubernetes administration and troubleshooting experience.
Tools and Technologies
- Terraform
- Kubernetes
- Python
- Bash
- Helm
- Argo CD
- Go
- Datadog
- AWS
- Azure
- Google Cloud
- CloudFormation
- ARM/Bicep
Benefits
- 100% employer-paid medical insurance, subject to country and worker type.
- Generous paid time off, paid sick time, inclusive parental leave, holidays, a year-end Global Week of Rest, and volunteer days off.
- RSU stock grants, subject to country and worker type.
- Professional development and training opportunities.
- Virtual happy hours, free food, and team-building activities.
- Monthly cell phone stipend.
- Access to a mental health support platform offering therapy, coaching, and self-guided mindfulness resources for covered employees and dependents.
Skills
AWS, Azure, GCP, Terraform, Kubernetes, Python, Bash, Helm, Argo Cd, Go, Datadog, CloudFormation, Arm/Bicep
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.