Senior Software Engineer, Accelerated Delivery
Build and evolve continuous deployment, progressive delivery, and AI-augmented release platforms for safe, large-scale multi-cloud rollouts on Kubernetes at Snowflake. Requires strong systems programming (Golang/Java/C++), Kubernetes, observability, and DevOps experience.
About the job
Responsibilities
- Design and build continuous deployment and rollout infrastructure that safely ships changes across Snowflake’s large-scale, multi-cloud production environment.
- Build and evolve platform capabilities for progressive delivery, including staged rollouts, canarying, automated health checks, rollback controls, and guardrails that reduce blast radius during production change events.
- Improve engineering velocity by removing friction from release pipelines and replacing manual workflows with durable platform abstractions and automation.
- Build internal platforms that support large-scale release orchestration, application rollouts on Kubernetes, and broader production change workflows.
- Partner with product and infrastructure teams to make their services easier to deploy, validate, observe, and operate through well-designed platform capabilities.
- Implement and evolve deployment methodologies such as GitOps-inspired workflows, infrastructure as code, policy-driven automation, and progressive delivery patterns appropriate for Snowflake’s environment.
- Build systems that evaluate rollout health using metrics, logs, alerts, and operational signals to detect regressions early and trigger safe mitigation or rollback paths.
- Develop self-service developer tooling that enables teams across Snowflake to adopt safe deployment patterns without requiring deep release expertise.
- Build automation and guardrails that reduce operational toil and make production change workflows more consistent, scalable, and resilient.
- Design and build AI-assisted, agentic-driven, and increasingly autonomous release workflows that improve rollout intelligence, developer productivity, and deployment safety.
Requirements
- Experience building or operating continuous deployment, release engineering, or production change platforms at scale.
- Experience with Kubernetes-based systems and safely rolling changes across distributed production environments.
- Strong software engineering skills in Golang, Java, C++, or similar systems languages, along with Python, Bash, or similar scripting languages.
- Experience with distributed systems, infrastructure automation, CI/CD pipelines, and cloud environments.
- Data-driven mindset and experience using observability platforms such as Prometheus, Datadog, or Grafana to evaluate system and rollout health.
- Deep care about safe production rollouts, developer experience, minimizing blast radius, and building systems that make the right operational path the easiest one.
- Enjoy building internal platforms and self-service systems that improve developer productivity across a large engineering organization.
- Combined software engineering and DevOps mindset to design, build, and continuously improve large-scale delivery platforms in production.
- Excitement about applying AI and intelligent automation to release operations, deployment safety, and autonomous workflows.
Nice-to-Haves
- Experience applying AI and intelligent automation to release operations, deployment safety, and autonomous workflows.
Skills
Go, Java, C++, Python, Kubernetes, CI/CD, Prometheus, Datadog, Grafana, GitOps, Infrastructure As Code
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Leads design, deployment, and operation of secure distributed cloud systems for public-sector and air-gapped environments. Requires active or obtainable TS/SCI clearance with polygraph, U.S. citizenship, and 7+ years of production experience.