Skip to content
DecagonDecagonSan Francisco, CA

Senior Software Engineer, Cloud Infrastructure

Build and operate scalable cloud infrastructure platforms, abstractions over Kubernetes and major cloud providers, and secure enterprise deployments for Decagon's agentic AI systems. Requires 4+ years in infrastructure/DevOps with deep Terraform, Kubernetes, and cloud networking experience.

200k – 400k/yr
On-site4+ YOEDevOps / SRE

About the role

What you'll do

  • Build the development and production platforms that power our products, and the abstractions over cloud infrastructure, Kubernetes, and networking that let engineers ship without becoming infrastructure experts. Ensure it scales to the next order of magnitude as usage grows.
  • Take end-to-end ownership of deployment architecture in customer-owned cloud environments (VPC configuration, permissioning, networking, provisioning) and the full lifecycle: setup, upgrades, scaling, and incident support. Build runbooks and automation to make it repeatable.
  • Treat monitoring, alerting, and rollback as first-class parts of anything shipped. Own the reliability of systems our AI agents depend on in production, where latency, availability, and graceful degradation shape customer experience.
  • Partner directly with customers' platform, security, and DevOps teams to navigate infrastructure and compliance constraints, and with Product, Security, Sales, and Customer Success teams to turn requirements into deployment plans.

Requirements

  • 4+ years building and operating core infrastructure, platform engineering, or infrastructure/DevOps, ideally with customer-facing deployment experience.
  • Deep experience with a major cloud provider (GCP, AWS, or Azure), along with Terraform and Kubernetes at scale.
  • Strong grasp of cloud networking fundamentals (VPCs, IAM, DNS, load balancing) and how they surface as deployment constraints.
  • Track record operating production systems reliably: monitoring, on-call, incident response, and reasoning about failure modes upfront.
  • Comfort navigating ambiguity across stakeholders from engineers to security and compliance teams, turning conversations into actionable plans.
  • Clear technical writing and track record of driving adoption across teams.
  • Comfortable in a fast-moving environment with rapid change.

Nice-to-haves

  • Experience managing deployments in customer-owned cloud environments, including security reviews, compliance requirements, and change management.
  • Experience building internal platforms or paved roads: service templates, self-serve environments, CI/CD pipeline design, deployment automation.
  • Familiarity with observability and incident management in distributed systems (Prometheus, Grafana, Datadog, or similar).
  • Infrastructure-as-code with a security-minded approach to supply chain (provenance, secrets, least privilege).
  • Experience operating latency-sensitive or ML/AI-serving workloads in production.
  • Experience using AI-assisted tooling to make yourself and your team more effective.

Compensation and Benefits

Compensation $200K – $400K + equity. This range reflects expected compensation. Determined based on experience, skills, and scope of responsibilities, with flexibility for exceptional impact. In addition to base salary, competitive equity is offered. Final compensation may vary based on location within the United States.

Benefits

  • Medical, Dental, and Vision benefits for you and your family
  • Life Insurance and Disability Benefits
  • Retirement Plan (e.g., 401K, pension)
  • Parental Leave
  • Fertility and family building benefits through Carrot
  • Monthly stipend to support wellness, lifestyle, and work-life balance
  • Daily lunches and snacks in the office
  • Take what you need vacation policy (subject to local requirements)

Skills

KubernetesTerraformAWSGCPAzurecloud networkingvpcIAMCI/CDMonitoringObservabilityPrometheusGrafanaDatadog

Similar roles

DevOps / SRE jobs
Glean

Lead Site Reliability Engineer

GleanPalo Alto, CA

Leads SRE team to ensure high availability, scalability, and reliability of cloud services through automation, incident management, and technical leadership. Requires 8+ years SRE experience, team management, and expertise in cloud platforms and containerization.

200k – 260k/yrHybrid8+ YOEDevOps / SRE
Snowflake

Senior Software Engineer, Snowpark Container Service

SnowflakeBellevue, WA +1

Senior Software Engineer building Snowflake's Snowpark Container Services, a managed Kubernetes-based platform for running containerized applications inside Snowflake's Data Cloud. Lead engineering efforts on highly scalable, reliable, multi-tenant container compute infrastructure.

200k – 288k/yrHybrid7+ YOEDevOps / SRE
Together AI

Senior Software Engineer, Observability

Together AISan Francisco, CA

Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.

200k – 280k/yrOn-site5+ YOEDevOps / SRE
Baseten

Software Engineer

BasetenSan Francisco, CA

Lead Software Engineer building real-time observability and root-cause analysis for Baseten's large-scale GPU fabrics and high-performance networks. Requires strong distributed systems and infrastructure experience to correlate telemetry from switches, NICs, GPUs, and inference services.

200k – 380k/yrHybrid7+ YOEDevOps / SRE
Snowflake

Senior Software Engineer, Accelerated Delivery

SnowflakeMenlo Park, CA

Build and evolve continuous deployment, progressive delivery, and AI-augmented release platforms for safe, large-scale multi-cloud rollouts on Kubernetes at Snowflake. Requires strong systems programming (Golang/Java/C++), Kubernetes, observability, and DevOps experience.

200k – 288k/yrHybrid5+ YOEDevOps / SRE