Skip to content
6sense6sense

Staff Software Engineer - Infrastructure/DevOps

The role leads the design, automation, security, and reliability of multi-region cloud infrastructure and Kubernetes platforms. It requires extensive software and infrastructure engineering experience, strong AWS and infrastructure-as-code expertise, and proficiency in Python, Go, or Bash.

About the job

Responsibilities

  • Architect, build, and scale core infrastructure systems.
  • Automate infrastructure lifecycle to minimize human intervention.
  • Design secure connectivity across services, VPCs, accounts, and regions.
  • Debug and resolve complex production issues across the stack.
  • Write production-quality code, tools, and frameworks.
  • Collaborate with engineering teams to standardize infrastructure usage.

Core Infrastructure and Cloud Platform

  • Design and evolve infrastructure on Amazon Web Services and Kubernetes.
  • Build and scale multi-region and multi-account architectures.
  • Implement secure and scalable service connectivity, including VPC, cross-account, and cross-region connectivity.

Security and Zero Trust

  • Drive identity-based access and eliminate shared credentials.
  • Implement secrets management with HashiCorp Vault.
  • Enforce policies with Open Policy Agent.
  • Design secure service-to-service communication and fine-grained IAM and access control systems.

Infrastructure as Code and Automation

  • Build infrastructure using Terraform, Pulumi, and Ansible.
  • Create reusable modules, abstractions, and fully automated provisioning workflows.
  • Integrate infrastructure automation with GitHub Actions and Jenkins.

Cluster and Infrastructure Lifecycle

  • Own the lifecycle of Kubernetes clusters, including EKS, and databases such as RDS, Aurora, and ElastiCache.
  • Automate cluster upgrades, migrations, node scaling, and patching.

Multi-Region and Migration Engineering

  • Design region-agnostic infrastructure patterns.
  • Enable automated region bootstrapping and failover.
  • Lead region, account, and cloud migration initiatives.

Observability and Reliability

  • Build and operate metrics, logging, and tracing systems using Prometheus, Grafana, the ELK Stack, and Datadog.
  • Improve system reliability, alerting, and incident response.

Governance and Standards

  • Define and enforce resource-tagging strategies and data-classification policies.
  • Ensure cost visibility and auditability.

Developer Productivity and Platform

  • Build self-service infrastructure workflows.
  • Improve developer experience through automation and tooling.
  • Contribute to internal platform evolution.

Requirements

  • 10+ years of experience in software engineering or equivalent experience.
  • 5+ years of experience in infrastructure engineering.
  • Strong hands-on experience with Amazon Web Services, Kubernetes, and Terraform or Pulumi.
  • Strong understanding of networking, including VPC, routing, DNS, and connectivity.
  • Strong understanding of IAM and access control.
  • Experience automating infrastructure workflows.
  • Experience with highly available and scalable systems.
  • Proficiency in Python, Go, or Bash.

Nice-to-Haves

  • Multi-region or multi-account architecture experience.
  • Experience automating infrastructure migrations and upgrades, including EKS upgrades.
  • Experience with secrets-management platforms such as HashiCorp Vault.
  • Familiarity with policy-as-code tools such as Open Policy Agent.
  • Familiarity with observability systems such as Prometheus, Grafana, and Datadog.
  • Exposure to data infrastructure such as Hadoop, Trino, and Spark.
  • Exposure to service meshes such as Istio.

Compensation and Benefits

  • Health coverage.
  • Paid parental leave.
  • Paid time off and holidays.
  • Quarterly self-care days off.
  • Stock options.
  • Equipment and support for working from home or in an office.
  • Learning and development initiatives, including LinkedIn Learning.
  • Wellness education sessions and employee resource group events.

Skills

Amazon Web Services, Kubernetes, Terraform, Pulumi, Ansible, Hashicorp Vault, Open Policy Agent, GitHub Actions, Jenkins, Prometheus, Grafana, Elk Stack, Datadog, Python, Go

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Okta

Okta

Bengaluru, India

Staff DevSecOps Engineer, Enterprise Technology
No salary listedOn-site8+ YOEDevOps / SRE

Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Okta

Okta

Bengaluru, India

Staff Software Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.