Skip to content
RSARSA

Principal DevOps Engineer

Principal DevOps Engineer responsible for platform engineering across bare-metal and Azure environments, including Kubernetes, CI/CD automation, Jenkins lifecycle management, Infrastructure as Code, networking, reliability, and disaster recovery. Requires 8+ years of core DevOps experience and deep hands-on expertise.

About the job

Responsibilities

  • Architect, build, and maintain Kubernetes clusters across bare-metal environments using kubeadm and Rancher, and Azure Kubernetes Service (AKS).
  • Design scalable and resilient platform infrastructure.
  • Build and maintain CI/CD pipelines using Jenkins, Groovy, and GitHub Actions.
  • Own the Jenkins lifecycle, including upgrades, patching, plugin management, migration, and modernization.
  • Manage GitHub and GitHub Cloud repositories and workflows.
  • Ensure timely upgrades of DevOps tools, platforms, and dependencies.
  • Define and implement disaster recovery, backup, and high-availability strategies.
  • Implement Infrastructure as Code practices for provisioning and configuration.
  • Ensure system security, reliability, and performance.
  • Troubleshoot complex distributed-system and networking issues.
  • Collaborate with engineering teams on cloud-native adoption.
  • Drive DevOps best practices and platform standardization.

Requirements

  • 8+ years of core DevOps experience.
  • Strong experience with bare-metal Kubernetes, including kubeadm and Rancher.
  • Hands-on experience with Azure Kubernetes Service (AKS).
  • Deep knowledge of Kubernetes architecture and lifecycle.
  • Expert Jenkins experience, including pipelines and Groovy.
  • Experience with Jenkins upgrades, migrations, and plugin dependency management.
  • Strong GitHub Actions experience.
  • Expertise with GitHub and GitHub Cloud.
  • Strong Azure knowledge and experience with hybrid environments.
  • Strong Infrastructure as Code experience with Terraform, ARM, Bicep, or equivalent.
  • Knowledge of Kubernetes networking, including CNI plugins such as Calico, Flannel, and Cilium.
  • Experience with ingress controllers such as NGINX, Traefik, and Azure Application Gateway Ingress Controller.
  • Knowledge of L4/L7 load balancing, including MetalLB for bare-metal environments.
  • Understanding of pod-to-pod and pod-to-service communication.
  • Experience with DNS management, including CoreDNS, service discovery, and external DNS integration.
  • Understanding of TLS/SSL fundamentals, certificate management, certificate rotation, and mutual TLS concepts.
  • Knowledge of Kubernetes network policies and segmentation.
  • Understanding of hybrid networking between on-premises environments and Azure, including VPN and ExpressRoute.
  • Strong ownership, troubleshooting, independent execution, and communication skills.

Nice to Have

  • Experience with Prometheus, Grafana, and ELK.
  • Experience with GitOps tools such as Argo CD and Flux.
  • Kubernetes or Azure certifications.
  • Experience modernizing CI/CD platforms.
  • Exposure to disaster recovery planning and execution.
  • Experience with AI/ML implementation in DevOps, including AIOps, intelligent automation, or pipeline optimization.
  • Experience with service meshes such as Istio or Linkerd.
  • Platform engineering experience.

Skills

Kubernetes, Azure Kubernetes Service, Jenkins, Groovy, GitHub Actions, GitHub, Azure, Terraform, Bicep, Rancher, Calico, Cilium, Nginx, Prometheus, Grafana

Fluidstack

Fluidstack

Remote

Principal Operations Engineer, Mechanical
$150k+/yrRemote10+ YOEDevOps / SRE

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Okta

Okta

Bengaluru, India

Staff DevSecOps Engineer, Enterprise Technology
No salary listedOn-site8+ YOEDevOps / SRE

Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Okta

Okta

Bengaluru, India

Staff Software Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.