Principal DevOps Engineer
Principal DevOps Engineer responsible for platform engineering across bare-metal and Azure environments, including Kubernetes, CI/CD automation, Jenkins lifecycle management, Infrastructure as Code, networking, reliability, and disaster recovery. Requires 8+ years of core DevOps experience and deep hands-on expertise.
About the job
Responsibilities
- Architect, build, and maintain Kubernetes clusters across bare-metal environments using kubeadm and Rancher, and Azure Kubernetes Service (AKS).
- Design scalable and resilient platform infrastructure.
- Build and maintain CI/CD pipelines using Jenkins, Groovy, and GitHub Actions.
- Own the Jenkins lifecycle, including upgrades, patching, plugin management, migration, and modernization.
- Manage GitHub and GitHub Cloud repositories and workflows.
- Ensure timely upgrades of DevOps tools, platforms, and dependencies.
- Define and implement disaster recovery, backup, and high-availability strategies.
- Implement Infrastructure as Code practices for provisioning and configuration.
- Ensure system security, reliability, and performance.
- Troubleshoot complex distributed-system and networking issues.
- Collaborate with engineering teams on cloud-native adoption.
- Drive DevOps best practices and platform standardization.
Requirements
- 8+ years of core DevOps experience.
- Strong experience with bare-metal Kubernetes, including kubeadm and Rancher.
- Hands-on experience with Azure Kubernetes Service (AKS).
- Deep knowledge of Kubernetes architecture and lifecycle.
- Expert Jenkins experience, including pipelines and Groovy.
- Experience with Jenkins upgrades, migrations, and plugin dependency management.
- Strong GitHub Actions experience.
- Expertise with GitHub and GitHub Cloud.
- Strong Azure knowledge and experience with hybrid environments.
- Strong Infrastructure as Code experience with Terraform, ARM, Bicep, or equivalent.
- Knowledge of Kubernetes networking, including CNI plugins such as Calico, Flannel, and Cilium.
- Experience with ingress controllers such as NGINX, Traefik, and Azure Application Gateway Ingress Controller.
- Knowledge of L4/L7 load balancing, including MetalLB for bare-metal environments.
- Understanding of pod-to-pod and pod-to-service communication.
- Experience with DNS management, including CoreDNS, service discovery, and external DNS integration.
- Understanding of TLS/SSL fundamentals, certificate management, certificate rotation, and mutual TLS concepts.
- Knowledge of Kubernetes network policies and segmentation.
- Understanding of hybrid networking between on-premises environments and Azure, including VPN and ExpressRoute.
- Strong ownership, troubleshooting, independent execution, and communication skills.
Nice to Have
- Experience with Prometheus, Grafana, and ELK.
- Experience with GitOps tools such as Argo CD and Flux.
- Kubernetes or Azure certifications.
- Experience modernizing CI/CD platforms.
- Exposure to disaster recovery planning and execution.
- Experience with AI/ML implementation in DevOps, including AIOps, intelligent automation, or pipeline optimization.
- Experience with service meshes such as Istio or Linkerd.
- Platform engineering experience.
Skills
Kubernetes, Azure Kubernetes Service, Jenkins, Groovy, GitHub Actions, GitHub, Azure, Terraform, Bicep, Rancher, Calico, Cilium, Nginx, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsAs a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.