Senior DevOps Engineer
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.
About the job
Responsibilities
- Design and manage multi-cloud environments across AWS and Google Cloud using Terraform.
- Operate production Kubernetes clusters, including cluster health, resource optimization, and application deployment.
- Build and maintain infrastructure for AI workloads.
- Manage vector search databases, PostgreSQL, MySQL, and MongoDB.
- Implement monitoring and alerting with Datadog and Prometheus.
- Develop and maintain automation using Python, Go, or Bash, including LLM-assisted workflows.
- Manage cloud networking, including VPCs, load balancing, service meshes, and caching with Redis.
- Lead incident response, debugging, root cause analysis, and on-call operations.
Requirements
- 7+ years of experience in infrastructure, DevOps, or site reliability engineering, including at least 5 years in cloud-native environments.
- Strong Linux administration and performance-tuning skills.
- Expert experience with Terraform or OpenTofu, including infrastructure state management at scale.
- Production Kubernetes experience with EKS, GKE, or self-managed clusters.
- Strong Python or Go programming proficiency, with software-engineering practices such as testing, code review, and CI/CD.
- Hands-on experience managing relational and non-relational databases.
- Experience leveraging and securing PaaS and serverless offerings.
- Practical experience using LLM tools such as GitHub Copilot, ChatGPT, or Claude.
- Ability to learn technologies independently and communicate complex technical issues clearly in English.
Compensation
- Compensation details are not provided.
Skills
AWS, GCP, Terraform, Kubernetes, Linux, Python, Go, Bash, Datadog, Prometheus, Postgres, MySQL, MongoDB, Redis, CI/CD
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Leads the design, automation, and reliability of large-scale, multi-cloud infrastructure supporting search, NoSQL, and AI-driven workloads. Requires 7+ years in infrastructure, DevOps, or SRE, plus deep Kubernetes, Terraform, Linux, and distributed data-systems expertise.