DevOps II
Builds and manages scalable infrastructure, CI/CD pipelines, containerized deployments, and monitoring systems. The role requires 4+ years of DevOps experience, strong Kubernetes and scripting skills, cloud platform knowledge, and familiarity with AI/ML deployment workflows.
About the job
Responsibilities
- Design, implement, and manage CI/CD pipelines for application deployment.
- Work with containerization and orchestration tools, especially Kubernetes.
- Automate infrastructure provisioning and configuration management.
- Monitor system performance, reliability, and availability.
- Collaborate with development and QA teams for seamless releases.
- Troubleshoot and resolve infrastructure and deployment issues.
- Support AI/ML model deployment and basic MLOps workflows.
Requirements
- 4+ years of experience in DevOps or related roles.
- Strong hands-on experience with Kubernetes.
- Proficiency in scripting languages such as Python, Bash, or Shell.
- Experience with CI/CD tools, including Jenkins, GitHub Actions, or GitLab CI.
- Knowledge of cloud platforms such as AWS, Azure, or GCP.
- Experience with containerization tools such as Docker.
- Understanding of Infrastructure as Code tools such as Terraform or CloudFormation.
Nice-to-Haves
- Basic knowledge of AI/ML concepts and model deployment.
- Familiarity with MLOps tools and workflows.
- Experience with monitoring tools such as Prometheus, Grafana, or the ELK stack.
- Knowledge of security best practices in DevOps.
Skills
Kubernetes, Python, Bash, Shell, Jenkins, GitHub Actions, Gitlab Ci, AWS, Azure, GCP, Docker, Terraform, CloudFormation, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.