Skip to content
OPSWATOPSWAT

DevOps Engineer

The DevOps Engineer designs, automates, and supports secure, scalable AWS infrastructure, Kubernetes environments, CI/CD pipelines, and monitoring systems. The role requires 2–4 years of DevOps experience plus expertise in infrastructure as code, testing, troubleshooting, and production support.

About the job

Responsibilities

  • Design, deploy, and manage AWS cloud infrastructure using Terraform and AWS CloudFormation, including EC2, S3, and RDS.
  • Implement and maintain Kubernetes clusters for scalable, highly available container orchestration.
  • Develop and maintain CI/CD pipelines for automated software delivery and deployment.
  • Manage infrastructure as code using Terraform and AWS CloudFormation.
  • Design and implement load-test infrastructure to validate system performance under varying conditions.
  • Collaborate with development and QA teams to integrate automated testing into CI/CD pipelines.
  • Implement monitoring and logging with Prometheus, Grafana, and ELK Stack.
  • Analyze and troubleshoot infrastructure issues to ensure reliability and uptime.
  • Provide production support and incident response, resolving issues promptly to minimize downtime.

Requirements

  • 2–4 years of experience as a DevOps Engineer or in a similar role, focused on AWS cloud infrastructure.
  • Experience administering Kubernetes and working with containerization concepts.
  • Strong understanding of CI/CD principles and hands-on experience with related tools.
  • Experience with infrastructure as code using Terraform and AWS CloudFormation.
  • Experience with test automation frameworks and methodologies.
  • Experience with load, endurance, scalability, and stress testing.
  • Knowledge of Python or Bash for automation.
  • Experience implementing monitoring and logging with Prometheus, Grafana, and ELK Stack.
  • Excellent troubleshooting skills and the ability to diagnose and resolve complex issues.
  • Experience with production support and incident response.
  • Strong communication and collaboration skills.

Nice to Have

  • Experience with Argo.
  • Familiarity with email infrastructure and domain administration.
  • Microsoft 365 or Exchange Online experience, or a prior email administration background.

Benefits

  • Innovative environment focused on cybersecurity.
  • Collaborative culture with valued team input.
  • Ongoing training and professional development opportunities.
  • Mission-driven work protecting critical infrastructure and digital assets.
  • Fun and dynamic atmosphere with team-building and social activities.

Skills

AWS, Kubernetes, Terraform, Aws Cloudformation, CI/CD, Python, Bash, Prometheus, Grafana, Elk Stack, Argo, Load Testing, Test Automation, Amazon Ec2, Amazon S3

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

PostHog

PostHog

Remote

ClickHouse Operations Engineer
No salary listedRemoteDevOps / SRE

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.