Senior SRE
The Senior SRE will design, automate, and operate high-throughput AWS and Kubernetes infrastructure, improving reliability, observability, CI/CD, and cost efficiency. The role requires 4–6 years of production SRE, DevOps, or systems engineering experience and strong Terraform, Linux, scripting, and incident-response skills.
About the job
Responsibilities
- Own and evolve AWS infrastructure as code using Terraform, ensuring environments are reproducible and consistently provisioned.
- Administer production Kubernetes clusters with a focus on cost optimization using Spot instances and fault-tolerant workload design.
- Participate in the on-call rotation, triage live incidents, and drive root-cause analysis through resolution.
- Build and maintain CI/CD automation pipelines using GitLab CI or Jenkins to improve deployment velocity and reliability.
- Strengthen observability across the stack through monitoring, alerting, and dashboards using Prometheus, Grafana, or Datadog.
- Harden core Linux systems.
- Write scripts and tooling in Python, Go, or Bash to eliminate manual toil across the infrastructure organization.
- Use AI developer tools such as Claude Code, GitHub Copilot, and Cursor to accelerate troubleshooting, scripting, and workflow automation.
Requirements
- 4–6 years of SRE, DevOps, or systems engineering experience managing production environments.
- 3+ years of hands-on experience with AWS and Kubernetes administration.
- Proven ability to provision and manage AWS infrastructure reproducibly with Terraform.
- Production Kubernetes administration experience, including cost optimization and fault-tolerance work on Spot instances.
- Track record of on-call ownership, production incident triage, and root-cause analysis.
- Experience designing, automating, and maintaining complex CI/CD pipelines with GitLab CI or Jenkins, including strong Groovy proficiency.
- Deep Linux systems fundamentals and scripting fluency in Python, Go, or Bash.
- Hands-on monitoring-stack experience with Prometheus, Grafana, or Datadog.
- Fluency with AI developer tools such as Claude Code, GitHub Copilot, and Cursor.
Nice-to-haves
- Production Aerospike administration, including memory optimization and data rebalancing; equivalent experience with ScyllaDB or Cassandra is also relevant.
- Apache Kafka operational experience with high-throughput streaming, topic partitioning, and consumer lag; equivalent experience with Pulsar, Kinesis, or RabbitMQ is also relevant.
- ClickHouse administration and tuning for real-time OLAP analytics.
Compensation and Benefits
- Competitive salary including an equity package and quarterly tax assistance for Ukraine tax residents.
- Competitive time off and company wellness days throughout the year.
- Team-building events.
- Annual continued education benefit.
- Newbie Productivity Perk.
- Monthly stipend.
- The role is currently remote but may become hybrid when there is a physical office.
Skills
AWS, Terraform, Kubernetes, Linux, Gitlab Ci, Jenkins, Groovy, Prometheus, Grafana, Datadog, Python, Go, Bash, Apache Kafka, ClickHouse
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
The Senior Cloud Operations Engineer designs, operates, and improves highly available multi-cloud infrastructure, observability, automation, and distributed systems across AWS and GCP. The role requires 5–7 years of DevOps or SRE experience, strong Kubernetes and infrastructure-as-code expertise, and participation in on-call operations.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.