# Senior SRE Engineer

**Company:** [Stellar Cyber](https://hotfix.jobs/companies/stellar-cyber)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, AWS, GCP, Azure, Oci, Prometheus, Grafana, Loki, Alertmanager, Elasticsearch, MongoDB, Spark, Kafka, Redis, Python
**Posted:** 2026-05-06

> Senior SRE responsible for operating and improving highly available cloud platforms, distributed data systems, observability, incident response, and deployment automation. Requires 5+ years of SRE, DevOps, or platform engineering experience, advanced Kubernetes expertise, cloud proficiency, and strong Python and Bash skills.

## Job Description

## Responsibilities
- Administer and maintain container orchestration platforms and containerized workloads.
- Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability.
- Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms.
- Administer and optimize cloud-based environments across multiple providers.
- Manage and support distributed data platforms and real-time processing systems.
- Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments.
- Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability.
- Automate and orchestrate infrastructure using programming and scripting languages.
- Perform system administration and networking tasks to support internal and external environments.
- Collaborate effectively with engineers and stakeholders across different time zones.

## Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
- Proven success leading large-scale production systems in cloud environments, including AWS, GCP, Azure, or OCI.
- Demonstrated leadership in driving incident response, on-call best practices, and a reliability-focused culture.
- Strong experience with production on-call operations and incident management.
- Advanced proficiency in Kubernetes administration and troubleshooting.
- Hands-on experience with Prometheus, Grafana, Loki, and Alertmanager.
- Knowledge of chat-based operations interfaces and/or auto-remediation controllers using AI agentic frameworks.
- Understanding of AI agents for auto-triaging alerts, correlating signals, and suggesting root-cause hypotheses.
- Expertise in operating data platforms including Elasticsearch, MongoDB, Spark, Kafka, and Redis.
- Proficiency with public cloud services such as AWS, Azure, GCP, or OCI.
- Strong programming and automation skills in Python and Bash.
- Deep understanding of Infrastructure as Code, distributed systems, databases, networking, and Linux administration.
- Experience with Terraform, Helm, and CI/CD pipelines using GitHub Actions, Bitbucket, or ArgoCD.
- Excellent problem-solving, communication, and leadership abilities.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.

## Nice-to-haves
- Certifications in AWS, GCP, observability, Linux, or Kubernetes.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Database Administrator - Core Infrastructure](https://hotfix.jobs/jobs/84ebc795-fb55-4998-94cf-dd630cc51e8a) - Kraken - Remote
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/6d7b2812-de3f-4dfa-8586-9010e1e594e7) - Clickhouse - Remote
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/56b82d25-159d-41f5-94fb-fb1f71dd6a5c) - Clickhouse - Remote

**Apply:** https://hotfix.jobs/jobs/92bc314a-111d-4492-809a-7b3b73761a19
**Canonical:** https://hotfix.jobs/jobs/92bc314a-111d-4492-809a-7b3b73761a19