# Site Reliability Engineer

**Company:** [Axle](https://hotfix.jobs/companies/axle)
**Location:** Frederick, MD
**Role:** DevOps / SRE
**Salary:** $140k – $155k/yr
**Experience:** 6+ years
**Skills:** SRE, DevOps, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitHub Actions, Splunk, Grafana, OpenTelemetry, Prometheus, elk, Python, Linux
**Posted:** 2026-07-30

> Site Reliability Engineer modernizing a multi-cloud (AWS/Azure/GCP) environment into a scalable, observable Kubernetes-based platform using DevOps/SRE practices, AIOps, IaC, and AI-driven automation to support scientific and clinical research programs. Requires 6+ years SRE/DevOps experience with strong Linux, IaC, observability, and scripting skills.

## Job Description

## Responsibilities
- Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and OpenTelemetry tools.
- Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements.
- Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments.
- Automate workload onboarding and offboarding processes, ensuring standardization and governance.
- Track system ownership, dependencies, and lifecycle states for operational transparency.
- Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact.
- Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments.
- Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP).
- Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows.
- Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards.
- Develop “golden paths” and standardized platform templates for consistent workload deployment.
- Automate provisioning, patching, configuration management, and environment lifecycle.
- Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms.
- Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights.
- Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations.
- Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering.
- Build developer-friendly platforms (“golden paths”) that simplify deployments, reduce friction, and improve velocity.
- Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads.
- Build and manage containerized and orchestrated platforms (Docker, Kubernetes).
- Support cloud migration, modernization, and platform standardization initiatives.
- Ensure systems meet security, compliance, backup, and disaster recovery requirements.
- Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities.
- Stay abreast of new technologies in areas including AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering etc.

## Requirements
- 6+ years experience in DevOps / SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads.
- 4+ years of hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management and administration support.
- 4+ years of experience automating Infrastructure-as-Code (IaC) deployments to Amazon AWS, Google GCP or Microsoft Azure.
- 4+ years with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions.
- Strong scripting skills (Python, Bash, PowerShell or similar).
- Proficiency using vibe coding and coding assistants to develop scripts, tools and applications for the DevOps and SRE use cases.
- Proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory.
- Willingness to learn new technologies, adopt and adapt to emerging technologies or needs from a project to a project.
- Cloud certifications preferred.
- Certifications in Grafana, Splunk, Docker, Kubernetes preferred but optional.

## Nice-to-Haves
- Experience optimizing infrastructure for AI/ML workloads.
- Familiarity with zero-trust principles and multi-cloud environments (AWS, Azure, GCP).
- ITIL process knowledge.

## Similar roles

- [Senior Software Engineer](https://hotfix.jobs/jobs/906e022e-4e8b-4773-984a-7ab3022e3a59) - ZoomInfo - Remote - $140k – $220k/yr
- [DevOps Engineer](https://hotfix.jobs/jobs/3fae50b7-bb1e-40a6-81dc-d263d7391dc1) - Pump.co - San Francisco, CA - $140k – $200k/yr
- [Senior Network Systems Engineer](https://hotfix.jobs/jobs/1b0a2131-90b5-452d-abcc-3601f5ded0b4) - Forterra - East Palo Alto, CA - $140k – $185k/yr
- [Senior Infrastructure Engineer](https://hotfix.jobs/jobs/0c56c557-0f7b-49c3-8e76-b8b2242895c0) - Scrunch - Remote - $140k – $200k/yr
- [Senior Infrastructure Engineer - Postgres](https://hotfix.jobs/jobs/9fe70ed5-6992-4761-81b2-6fd1540bd5ab) - Clickhouse - Remote - $140k – $230k/yr

**Apply:** https://hotfix.jobs/jobs/9c80499b-cf69-4ede-bc6b-5a18bc5e7463
**Canonical:** https://hotfix.jobs/jobs/9c80499b-cf69-4ede-bc6b-5a18bc5e7463