Skip to content
AxleAxleFrederick, MD

Site Reliability Engineer

Site Reliability Engineer modernizing a multi-cloud (AWS/Azure/GCP) environment into a scalable, observable Kubernetes-based platform using DevOps/SRE practices, AIOps, IaC, and AI-driven automation to support scientific and clinical research programs. Requires 6+ years SRE/DevOps experience with strong Linux, IaC, observability, and scripting skills.

140k – 155k/yr
On-site6+ YOEDevOps / SRE

About the role

Responsibilities

  • Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and OpenTelemetry tools.
  • Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements.
  • Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments.
  • Automate workload onboarding and offboarding processes, ensuring standardization and governance.
  • Track system ownership, dependencies, and lifecycle states for operational transparency.
  • Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact.
  • Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments.
  • Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP).
  • Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows.
  • Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards.
  • Develop “golden paths” and standardized platform templates for consistent workload deployment.
  • Automate provisioning, patching, configuration management, and environment lifecycle.
  • Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms.
  • Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights.
  • Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations.
  • Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering.
  • Build developer-friendly platforms (“golden paths”) that simplify deployments, reduce friction, and improve velocity.
  • Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads.
  • Build and manage containerized and orchestrated platforms (Docker, Kubernetes).
  • Support cloud migration, modernization, and platform standardization initiatives.
  • Ensure systems meet security, compliance, backup, and disaster recovery requirements.
  • Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities.
  • Stay abreast of new technologies in areas including AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering etc.

Requirements

  • 6+ years experience in DevOps / SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads.
  • 4+ years of hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management and administration support.
  • 4+ years of experience automating Infrastructure-as-Code (IaC) deployments to Amazon AWS, Google GCP or Microsoft Azure.
  • 4+ years with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions.
  • Strong scripting skills (Python, Bash, PowerShell or similar).
  • Proficiency using vibe coding and coding assistants to develop scripts, tools and applications for the DevOps and SRE use cases.
  • Proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory.
  • Willingness to learn new technologies, adopt and adapt to emerging technologies or needs from a project to a project.
  • Cloud certifications preferred.
  • Certifications in Grafana, Splunk, Docker, Kubernetes preferred but optional.

Nice-to-Haves

  • Experience optimizing infrastructure for AI/ML workloads.
  • Familiarity with zero-trust principles and multi-cloud environments (AWS, Azure, GCP).
  • ITIL process knowledge.

Skills

SREDevOpsKubernetesDockerTerraformAnsibleJenkinsGitHub ActionsSplunkGrafanaOpenTelemetryPrometheuselkPythonLinux

Similar roles

DevOps / SRE jobs
ZoomInfo

Senior Software Engineer

ZoomInfoUnited States

Build and operate production-critical GitOps deployment platforms, shared service tooling, and infrastructure automation in Go and TypeScript. The role requires 8+ years of software engineering experience plus expertise with Argo CD, Helm, Kubernetes, cloud platforms, and scalable APIs.

140k – 220k/yrRemote8+ YOEDevOps / SRE
Pump.co

DevOps Engineer

Pump.coSan Francisco, CA

Hands-on DevOps role owning AWS infrastructure, building developer tooling, and driving technical roadmap at an early-stage YC startup. Requires 6+ years infra/DevOps experience and strong AWS/K8s/Terraform skills.

140k – 200k/yrOn-site6+ YOEDevOps / SRE
Forterra

Senior Network Systems Engineer

ForterraEast Palo Alto, CA +2

Deploys, operates, and troubleshoots network infrastructure including routers, switches, Linux appliances, and AWS resources for edge-deployed communications in DDIL environments. Requires 5+ years network engineering experience, Linux proficiency, IaC automation, and 50% domestic travel.

140k – 185k/yrHybrid5+ YOEDevOps / SRE
Scrunch

Senior Infrastructure Engineer

ScrunchNew York, NY +16

Senior Infrastructure Engineer designs, builds, and operates cloud infrastructure, developer tooling, observability, and reliability systems at scale, primarily on GCP. Requires high-velocity dev experience, IaC, database scaling, workflow orchestration, and production Python coding.

140k – 200k/yrRemoteDevOps / SRE
Clickhouse

Senior Infrastructure Engineer - Postgres

ClickhouseUnited States

Senior Infrastructure Engineer owns reliability, operations, and automation for ClickHouse's Postgres integration across multi-cloud environments. Requires 7+ years SRE/DevOps experience, Postgres expertise, Terraform/Kubernetes proficiency, and strong Go skills.

140k – 230k/yrRemote7+ YOEDevOps / SRE