Skip to content
OnebriefOnebrief

Senior Site Reliability Engineer, Colorado Springs

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.

About the job

Responsibilities

  • Own the reliability, scalability, and security of production applications and platforms.
  • Design, implement, and manage monitoring, logging, and alerting using tools such as Prometheus, Loki, Alloy, and Grafana.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), alerting, and error budgets.
  • Lead incident response, serve as incident commander when needed, and conduct blameless postmortems and After Action Reviews.
  • Build automation to eliminate operational toil and improve deployment and management of on-premise systems.
  • Partner with platform and application teams to design and operate secure Kubernetes clusters and AWS cloud/on-premise environments.
  • Embed RMF and STIG security and compliance controls into infrastructure automation.
  • Share best practices for air-gapped environments and production readiness.

Requirements

  • Active Top Secret clearance; SCI eligibility is a plus.
  • 5+ years of experience in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
  • Experience with incident response, root-cause analysis, and continuous improvement.
  • Expertise with Terraform or CloudFormation and Ansible.
  • Experience designing, deploying, and operating Kubernetes environments.
  • Experience building and maintaining CI/CD pipelines with GitLab CI/CD, Jenkins, or GitHub Actions.
  • Proficiency in at least one of Python, Go, or Bash.
  • Familiarity with AWS or AWS GovCloud.
  • Experience with observability tools such as the Grafana stack, ELK stack, or Datadog.
  • Understanding of networking fundamentals, core protocols, and secure configurations.
  • Ability to work on-site at customer locations in Colorado Springs, Colorado; relocation assistance is available for candidates outside commuting distance.

Nice-to-Haves

  • Experience in Department of Defense environments and compliance frameworks such as RMF, STIGs, or ICD 503.
  • GitOps practices and toolchains.
  • Security-minded design for sensitive environments.
  • Experience designing meaningful SLIs and SLOs for complex distributed systems.
  • Familiarity with on-premise virtualization platforms such as VMware, Proxmox, Nutanix, or Hyper-V.
  • Service mesh experience with Istio or Linkerd.
  • Certifications such as AWS DevOps Engineer, CKA, or CKAD.
  • Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain one within three months of employment.

Skills

Kubernetes, Terraform, Ansible, AWS, Aws Govcloud, Prometheus, Grafana, Loki, Gitlab Ci/Cd, Jenkins, GitHub Actions, Python, Go, Bash, Datadog

Runpod

Runpod

United States

Senior HPC Storage Engineer
$180k+/yrRemote8+ YOEDevOps / SRE

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Lightning AI

Lightning AI

New York, NY

Senior Infrastructure Software Engineer
$180k+/yrHybrid8+ YOEDevOps / SRE

Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.