Senior Site Reliability Engineer, Colorado Springs
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.
About the job
Responsibilities
- Own the reliability, scalability, and security of production applications and platforms.
- Design, implement, and manage monitoring, logging, and alerting using tools such as Prometheus, Loki, Alloy, and Grafana.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), alerting, and error budgets.
- Lead incident response, serve as incident commander when needed, and conduct blameless postmortems and After Action Reviews.
- Build automation to eliminate operational toil and improve deployment and management of on-premise systems.
- Partner with platform and application teams to design and operate secure Kubernetes clusters and AWS cloud/on-premise environments.
- Embed RMF and STIG security and compliance controls into infrastructure automation.
- Share best practices for air-gapped environments and production readiness.
Requirements
- Active Top Secret clearance; SCI eligibility is a plus.
- 5+ years of experience in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
- Experience with incident response, root-cause analysis, and continuous improvement.
- Expertise with Terraform or CloudFormation and Ansible.
- Experience designing, deploying, and operating Kubernetes environments.
- Experience building and maintaining CI/CD pipelines with GitLab CI/CD, Jenkins, or GitHub Actions.
- Proficiency in at least one of Python, Go, or Bash.
- Familiarity with AWS or AWS GovCloud.
- Experience with observability tools such as the Grafana stack, ELK stack, or Datadog.
- Understanding of networking fundamentals, core protocols, and secure configurations.
- Ability to work on-site at customer locations in Colorado Springs, Colorado; relocation assistance is available for candidates outside commuting distance.
Nice-to-Haves
- Experience in Department of Defense environments and compliance frameworks such as RMF, STIGs, or ICD 503.
- GitOps practices and toolchains.
- Security-minded design for sensitive environments.
- Experience designing meaningful SLIs and SLOs for complex distributed systems.
- Familiarity with on-premise virtualization platforms such as VMware, Proxmox, Nutanix, or Hyper-V.
- Service mesh experience with Istio or Linkerd.
- Certifications such as AWS DevOps Engineer, CKA, or CKAD.
- Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain one within three months of employment.
Skills
Kubernetes, Terraform, Ansible, AWS, Aws Govcloud, Prometheus, Grafana, Loki, Gitlab Ci/Cd, Jenkins, GitHub Actions, Python, Go, Bash, Datadog
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.