Skip to content
QuindarQuindarArvada, CO

Site Reliability Engineer, US Gov

Build and operate secure, highly available cloud infrastructure for mission-critical government and space systems across AWS GovCloud and C2E environments. The role requires Kubernetes, Terraform, Python, observability, networking, compliance, and an active U.S. security clearance.

160k – 200k/yr
Hybrid3+ YOEDevOps / SRE

About the role

Responsibilities

  • Architect, automate, test, deploy, and maintain highly available cloud infrastructure in AWS GovCloud and AWS C2E, emphasizing security, compliance, and operational excellence.
  • Ensure adherence to SOC 2, NIST 800-171, and FedRAMP Moderate controls while supporting cleared environments and workflows.
  • Manage deployment and ongoing maintenance of Quindar deployments at government sites.
  • Define best practices for availability, latency, and performance across services.
  • Drive Quindar’s readiness and deployment onto AWS C2E, including control implementation, continuous container hardening, deployment configuration, and networking requirements.
  • Participate in incident management, a 24/7 on-call rotation, and technical troubleshooting of complex enterprise and mission-critical systems.
  • Create systems and automation that reduce manual intervention.
  • Collaborate with frontend, backend, and flight/mission operations engineers to ensure the Quindar system meets required performance metrics.

Requirements

  • Deep experience with Kubernetes, containerized workloads, and serverless architectures.
  • Expertise managing Kubernetes clusters with AWS EKS or Rancher and integrating AWS cloud services.
  • Hands-on experience supporting GovCloud, IL-enclave deployments, or C2E environments.
  • Experience with observability stacks such as Grafana LGTM and Datadog.
  • Proficiency in Python and Terraform or similar infrastructure-as-code tooling.
  • Strong understanding of VPNs, NLBs/ALBs, HTTPS, TLS, VPC peering, and CDN integration.
  • Knowledge of API services, distributed NoSQL and relational databases, caching systems, event-driven architectures, and multi-tier systems.
  • Experience with task automation and CI/CD pipeline development; GitLab Workflows preferred.
  • Understanding of Unix/Linux operating systems.
  • Knowledge of cloud security best practices, enclave boundary protections, enclave-to-enclave interconnects, and cost-efficient architectures.
  • Experience with identity and access management, including Auth0, Keycloak, AWS IAM, or ICAM patterns.
  • Strong Git fundamentals.
  • Bachelor’s degree in Computer Science or a related field.
  • 3+ years of professional experience as an SRE, DevOps, reliability, infrastructure, or platform engineer.
  • Active U.S. security clearance; Secret or higher required, with TS/SCI preferred.
  • U.S. citizenship required.

Nice to Have

  • Experience working toward ATO or authorization in federal, DoD, or IC environments.
  • Experience supporting deployments in GovCloud, C2S/C2E, or IL-enclave environments.
  • Experience managing software product deployment across multiple classification levels.

Compensation and Benefits

  • $160,000–$200,000 annual salary.
  • Unlimited PTO with a required minimum of 15 days off.
  • Most U.S. federal government holidays.
  • Quarterly health and wellness benefits.
  • Comprehensive health insurance for employees and families, with 100% employee coverage.
  • 4% 401(k) matching.
  • Annual four-day company offsite.

Skills

AWSaws govcloudaws eksKubernetesrancherTerraformPythongrafana lgtmDatadogCI/CDgitlab workflowsLinuxaws iamvpc peeringtls

Similar roles

DevOps / SRE jobs
Office Hours

Platform Engineer

Office HoursSan Francisco, CA +1

Senior Platform Engineer owning AWS/EKS infrastructure, IaC with Terraform, GitOps, security, observability, and data systems for a fast-growing expert network marketplace. Requires 5+ years production Kubernetes/AWS experience, strong IaC and security skills.

160k – 180k/yrRemote5+ YOEDevOps / SRE
Sola

Software Engineer, Desktop Automation

SolaNew York, NY

Build and own desktop automation execution platform integrating AI agents with Windows sessions via remote protocols like VNC/RDP. Requires deep systems knowledge in OS APIs, accessibility frameworks, and automation infrastructure for reliable enterprise workflows.

160k – 300k/yrRemoteDevOps / SRE
Hebbia

Software Engineer, Site Reliability

HebbiaNew York, NY +1

Site Reliability Engineer who owns production services end-to-end, writes production code, builds observability and internal tooling, and embeds with product teams to improve reliability and performance. Requires 5+ years experience and strong systems programming skills.

160k – 300k/yrOn-site5+ YOEDevOps / SRE
Rad AI

Software Engineer, Infrastructure (All Levels)

Rad AISan Francisco, CA

Designs, builds, and operates scalable cloud infrastructure on AWS with Kubernetes and serverless tech to support AI healthcare products. Requires 4+ years experience in cloud-native platforms, IaC, automation, and reliability practices for regulated environments.

160k – 300k/yrOn-site4+ YOEDevOps / SRE
Hebbia

Platform Engineer, Document Intelligence

HebbiaNew York, NY +1

Platform Engineer building high-scale distributed document indexing and search systems for an AI platform serving top financial institutions. Requires 5+ years experience in backend/distributed systems with Python/Java/Go and cloud infrastructure.

160k – 300k/yrOn-site5+ YOEDevOps / SRE