Skip to content
TwentyTwenty

Forward Deployed Site Reliability Engineer

On-site SRE ensuring reliability of mission-critical platform in air-gapped AWS environment at government site. Defines SLOs/SLIs, leads incident response, manages deployments with Docker/Terraform, and liaises between customer and engineering team. Requires 5+ years SRE experience and TS/SCI clearance.

About the job

Responsibilities

Reliability Engineering

  • Define, track, and report on SLIs and SLOs for platform services running in the customer environment.
  • Use error budgets to drive reliability conversations with the engineering team, translating operational data into prioritized engineering work.
  • Identify and eliminate toil: build automation for repetitive operational tasks within the constraints of the secure environment.
  • Conduct post-incident reviews, own root cause analysis, and drive durable fixes in partnership with the engineering team.

Observability & Incident Response

  • Own the observability posture for the on-site deployment — dashboards, alerting thresholds, and log pipelines using the LGTM stack (Grafana, Loki, Tempo, Mimir).
  • Lead incident response on-site: triage, containment, coordination with engineering team, and customer communication.
  • Maintain and continuously improve runbooks for operational procedures and emergency response protocols.
  • Serve as the on-call anchor for the customer environment, with clear escalation paths to the engineering team.

Deployment & Infrastructure Operations

  • Work with the customer deployment team to get platform stood up and updated within the restricted environment.
  • Manage containerized services (Docker, Docker Compose) across deployment lifecycle — configuration, updates, rollbacks.
  • Apply and validate Terraform-based infrastructure changes within the enclave, in coordination with the DSO engineer.
  • Perform capacity planning and flag scaling requirements to the engineering team before they become incidents.

Customer Liaison & Engineering Feedback

  • Serve as the primary technical interface between the government customer and engineering team — translating operational requirements, constraints, and issues in both directions.
  • Represent the operational environment accurately in engineering discussions.
  • Partner with the DevSecOps engineer on compliance, logging, and audit requirements specific to the customer environment.
  • Provide technical guidance and support to customer stakeholders on system behavior and troubleshooting procedures.

Requirements

Must Have

  • 5+ years of professional experience in site reliability engineering, production operations, or a closely related infrastructure role.
  • Proven experience defining and tracking SLIs, SLOs, and error budgets in a production environment.
  • Hands-on experience with Docker, Docker Compose, and AWS (EC2, ECS, RDS, VPCs, security groups) in production deployments.
  • Solid Linux/Unix systems administration skills; productive in constrained environments where GUI tooling may be limited or unavailable.
  • Experience with Terraform for infrastructure provisioning and configuration, working within DSO-provided policy guardrails.
  • Experience with the LGTM observability stack or equivalent (Grafana, Loki, Prometheus/Mimir, distributed tracing).
  • Strong incident response experience: you've led responses, written post-mortems and runbooks, and shipped the preventive fix.
  • Scripting proficiency in Python or Bash for operational automation, with familiarity in Go a plus; experience with PagerDuty or equivalent on-call tooling.
  • Experience working in or directly supporting government or defense environments, including air-gapped or enclave deployments.
  • Must possess and be able to maintain a TS/SCI security clearance with appropriate polygraph.
  • U.S. citizenship required.
  • Willingness to travel occasionally for customer engagements and operational support.

Nice To Have

  • Experience with NATS or similar pub/sub messaging systems in production.
  • Background in cyber operations, intelligence systems, or signals environments.
  • AWS certifications (Solutions Architect, SysOps, or DevOps Engineer).

Skills

SLOs, Slis, Docker, Docker Compose, AWS, Terraform, Linux, Grafana, Loki, Tempo, Mimir, Python, Bash, Pagerduty

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.