Forward Deployed Site Reliability Engineer (TS/SCI Required)
Forward Deployed Site Reliability Engineer responsible for owning reliability, observability, incident response, and deployments of a mission-critical platform in a restricted, air-gapped government AWS environment. Requires 5+ years SRE/production ops experience, strong Linux/Docker/Terraform skills, LGTM stack proficiency, and active TS/SCI clearance.
Salary not listed
On-site5+ YOEDevOps / SRE
About the role
What You'll Do
Reliability Engineering
Define, track, and report on SLIs and SLOs for platform services running in the customer environment.
Use error budgets to drive reliability conversations with the Arlington engineering team, translating operational data into prioritized engineering work.
Identify and eliminate toil: build automation for repetitive operational tasks within the constraints of the secure environment.
Conduct post-incident reviews, own root cause analysis, and drive durable fixes in partnership with the engineering team.
Observability & Incident Response
Own the observability posture for the on-site deployment — dashboards, alerting thresholds, and log pipelines using the LGTM stack (Grafana, Loki, Tempo, Mimir).
Lead incident response on-site: triage, containment, coordination with Arlington, and customer communication.
Maintain and continuously improve runbooks for operational procedures and emergency response protocols.
Serve as the on-call anchor for the customer environment, with clear escalation paths to the engineering team.
Deployment & Infrastructure Operations
Work with the customer deployment team to get Twenty's platform stood up and updated within the restricted environment.
Apply and validate Terraform-based infrastructure changes within the enclave, in coordination with the DSO engineer who owns IaC policy and guardrails.
Perform capacity planning and flag scaling requirements to the Arlington team before they become incidents.
Customer Liaison & Engineering Feedback
Serve as the primary technical interface between the government customer and Twenty's engineering team — translating operational requirements, constraints, and issues in both directions.
Represent the operational environment accurately in engineering discussions: what the team in Arlington can't see, you make visible.
Partner with the DevSecOps engineer on compliance, logging, and audit requirements specific to the customer environment.
Provide technical guidance and support to customer stakeholders on system behavior and troubleshooting procedures.
Must Have
5+ years of professional experience in site reliability engineering, production operations, or a closely related infrastructure role.
Proven experience defining and tracking SLIs, SLOs, and error budgets in a production environment.
Hands-on experience with Docker, Docker Compose, and AWS (EC2, ECS, RDS, VPCs, security groups) in production deployments.
Solid Linux/Unix systems administration skills; productive in constrained environments where GUI tooling may be limited or unavailable.
Experience with Terraform for infrastructure provisioning and configuration, working within DSO-provided policy guardrails.
Experience with the LGTM observability stack or equivalent (Grafana, Loki, Prometheus/Mimir, distributed tracing).
Strong incident response experience: you've led responses, written post-mortems and runbooks, and shipped the preventive fix.
Scripting proficiency in Python or Bash for operational automation, with familiarity in Go a plus; experience with PagerDuty or equivalent on-call tooling.
Experience working in or directly supporting government or defense environments, including air-gapped or enclave deployments.
Nice To Have
Experience with NATS or similar pub/sub messaging systems in production.
Background in cyber operations, intelligence systems, or signals environments.
AWS certifications (Solutions Architect, SysOps, or DevOps Engineer).
Security Requirements
Must possess and be able to maintain a TS/SCI security clearance with appropriate polygraph.
U.S. citizenship required.
Willingness to travel occasionally for customer engagements and operational support.
Benefits
Health: Medical, dental, and vision plan options. Life / AD&D, disability coverage options.
Family: Paid parental leave for eligible full-time employees. 12 weeks for birthing parents, 4 for non-birthing parents, 6 weeks for adoptive, foster, or intended parents through surrogacy.
Vacation: Paid holidays and flexible PTO. Take what you need.
Retirement: 401(k) with pre-tax and Roth options. HSA/FSA options, dependent care FSA.
Benefits vary by location, role, and eligibility. Full plan details provided during the interview and offer process.
Skills
site reliability engineeringslisloerror budgetsDockerdocker composeAWSLinuxTerraformGrafanalokimimirPrometheusIncident Responsepython scripting
Build and operate Kubernetes-based compute and runtime infrastructure powering AI search, assistant, and agent workloads across multi-cloud environments. Own reliability, scalability, cost-efficiency, and on-call for production platform services.
140k – 220k/yrHybrid5+ YOEDevOps / SRE
Software Engineer, Cloud Infrastructure
GleanMountain View, CA
Software Engineer building automated customer deployment systems, IaC templates, and complex cloud networking setups (AWS/GCP) for Glean's Work AI platform. Requires 4+ years experience in cloud infrastructure, Terraform, and strong networking fundamentals.
Build reliable, productized systems for customer cloud setup and deployment across AWS and GCP. The role requires 4+ years of software engineering experience, backend fundamentals, infrastructure-as-code expertise, and strong networking knowledge.
200k – 270k/yrHybrid4+ YOEDevOps / SRE
Infrastructure engineer
WriterNew York, NY +2
Infrastructure Engineer building and operating scalable, reliable production systems for an enterprise AI platform. Owns end-to-end reliability, automates with Python/Go, integrates AI agents into workflows, leads incident response, and collaborates cross-functionally on high-availability infrastructure using Kubernetes, Terraform, and multi-cloud tooling. Requires 5+ years experience and daily use of AI tooling.
140k – 274k/yrHybrid5+ YOEDevOps / SRE
Site Reliability Engineer
ConductorOnePortland, ME
Own reliability and scalability of a horizontal identity platform, including core cloud and FedRAMP environments. Build observability, automate operations, lead incident response, and ensure new features are reliable from the start. Requires production SRE experience at scale with Kubernetes, IaC, and strong programming skills.