Join Onebrief's Infrastructure & Security team as an SRE focused on improving application reliability directly in the TypeScript codebase for mission-critical military planning software. Requires active Secret clearance, 5+ years shipping application code, strong observability and incident response skills, and willingness to work onsite in Arlington, VA.
180k – 220k/yr
Hybrid5+ YOEDevOps / SRE
About the role
What You'll Do
Improve the application by working directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. Partner with product engineers on design decisions and review code with reliability and security in mind.
Build observability that developers actually use: Design and run monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). Create alerts and dashboards tied to real application behavior.
Own reliability targets: Define and measure SLIs and SLOs, wire up alerting, and prove what "reliable" means for the systems with data.
Lead incident response: Act as incident responder and commander when needed. Run blameless post-mortems (AARs) that identify root causes and turn them into code or process fixes.
Automate away toil: Identify repetitive operational work and write software to eliminate it. Share solutions with other teams, including those in air-gapped environments.
Requirements
Active Secret clearance
5+ years in software engineering, SRE, or related role, with real time spent writing and shipping application code
Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript)
Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage
Experience with incident response, root cause analysis, and turning findings into lasting fixes
Collaborator who works well across product, platform, and DevOps teams and shares context openly
Technical Expertise
Application development in TypeScript (Node and/or a modern front-end framework)
CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins)
Testing and quality practices as part of the delivery process
Comfort with at least one of Python, Go, or Bash for tooling and automation
Working knowledge of containers and Kubernetes (enough to debug and deploy)
Networking fundamentals and secure configuration basics
Nice-to-Haves
Observability: Grafana stack, ELK, or Datadog
Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud)
Kubernetes cluster design and operations
Designing meaningful SLIs/SLOs with error budgets for distributed systems
GitOps practices and toolchains
DoD environments and compliance frameworks (RMF, STIGs, ICD 503)
Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.
180k – 260k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure
AcuityMDBoston, MA
Senior Infrastructure Engineer responsible for building and operating platform primitives including Kubernetes, CI/CD, observability, and developer tooling at a high-growth AI and data platform company.
180k – 250k/yrHybrid5+ YOEDevOps / SRE
Sr. Software Engineer, DevOps
AKASASouth San Francisco, CA
Builds and scales reliable infrastructure for SaaS applications using Kubernetes, Terraform, and GitHub CI/CD. Focuses on observability with Grafana/Prometheus, automation to reduce toil, production troubleshooting, and cross-team collaboration. Requires 5+ years Python experience.
180k – 220k/yrHybrid5+ YOEDevOps / SRE
Senior DevEx Engineer
ReplitFoster City, CA
Steward Replit's TypeScript monorepo, Go services, and developer tooling to accelerate engineering velocity and reduce friction. Partner with AI team to enhance agent-generated code, requiring senior-level expertise in build systems and large-scale codebases.
180k – 250k/yrHybridDevOps / SRE
Network Development Engineer, ML Infrastructure (High-Speed Interconnects)
xAIPalo Alto, CA
Designs, builds, and optimizes high-speed copper and optical interconnects for large-scale AI/ML clusters. Requires 8+ years experience in high-speed networking, deep knowledge of SerDes, photonics, and Master's/PhD in EE/Photonics/Physics.