# Forward Deployed Site Reliability Engineer

**Company:** [Palantir](https://hotfix.jobs/companies/palantir)
**Location:** Washington, DC
**Role:** DevOps / SRE
**Experience:** 4+ years
**Skills:** Linux, rhel, Kubernetes, openshift, Docker, podman, Prometheus, Python, Go, Java, Bash, JavaScript
**Posted:** 2025-05-29

> Forward Deployed Site Reliability Engineer responsible for building, operating, and maintaining scalable infrastructure in air-gapped on-prem environments for US Government customers. Requires 4+ years Linux admin experience, hardware/networking knowledge, scripting skills, 50% travel availability, and active Top Secret clearance.

## Job Description

## Core Responsibilities
- Maintaining availability of physical Linux servers that power the Palantir platform in air-gapped production environments
- Design, deploy, and operate infrastructure to support customer & product requirements via modern orchestration & monitoring platforms
- Collaborate closely with product teams on requirements & SLOs for deploying software into air-gapped environments
- Identifying, troubleshooting, and solving network & systems issues
- Scripting to automate away routine operational tasks
- Provide technical troubleshooting support for production issues, ensuring timely resolution and minimal impact on operations. Participate in a support on-call schedule

## What We Value
- Confidence in troubleshooting complex systems issues independently using stack traces and observability & systems tools
- Comfort with configuration management, load balancing, monitoring & alerting infrastructure, and container orchestration on small hardware form factors
- Demonstrated ability to continuously learn and work independently, making decisions with minimal supervision while working in secure facilities
- Experience with containers (Docker/Podman) and orchestration (OpenShift/Kubernetes) at scale is a plus
- Preferred Certifications: DOD 8570 IAT Level II or greater (CISSP, Sec+), Unix/Linux Computing Environment (e.g Linux+, RHCE)

## What We Require
- Available for 50% travel (domestic and international)
- 4+ years of experience with Linux system administration (RHEL or equivalent preferred)
- Experience with hardware environments, including setup, configuration, and management of physical servers and networking equipment
- Familiarity with monitoring systems using tools like Prometheus and writing health checks
- Proficiency with at least one programming or scripting language, such as Java, Go, Python, JavaScript, Bash, or similar languages
- Strong engineering background, preferred in fields such as Computer Science, Mathematics, Software Engineering, Physics, and Data Science
- Active US Security Clearance at or above the Top Secret level

## Similar roles

- [API Deployment Manager](https://hotfix.jobs/jobs/3ef6b6ec-4b8c-4fb4-8c5b-9d761d738d10) - Runway - Remote - $145k – $270k/yr
- [Software Engineer, Infrastructure, Interpretability](https://hotfix.jobs/jobs/b167b8e0-6f36-44a7-92b2-7803a2869b2a) - Anthropic - San Francisco, CA - $320k – $485k/yr
- [Software Engineer - Continuous Delivery](https://hotfix.jobs/jobs/0ceaa410-df08-45d6-80b3-68328c769297) - Baseten - San Francisco, CA - $165k – $330k/yr
- [Platform Engineer II](https://hotfix.jobs/jobs/867b3afd-1415-4d60-86ca-26742bdd629b) - Bestow - Remote - $115k – $130k/yr
- [DevOps Engineer](https://hotfix.jobs/jobs/0ba7c9d4-0bad-447f-a491-48b3469ea0e4) - Fusion Health - Woodbridge, NJ - $120k – $140k/yr

**Apply:** https://hotfix.jobs/jobs/0fbdc6ce-9878-4c74-aa9e-080b3ca268d5
**Canonical:** https://hotfix.jobs/jobs/0fbdc6ce-9878-4c74-aa9e-080b3ca268d5