Site Reliability Engineer - US Government
Site Reliability Engineer builds, operates, and maintains scalable infrastructure for air-gapped production environments, focusing on Linux servers, cloud/on-prem systems, automation, and troubleshooting. Requires 4+ years Linux admin experience, active security clearance, and proficiency in programming/scripting.
About the job
Core Responsibilities
- Maintaining availability of cloud & physical Linux servers that power the Palantir platform in air-gapped production environments
- Design, deploy, and operate infrastructure to support customer & product requirements via modern orchestration & monitoring platforms.
- Collaborate closely with product teams on requirements & SLOs for deploying software into air-gapped environments.
- Identifying, troubleshooting, and solving network & systems issues
- Scripting to automate away routine operational tasks
- Provide technical troubleshooting support for production issues, ensuring timely resolution and minimal impact on operations. Participate in a support on-call schedule
What We Value
- Confidence in troubleshooting complex systems issues independently using stack traces and observability & systems tools
- Comfort with managing large scale production systems and technologies with configuration management, load balancing, monitoring & alerting infrastructure, and container orchestration
- Demonstrated ability to continuously learn and work independently, making decisions with minimal supervision while working in secure facilities
- Experience with containers (Docker/Podman) and orchestration (OpenShift/Kubernetes) at scale is a plus
- Preferred Certifications: DOD 8570 IAT Level II or greater (CISSP, Sec+), Unix/Linux Computing Environment (e.g Linux+, RHCE)
What We Require
- Active security clearance
- 4+ years of experience with Linux system administration (RHEL or equivalent preferred)
- Experience with cloud-based hosting platforms like AWS, Azure, or GCP and/or experience with hardware-based environments
- Familiarity with monitoring systems using tools like Prometheus and writing health checks
- Proficiency with at least one programming language, such as Java, Go, Python, JavaScript, Bash, or similar languages.
- Strong engineering background, preferred in fields such as Computer Science, Mathematics, Software Engineering, Physics, and Data Science
Skills
Linux, Rhel, Kubernetes, Openshift, Docker, Podman, AWS, Azure, GCP, Prometheus, Python, Go, Java, Bash, JavaScript
Similar jobs
DevOps / SRE jobsDesigns and operates foundational developer-infrastructure services for CI, builds, deployments, and testing. The role requires senior-level systems engineering, end-to-end service ownership, and cross-functional technical leadership.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.